10,000 Matching Annotations
  1. Aug 2026
    1. Author response:

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #1 (Public review):

      (1) The study used a WNK643 inhibitor as the only tool to manipulate WNK1-4 activity. This inhibitor seems selective; however, it has been reported that it exhibits different efficiency in inhibiting the individual WNK kinases among each other (e.g. PMID: 31017050, PMID: 36712947). Additionally, the authors do not analyze nor report the expression profiles or activity levels of WNK1, WNK2, WNK3, and WNK4 within the relevant brain regions (i.e. hippocampus, cortex, amygdala). Combined, these weaknesses raise concerns about the direct involvement of WNK kinases within the selected brain regions and behavior circuits. It would be beneficial if the authors provided gene profiling for WNK1, 2, 3, and -4 (e.g. using Allen brain atlas). To confirm the observations, the authors should either add results from using other WNK inhibitors or, preferentially, analyze knock-down or knock-out animals/tissue targeting the single kinases.

      Revisions 1: The authors added Fig. S1A during the revisions to show expression of Wnk1-4. While the expression data from humans is interesting, the experimental part of the study is performed in mice. It would be more informative for the authors to add expression profiles from mice or overview the expression pattern with suitable references in the introduction to address this point. The authors did not add data from knock down or knockout tissue targeting the single kinases.

      Thank you for the excellent suggestion. We have added mouse in situ hybridization data curated from Allen Brain Atlas and found mRNA encoding WNK1 and WNK2 highly expressed in the hippocampus compared to WNK3 and WNK4. We also have included WNK1 knockdown data from cell lines (Figure S7A-F).

      Whole body WNK1 knockout is embryonically lethal, and we do not have access to brain tissue specific WNK knockout animal models. In addition, knockout of WNKs from brain tissue samples from animals is not very efficient from our experience and therefore, we included data from cell lines. In other non-neuronal cell lines, WNK1 knockdown replicates the effect of WNK463 (Figure S7A-D). However, in SHSY5Y cells, WNK1 knockdown did not replicate the effects of WNK463 on pAKT levels (Figure S7EF). This suggests tissue-specific effects of WNKs, and it also supports our suggestion that cooperativity among WNK family members is required in neuronal cells. This further supports our conclusion that WNK463 is an ideal tool to test our hypothesis in this study as it targets all 4 WNKs (WNK1-4) and furthermore, WNK463 was reported in the literature to inhibit only the four WNKs out of more than 400 kinases tested, indicating more selectivity than many small molecules used to target other enzymes.

      (2) The authors do not report any data on whether the global inhibition of WNKs affects insulin levels as such. Since the authors demonstrate the synergistic effect of simultaneous insulin treatment and WNK1-4 inhibition, such data are missing.

      Revisions 1: The authors added Fig. S5A to address this point. It is appreciated that authors performed the needed experiment. Unfortunately, no significant change was found, therefore, the authors still cannot conclude that they demonstrate a synergistic effect of simultaneous insulin treatment and WNT1-4 inhibition. It is a missed opportunity that the authors did not measure insulin in the CSF or tissue lysate to support the data.

      Thank you for the comment. As suggested, we tried to measure insulin in mouse hippocampal tissue lysate, and the levels fell way below the detectable range (78 - 5000 pg/mL) of the mouse insulin detection kit (Abcam: AB285341) used.

      (3) The study discovered that the Sortilin receptor binds to OSR1, leading the authors to speculate that Sortilin may be involved in the insulin-dependent GLUT4 surface trafficking. The authors conclude in the result section that "WNK/OSR1/SPAK influences insulin-sensitive GLUT4 trafficking by balancing GLUT4 sequestration in the TGN via regulation of Sortilin with GLUT4 release from these vesicles upon insulin stimulation via regulation of AS160." However, the authors do not provide any evidence supporting Sortilin's involvement in such regulation, thus, this conclusion should be removed from the section. Accordingly, the first paragraph of the discussion should be also rephrased or removed.

      Revisions 1: The authors added Fig. 5M-N to address this point. The new experiment is appreciated. However, the authors still do not show that sortilin is involved in insulin or WNK-dependent GLUT4 trafficking in their set up since the authors do not demonstrate any changes in GLUT4 sorting or binding. The conclusions should therefore be rephrased or included purely in the discussion. Moreover, the discussion was not adjusted either, leading to over interpretation based on the available data.

      Thank you for the suggestion. The conclusion has been rephrased as suggested.

      (4) The background relevant to Figure 5, as well as the results and conclusions presented in Figure 5 are quite challenging to follow due to the lack of a clear introduction to the signaling pathways. Consequently, understanding the conclusions drawn from the data is also difficult. It would be beneficial if the authors addressed this issue with either reformulations or additional sections in the introduction. Furthermore, the pulldown experiments in this figure lack some of the necessary controls.

      Revisions 1: The Authors insufficiently addressed this point during the revisions and did not rewrite the introduction as suggested.

      The background information related to figure 5 has been simplified as suggested. Response regarding the controls used is provided in the response to critique 5 as below.

      (5) The authors lack proper independent loading controls (e.g. GAPDH levels) in their immunoblots throughout the paper, and thus their quantifications lack this important normalization step. The authors also did not add knock-out or knock-down controls in their co-IPs. This is disappointing since these improvements were central and suggested during the revision process.

      GAPDH has been used as a loading control wherever applicable for Western blots on lysates (see Figures: 5D, 3F, 4E, 4C) In other cases, such as in Figure 5E, the analysis of pAS160 is normalized to total AS160 as this is more appropriate compared to GAPDH. For IP experiments such as Figure 5G, GAPDH is not an applicable control as we are using purified protein fragments in this case. For IP experiments (Figure 5K, 5L, 5C), IP proteins have been normalized to the input protein levels serving as a loading control for the IP because GAPDH is an intracellular protein which necessarily is not pulled down along with the proteins being IP’ed. Therefore, in this case, GAPDH is not a valid loading control. The choice of our loading controls used are very well supported by previous publications from our lab and other labs working on WNK pathways.

      (6) The schemes that represent only hypotheses (Fig. 1K, 4A) are unnecessary and confusing and thus should be omitted or placed at the end of each figure if the conclusions align.

      Thank you for the suggestion. Figure 1K is already at the end of the figure 1 and it shows the conclusion of that figure. Figure 4A have been placed at the end of the figures as suggested. Other schemes are only added at the end of the figures as suggested.

      (7) Low-quality images, such as Fig. 5H should be replaced with high-resolution photos, moved to the supplementary, or omitted.

      Thank you for your comment. The suggested images have been replaced with higher resolution ones.

      Reviewer #2 (Public review):

      This study by Jaykumar and colleagues seeks to expand the field's appreciation of insulin responses in the brain, specifically by implicating WNK kinase function in various neuronal responses, ranging from behavioral / memory changes to GLUT4 trafficking to the cell surface with subsequent glucose uptake. This revised study is now comprehensive and presents a logical and reasonably documented cascade of molecular interactions responsible in part for GLUT4 trafficking under the regulation of WKK and insulin. Additional data allow the authors to dissect a plausible WNK/OSR1/SPAK-sortilin pathway for the modulation of GLUT4 trafficking, in part by capitalizing on a overlay of various techniques and systems. The data - much of it in vivo or ex vivo - showing a potential role for WNK function in brain glucose utilization remains a compelling part of the story, with the dissection of the signaling cascade and a potential role for sortilin in mediating WNK function via effects on GLUT4 cellular localization now more convincing.

      Initially, the group shows that oral WNK463 treatment - an inhibitor of WNKs broadly - in mice augments a number of memory readouts. These findings fit within the context of the overall story the authors present: that WNK function is critical to brain glucose utilization, which impacts learning. Multiple approaches are used to show that WNK463 treatment, i.e. inhibition of WNKs, increases glucose uptake, including labeled 2deoxyglucose uptake in vivo in the brain and in isolated synaptosome, and uptake in ex vivo hippocampal slices. These findings are solid and consistent. With the exception of some relatively minor comments regarding the data presentation made to the authors and now fully addressed, the findings showing that WNK463 treatment increases GLUT4-mediated glucose uptake and surface localization of GLUT4 are reasonable, with the hippocampal slice data being particularly relevant.

      While the details of the WNK signaling cascade is dense, in the revised application one clearly appreciates the molecular interrogation and interactions the group is dissecting, supported by the use of multiple models. With the additional findings, these systems and the data now reinforce each other, presenting a strongly documented overall story.

      A limitation of the study with the initial submission was the authors' reliance upon a single pharmacological tool (WNK463) to inhibit WNK kinases. WNK463 apparently has substantial specificity for WNKs and WNK463 treatment lessened OSR1 phosphorylation (a WNK substrate). Nevertheless, the cohesiveness of the findings in terms of the broader pathway engagement (GLUT4 trafficking, glucose uptake) is consistent with the author's proposed mechanisms and conclusions. The authors have additionally addressed this concern in the revised manuscript with more information supporting the specificity of WNK463 as well as the multiple approaches to confirm the effect of WNK463 on the WNK signaling pathway of interest.

      The final few paragraphs of the discussion that weave the author's findings into the field more broadly, including Sortilin function and neurological disorders, are appreciated. Additional clarity in the Methods section is also helpful.

      Thank you for the positive response and acknowledging that we have satisfactorily addressed all of your critiques.

    1. eLife Assessment

      This manuscript presents a valuable computational tool for identifying 3-5 gene regulatory network topologies capable of generating oscillatory dynamics. The application of Monte Carlo Tree Search to circuit design is novel and effectively expands the scale at which non-linear behaviours can be explored in silico. The efficiency of the proposed algorithm is convincing, and the work will be of interest to the systems and synthetic biology communities. While the generality of the identified circuit properties is constrained by the simplifying modelling assumptions and parameter choices, the methodological contribution represents a significant advance in the field.

    2. Joint Public Reviews:

      This manuscript presents an algorithm for identifying network topologies that exhibit a desired qualitative behaviour, with a particular focus on oscillations. The approach is first demonstrated on 3-node networks-where results can be validated through exhaustive search-and then extended to 5-node networks, where the search space becomes intractable. Network topologies are represented as directed graphs, and their dynamical behaviour is classified using stochastic simulations based on the Gillespie algorithm. To efficiently explore the large design space, the authors employ reinforcement learning via Monte Carlo Tree Search (MCTS), framing circuit design as a sequential decision-making process.

      This work meaningfully extends the range of systems that can be explored in silico to uncover non-linear dynamics and represents a valuable methodological advance for the fields of systems and synthetic biology.

      Strengths:

      The evidence presented is strong and compelling. The authors validate their results for 3-node networks through exhaustive search, and the findings for 5-node networks are consistent with previously reported motifs, lending credibility to the approach. The use of reinforcement learning to navigate the vast space of possible topologies is both original and effective and represents a novel contribution to the field. The algorithm demonstrates convincing efficiency, and the ability to identify robust oscillatory topologies is particularly valuable. Expanding the scale of systems that can be systematically explored in silico marks a significant advance for the study of complex gene regulatory networks.

      Weaknesses:

      Although the proposed approach substantially expands the scale of tractable searches, the systems explored remain relatively small, being limited to five-node networks. The authors now discuss possible avenues for improving scalability, but extending the framework to substantially larger networks remains an important future challenge.

      Another important limitation concerns the assumption of identical reaction rates for all circuit connections. As the authors' own analysis shows, relaxing this assumption leads to significant qualitative and quantitative changes in oscillatory dynamics. Consequently, it remains unclear how the properties of the identified fault-tolerant oscillators translate to more biologically realistic regulatory circuits, where kinetic parameters vary across interactions.

      The conclusions should also be interpreted within the chosen modelling framework and parameter space. In particular, the sampled parameter ranges and restriction to relatively low Hill coefficients define the subset of regulatory architectures explored. Whether broader parameter regimes, including higher Hill coefficients, would reveal additional oscillatory architectures remains unclear.

    3. Author response:

      The following is the authors’ response to the original reviews.

      Joint Public Review:

      (1) The principal weakness of the manuscript lies in the interpretation of biological robustness. The authors identify network topologies that sustain oscillatory behaviour despite perturbations to the system or parameters. However, in many cases, this persistence is due to the presence of partially redundant oscillatory motifs within the network. While this observation is interesting and of clear value for circuit design, framing it as evidence of evolutionary robustness may be misleading. The “mutant” systems frequently exhibit altered oscillatory properties, such as changes in frequency or amplitude. From a functional cellular perspective, mere oscillation is insufficient — preservation of specific oscillation characteristics is often essential. This is particularly true in systems like circadian clocks, where misalignment with environmental cycles can have deleterious effects. Robustness, from an evolutionary standpoint, should therefore be framed as the capacity to maintain the functional phenotype, not merely the qualitative behaviour. 

      We agree with the reviewers that our framing conflated qualitative robustness (continued oscillation) with functional robustness (oscillation with preservation of properties such as frequency), and that the latter is a more meaningful definition of robustness in an evolutionary context. We have edited the manuscript to remove any suggestions that our results can explain the multiple-oscillator architecture of circadian clocks. We have also added a paragraph in the discussion that highlights the distinction between qualitative and functional robustness, as well as the other ways in which our training environment differs from an evolutionary context.

      Locations of changes:

      Abstract, second-to-last sentence

      Results, final section, first paragraph

      Discussion, first paragraph

      Discussion, new paragraph

      (2) A secondary limitation is that, despite the methodological advances, the scale of the systems explored remains modest. While moving from 3- to 5-node systems is non-trivial, five elements still represent a relatively small network. It is somewhat surprising that the algorithm does not scale further, particularly when considering the performance of MCTS in other domains — for instance, modern chess engines routinely explore far larger decision trees. A discussion on current performance bottlenecks and potential avenues for improving scalability would be valuable.

      We thank the reviewers for raising this important point. We have edited the manuscript to specify that we faced two distinct bottlenecks in scaling our experiments. The first is the runtime and scaling of the underlying Gillespie simulations, which become much more expensive as circuit size increases. The second is our use of the original (“vanilla”) MCTS algorithm, without the deep-learning value and policy networks that have driven the dramatic gains in domains such as Go and chess. We have added a new Discussion paragraph that identifies these two bottlenecks, followed by a paragraph on future methodological enhancements that goes into more detail about potential improvements to the algorithm. Using a power-law extrapolation of the current scaling, we estimate that without further methodological improvements the largest tractable network is approximately 7 nodes, while AlphaZero-style scaling could plausibly extend the approach to roughly 19-node circuits. We have also added an order-of-magnitude estimate of the 5-node search space (≈2×10<sup>9</sup> topologies) to give the reader a more concrete picture of the current scale.

      Locations of changes:

      Results, final section, end of paragraph 1 (Gillespie runtime as the practical bottleneck)

      Discussion, new paragraph 4 (bottlenecks and possible improvements)

      Discussion, new paragraph 5 (projected scaling under deep-learning-based extensions)

      Introduction, paragraph 4, and Discussion, paragraph 1 (search-space size)

      Methods, new section “Estimation of search space size”

      (3) It is worth noting that the emergence of oscillations in a model often depends not only on the topology but also critically on parameter choices and the nature of the nonlinearities. The use of Hill functions and high Hill coefficients is a common strategy to induce oscillatory dynamics. Thus, the reported results should be interpreted within the context of the modelling assumptions and parameter regimes employed in the simulations.

      We agree that the modeling assumptions substantially impact the interpretation of the results, and we have expanded the description of our modeling framework to make these assumptions explicit. To clarify, our model does not use Hill equations directly. Instead, cooperative binding is represented as a sequential, mass-action binding process. In addition to a new Methods section explaining our model in more detail, we have added a Methods section to show analytically that the effective Hill coefficient in our system is always ≤2. It also mentions an important practical benefit of using sequential binding rather than explicit TF dimerization, which is that it improves the size and scaling behavior of the reaction system. 

      Locations of changes:

      Results, section 2, paragraph 1 (clarification that no Hill function is imposed; effective Hill coefficient ≤2)

      Methods, section 1, new subsection “Sequential binding yields Hill coefficients ≤2”

      Recommendations for the authors:

      (4) It would be helpful to include the explicit reaction equations and corresponding reaction rates used in the simulations, to facilitate reproducibility and better understanding of the modelling assumptions.

      We have added the explicit reaction equations, the corresponding rate parameters, and the bounds used during random parameter sampling. The Results section describing our model now reports the rate parameters used in the simulations shown in the paper. Additionally the new subsection at the beginning of the Methods presents the full stochastic model. Finally, the bounds used for parameter sampling are now included in-line in the corresponding Methods subsection. The original tables of parameter values and sampling bounds (Tables 1 and 2) have been retained.

      Locations of changes:

      Results, section 2, paragraph 1 (rate parameters in main text; Table 1 retained)

      Methods, section 1, new subsection “Stochastic model of a transcription factor network”

      Methods, section “Random sampling” (explicit parameter bounds; Table 2 retained)

      (5) Sustained oscillations are notoriously difficult to observe in Gillespie simulations due to stochastic noise. Could the authors comment on whether simulation times were sufficiently long to distinguish sustained oscillations from transient dynamics?

      We agree this is an important methodological point. We have clarified that simulations were run for 11.1 hours. Because nearly all the oscillators we found had periods below 100 minutes and most oscillators had periods below 12 minutes, almost all were observed for at least 6.6 cycles, which we believe is sufficient to distinguish sustained oscillations from transient dynamics.

      Locations of changes:

      Results, section 2, paragraph 2

      (6) While the manuscript alludes to broader applications of the proposed method, it would be beneficial to elaborate on these possibilities. Clarifying how this tool could be extended to other types of dynamical behaviours or biological questions would strengthen the impact.

      We have rewritten the final paragraph of the Discussion to give concrete illustrative examples of how the framework can be extended beyond oscillator design, including the discovery of more complex design principles and the design of synthetic multicellular circuits such as multi-cell type cancer therapies and morphogenetic patterning circuits.

      Locations of changes:

      Discussion, final paragraph

      (7) It remains somewhat unclear whether the contribution lies primarily in the novel application of reinforcement learning or whether there are methodological innovations within the algorithm itself. Are there existing tools with similar objectives, and how does this work improve upon them in terms of performance or capabilities?

      We have edited the Discussion to state more directly that the principal contribution of this work is the novel application of reinforcement learning to the problem of network topology design problem, rather than a fundamentally new RL algorithm. We also contextualize CircuiTree relative to other topology-search approaches and articulate where the use of RL provides a concrete advantage in navigating large combinatorial search spaces.

      Locations of changes:

      Discussion, end of paragraph 3

      (8) Further details on the training of the reinforcement learning algorithm would be appreciated. Was it trained a priori, and if so, how much data was required?

      CircuiTree is not pre-trained. Like AlphaZero, it learns exclusively from simulations performed during the search itself, with no externally curated dataset or a priori domain knowledge. We have clarified this in two places in our manuscript.

      Locations of changes:

      Introduction, beginning of first paragraph

      Results, section 1, paragraph 4

      (9) Some discussion on how computational time scales with increasing network size would be valuable.

      As mentioned above, we have added (i) an order-of-magnitude estimate of the size of the 5-node search space (≈2×10⁹ topologies) to make the current scale concrete; (ii) a power-law extrapolation of the current algorithm’s scaling that suggests ≈7 nodes is the largest tractable network without further improvements; and (iii) a Discussion paragraph projecting that AlphaZero-style scaling, combined with faster simulators, could plausibly extend the approach to ≈19-node circuits. The methodology for estimating the search-space size is described in a new Methods section.

      Locations of changes:

      Introduction, paragraph 4 (search-space size)

      Discussion, paragraph 1 (search-space size)

      Discussion, new paragraph 4 (extrapolated scaling and 7-node bound)

      Discussion, new paragraph 5 (projected scaling under deep-learning extensions)

      Methods, new section “Estimation of search space size”

      (10) The discussion of knockouts that dampen or restore oscillations is interesting. Was this analysis performed systematically, or were the examples selected through observation? Clarifying the methodology here would add rigour to the interpretation.

      We have clarified that the topologies highlighted in this analysis were selected by observing the fault-tolerant properties of a few topologies that exhibited high mutational robustness in our screen, rather than via a systematic analysis of the screen results. The text now states this explicitly so the reader can interpret the examples accordingly.

      Locations of changes:

      Results, final section, paragraph 3

      Other changes made by the authors

      In addition to the changes above, we have made several minor edits to improve clarity and presentation. We corrected spelling and formatting errors throughout, and improved the clarity of language in the Discussion section. No changes were made to the Figures, Tables, Algorithms, or Supplementary Information.

      We are grateful to the reviewers for their input and believe the manuscript has been substantially improved as a result. We hope the revised version meets with the editors’ and reviewers’ approval.

    1. eLife Assessment

      This study reports the discovery of a new circuit mechanism for light-avoidance behavior in the marine annelid, Platynereis dumerilii. Using calcium imaging, molecular perturbations, behavioral measurements, and modeling, the authors provide compelling evidence that nitric oxide is released by postsynaptic neurons onto ciliary photoreceptors to prolong and enhance their response to ultraviolet light. The fundamental new role of nitric oxide described in this study may be conserved across animal phyla and thus will be of broad interests to neuroscientists and neuroendocrinologists.

    2. Reviewer #1 (Public review):

      Summary:

      The ciliary photoreceptor cells and its downstream neurons of larval annelid must be orchestrated in a specific pattern to promote downward swimming in response to long duration of UV exposure. The authors first conducted neuroanatomical examination of the circuit to identify NOS-expression neurons (INNOS) that are immediately downstream to the ciliary photoreceptor cells. The INNOS is activated by UV and produce NO. The NOS is required for UV avoidance by Platynereis larvae and neural dynamics of the photoreceptor cells and their downstream circuit. Following up the RNA-seq data with in-situ hybridization experiments, the authors found that two unconventional guanylate cyclases, NIT-GC1 and NIT-GC2, are expressed and localized in different subcellular domain of the photoreceptor cells. Experiments using the culture cells ang genetically encoded sensors demonstrated that NIT-GC1 can generate cGMP in response to nitric oxide. Finally, authors build mathematical model that fit the live imaging data and used it to predict how the magnitude of the photoreceptor activation varied by intensity and duration of UV light.

      Strengths:

      The authors conducted comprehensive interrogations of the UV avoidance pathway at the molecular and circuit levels and constructed mathematical model. The main conclusions are supported with layers of evidence from different assays.

      Weaknesses:

      The authors addressed these weaknesses in the previous version of the manuscript. Statistics are missing in both figure legends and methods. The perturbations of genes and molecules were not cell-type-specific and therefore the observed behavioral defect could be attributed to the malfunction of the circuit elsewhere not examined in this study. I suggest adding more explanation about the functions of other NOS-expressing cells and conducting a control experiment to test behavioral response to a non-visual stimulus.

    3. Reviewer #2 (Public review):

      Summary:

      This study is quite thorough, tackling this NO-dependent UV avoidance circuit with both breadth and depth. There are several novel discoveries throughout, but the whole package represents perhaps even more than the sum of these parts.

      Strengths:

      The presentation of the work is compelling. The introduction sets up the question and the state of the field very nicely. The discovery of the non-canonical NO receptor pathway in the ciliary photoreceptors is fascinating and will likely open up new avenues for future research into NO-pathways in different species. The use of genetic and pharmacological manipulations of circuit components was well thought-out. The authors applied different experimental techniques expertly throughout the study so that they could develop a comprehensive view from the molecular to the behavioral levels.

      Weaknesses:

      The authors have done an excellent job revising and explaining their model. No important weaknesses remain, in my opinion.

    4. Reviewer #3 (Public review):

      The transition from planktonic to benthic depends upon several physical and chemical cues. Nitric oxide (NO) is known as a critical player in the induction of larval metamorphosis in several invertebrates. Although NO is a widespread signalling molecule in a broad range of organisms regulating key physiological processes, internal regulatory mechanisms studies are scarce. While the UV sensing in larvae of the annelid Platynereis dumerilii using ciliary photoreceptors has been studied, the neuronal signalling mechanism remains unknown. In this study, Kei Jokura et al. investigated how annelid Platynereis dumerilii larvae detect UV sensing and modulate swimming behaviour through nitric oxide feedback. Using existing resources of Platynereis larval connectome/volume EM data, they identified NOS-expressing interneurons within the ciliary photoreceptors circuit (cPRCs). They demonstrated that NO is produced in cPRCs during UV/violet stimulation by using a fluorescent NO-reporter line. Further, they demonstrated that Nitric oxide signalling mediates UV-avoidance behaviour by using NOS-mutant larvae. Finally, they mapped out the signalled mechanisms of the cPRC circuit using published spatially mapped single-cell transcriptome data of Platynereis larvae, the Ca sensor lines, in situ HCR, and immunostaining. Additionally, by using their findings from Ca imagining data of cPRC, INNOS and INRGWa cells collected in wild-type, NOS knockout and NIT-GC2 morphant larvae, Kei Jokura et al. developed a mixed cellular-circuit-level mathematical model. However, my expertise in mathematical modelling is limited, so I cannot comment on this section.

      Comments on revised version.

      Thank you for the opportunity to re-evaluate this manuscript. I have reviewed the authors' responses and the revised manuscript. The authors have carefully and satisfactorily addressed all of my previous comments and concerns. The revisions have strengthened the paper, and I have no further suggestions.

    5. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public Review):

      Summary:

      The ciliary photoreceptor cells and its downstream neurons of larval annelid must be orchestrated in a specific pattern to promote downward swimming in response to long duration of UV exposure. The authors first conducted neuroanatomical examination of the circuit to identify NOS expression neurons (INNOS) that are immediately downstream to the ciliary photoreceptor cells. The INNOS is activated by UV and produces NO. The NOS is required for UV avoidance by Platynereis larvae and neural dynamics of the photoreceptor cells and their downstream circuit. Following up the RNA-seq data with in situ hybridization experiments, the authors found that two unconventional guanylate cyclases, NIT-GC1 and NIT-GC2, are expressed and localized in different subcellular domain of the photoreceptor cells. Experiments using the culture cells and genetically encoded sensors demonstrated that NIT-GC1 can generate cGMP in response to nitric oxide. Finally, authors build a mathematical model that fit the live imaging data and used it to predict how the magnitude of the photoreceptor activation varied by intensity and duration of UV light.

      Strengths:

      The authors conducted comprehensive interrogations of the UV avoidance pathway at the molecular and circuit levels, and constructed a mathematical model. The main conclusions are supported by layers of evidence from different assays.

      Weaknesses:

      Statistics are missing in both figure legends and methods. The perturbations of genes and molecules were not cell-type-specific and therefore the observed behavioral defect could be attributed to the malfunction of the circuit elsewhere not examined in this study. I suggest adding more explanation about the functions of other NOS-expressing cells and conducting a control experiment to test behavioral response to a non-visual stimulus.

      Thank you for this assessment of our work. We have now added additional panels with statistical tests to the figures and included explanatory text in the figure legends.

      Regarding the cell-type-specific effects, we would like to offer a more nuanced view of this. Some of the genes we studied (NIT-GC1 and NIT-GC2) are only expressed in 4 cells in an organism of ~10,000 cells (as we demonstrated by HCR, immunostainings and the analysis of single-cell RNAseq data) and we knocked-down these genes (validated by antibody staining) with two independent morpholinos (to be able to rule out off-target effects). This is as cell-type specific as it gets. NOS is also expressed in a very limited number of cells. In the larval stages, we detected NOS expression only in the four INNOS cells and the pigmented eyes. We could previously show that UV avoidance is only mediated by the cPRC circuit (including INNOS) whereas phototaxis is exclusively mediated by the pigmented eyes Verasztó et al. (2018). Due to this behavioural specificity, we are therefore confident that the effects of the NOS mutations on UV avoidance are due to the lack of NOS from the INNOS cells and not the pigmented eyes. Besides UV avoidance, we have characterised the speed of ciliary swimming without a light stimulus, as well as phototaxis []. In addition, we also tested phototaxis and detected a reduced phototactic reaction in NOS mutants, but not after chemically inhibition of NO production (Figure 3B and 3F). The additional effects of NO are thus well documented in the paper and independent of the function of NO in the UV reaction. In addition, we have now did further quantifications and added new data (Figure 3 – figure supplement 4) to show that NOS mutants show normal lunar periodicity of sexual maturation, similar to wild-type animals. NOS mutants are viable and fertile, are feeding, building tubes and are mating as wild-type animals (these behaviours were not quantified here, we only show the data for lunar periodicity).

      Reviewer #2 (Public Review):

      Summary:

      This study is quite thorough, tackling this NO-dependent UV avoidance circuit with both breadth and depth. There are several novel discoveries throughout, but the whole package represents perhaps even more than the sum of these parts.

      Strengths:

      The presentation of the work is compelling. The introduction sets up the question and the state of the field very nicely. The discovery of the non-canonical NO receptor pathway in the ciliary photoreceptors is fascinating and will likely open up new avenues for future research into NO pathways in different species. The use of genetic and pharmacological manipulations of circuit components was well thought-out. The authors applied different experimental techniques expertly throughout the study so that they could develop a comprehensive view from the molecular to the behavioral levels.

      Weaknesses:

      While I appreciate the intent of bringing together a large set of measurements from connectomics and calcium imaging in the framework of a model, the model seemed rather poorly constrained. How many parameters are in the model shown in Figure 6A? How many of them are well constrained by experimental measurements? The authors also don’t perform sensitivity analysis on the parameters of the model. And ultimately, the conclusion over the model in Figure 7 is somewhat trivial within the unitless construction: larger amplitude and longer duration stimuli lead to increased activation of the downstream neuron thought to lead to the downward swim behavior. I could imagine that a large family of models would arrive at this same result, and without units, there is no way to really test it with new behavioral experiments.

      We thank the reviewer for these comments. We have now thoroughly revised the model based on new experiments and carried out a sensitivity analysis. We are also more positive about the usefulness of the model though, for the following reasons.

      General usefulness of the model: With the model we can now reproduce all the qualitative dynamics of the circuit. The modelling also completely changed the way we were thinking about the system. For example, we needed to include a time-limited step in the cPRC transduction cascade leading to NOS activation to capture the time-invariance of the peak. In the future, we can also use this model to generate prediction e.g. about what the response to repeated stimulations can be. In the revised version, we included a new series of measurements of calcium dynamics in the cPRCs under varying duration and intensity of UV light. These important new results gave us further insights and necessitated a revision of the model.

      Unitless model: The model is indeed unitless, since already our input data from calcium imaging represent normalised data and the model was fit to these data. We would need a lot more information to build up a proper ground-up biophysical model (e.g. including capacitance, ionic concentrations etc.). This does not mean that the model is not useful.

      Sensitivity analysis: We have created a pipeline to carry out local and global sensitivity analysis and carried out a variance-based sensitivity and identifiability analysis. These data and code are included in the revised version. In the model, most parameters are identifiable. If we fix only two of the parameters all other parameters can be identified.

      Constrains and lots of possible models: In terms of the constrains, since we don’t have dimensioned quantitites, our constraints are bounds on the parameters. However, the structural identifibiality analysis tells us where we can identify parameters and can guard against sloppiness and overfitting. We only fitted the model on a subset of the data and can reproduce dynamics on a larger set of data. For important cellular interactions, the signs of the interactions are constrained. The model thus also allows us to rule out a large number of models - e.g. we could rule out a very simple model of progression from step to step as it was not possible to fit the data to such a model.

      We have updated the text to reflect these changes, e.g.: “Many of the couplings in our model are constrained (e.g. UV leads to INNOS activation) and e.g. reversing the sign of some of these couplings would not arrive at the same result. Several earlier variants of the model could not be fit to the data, The model is thus well constrained by our physiological experiments and the 106 circuit map.”

      Reviewer #3 (Public Review):

      The transition from planktonic to benthic depends upon several physical and chemical cues. Nitric oxide (NO) is known as a critical player in the induction of larval metamorphosis in several invertebrates. Although NO is a widespread signalling molecule in a broad range of organisms regulating key physiological processes, internal regulatory mechanisms studies are scarce. While the UV sensing in larvae of the annelid Platynereis dumerilii using ciliary photoreceptors has been studied, the neuronal signalling mechanism remains unknown. In this study, Kei Jokura et al. investigated how annelid Platynereis dumerilii larvae detect UV sensing and modulate swimming behaviour through nitric oxide feedback. Using existing resources of Platynereis larval connectome/volume EM data, they identified NOS-expressing interneurons within the ciliary photoreceptors circuit (cPRCs). They demonstrated that NO is produced in cPRCs during UV/violet stimulation by using a fluorescent NO-reporter line. Further, they demonstrated that Nitric oxide signalling mediates UV-avoidance behaviour by using NOS-mutant larvae. Finally, they mapped out the signalled mechanisms of the cPRC circuit using published spatially mapped single-cell transcriptome data of Platynereis larvae, the Ca sensor lines, in situ HCR, and immunostaining. Additionally, by using their findings from Ca imagining data of cPRC, INNOS and INRGWa cells collected in wild-type, NOS knockout and NIT-GC2 morphant larvae, Kei Jokura et al. developed a mixed cellular-circuit-level mathematical model. However, my expertise in mathematical modelling is limited, so I cannot comment on this section.

      No doubt, the study has been conducted extensively. However, I have a few comments, please see below.

      Page 4: “In contrast, both two- and three-day-old homozygous NOS-mutant larvae showed a strongly diminished UV avoidance response (Figure 3A, B and Figure 3-figure supplement 1B, C).” Instead of using subjective terms like “strongly,” it would be more relevant to provide statistical values. However, I could not locate any means of statistical analysis on larval behaviour. Can the authors indicate the statistical values for all behaviour studies?

      We thank the reviewer for these comments. We have changed the wording and also added the results of statistical analyses to the figures and explanations to the figure legends.

      Page 5: “(D) Vertical displacement in 30 sec bins of wild type and mutant (NOSΔ11/Δ11 and NOSΔ23/Δ23) three-day-old larvae stimulated with 395 nm light from the side, 488 nm light from the top and 395 nm light from the top.” The error bars for WT are too long at the end of the experiment. It is not clear how the authors decided to use this time frame. Did the authors try carrying this out for an extended time period? How did the authors decide on 120 seconds as the time frame for exposure? Authors should provide data on larval behaviour for an extended time.

      The 120 seconds time frame of exposure takes into account the reaction time and swimming speed of the larvae as well as the size of the assay chamber (160 mm water height).

      By the end of a 120 stimulation many larvae tend to accumulate at the bottom of the chamber due to downward swimming and cannot be further tracked. This effect leads to higher variability in the data towards the end of the experiment in the wild-type batches. We have showed both continuous vertical data as well as the binned data. The 30 sec bin was chosen for convenience and for better comparison with our previous paper on UV avoidance behaviour (Verasztó et al. 2018)

      Page 13: “During the UV response, prototroch cilia beat slower than trunk cilia, resulting in a head down stable state (‘rear-wheel drive’). In contrast, during the pressure response prototroch cilia beat faster than trunk cilia, leading to a head-up orientation (‘front-wheel drive’). Testing this hypothesis will require biophysical experiments and mathematical modelling.” Authors should carry out ciliary beating analysis under UV light in the current study with NOS mutant larvae. Since the pressure and UV detection systems are closely related, comparing the difference in ciliary beating is important to 155 demonstrate this hypothesis. Further, did the authors check the Ca sensor GCaMP6s under pressure conditions?

      We thank the reviewer for this suggestion. We have carried out further experiments and analysed the ciliary beat frequency (CBF) of larvae exposed to UV stimulation. We added these data to Figure 3—figure supplement 3. The results (increase of CBF under UV in wild-type but not NOS mutant larvae) were quite surprising to us and falsified our initial hypothesis. We have rewritten the discussion to reflect this important new finding.

      The response of cPRC cilia to changes in hydrostatic pressure has been extensively documented in our recent paper on the mechanism of barotaxis in the Platynereis larva (see Bezares-Calderón et al., https://doi.org/10.1101/2023.02.28.530398).

      Page 18: “strips. One strip contained UV (395 nm) LEDs (SMB1W-395, Roithner Lasertechnik) and the other infrared (810 nm) LEDs (SMB1W-810NR-I, Roithner Lasertechnik).” Authors should test larval swimming behaviour at different wavelengths. Even though they are performed in previous work, the experiment with different wavelengths is necessary to be conducted in NOS mutant larvae in parallel with a control. This will confirm that NOS is principally associated with UV. Further, to demonstrate that this mechanism is associated with ciliary movement, authors need to provide this evidence.

      The diving reaction by non-directional light can only be induced by UV/cyan light, as we have shown previously, and it is mediated by a single UV-opsin photopigment (c-opsin1). The avoidance experiments can thus only be done with UV/cyan light. We also measured swimming 174 behaviour with 480 nm directional light to test phototaxis (Figure 3D and Figure 3—figure 175 supplement 1F).

      Recommendations for the authors:

      Reviewer #1 (Recommendations For The Authors):

      (1) The current introduction focuses on the nitric oxide signaling. It would be helpful for readers to have an introductory section about Platynereis larvae (a total number of neurons, etc) and rationales to use its nervous system as a model to study mechanisms of NO signaling and gating of visual response.

      We added an extra paragraph to give more detail on the number of cells in the larva and why Platynereis larvae are interesting to study to understand how synaptic and volume signalling interact. “The 3-day-old larva has over 9,000 cells classified into 202 neuronal and 92 non neuronal cell types (Verasztó et al., 2025). The synapticly connected subset of the cells in the body form a connectome of over 2,000 cells. Besides synapses, neurons in the larva also signal by volume transmission mediated by a rich repertoire of neuropeptides (Williams et al., 2017) and other modulators (Bauknecht and Jékely, 2017). The transparent and experimentally accessible Platynereis larvae could therefore inform how synaptic and volume signalling interact to mediate behaviour (Jékely and Yuste, 2024).”

      (2) The readers would want to know why these larvae swim downward in response to UV and upward to 480nm light. Is that for maintaining the certain depth from the surface of the water?

      It has been suggested that the ciliary photoreceptor circuit, which senses UV light, and the rhabdomeric photoreceptor circuit, which senses blue light, exchange messages with each other and that the two work together to form a depth gauge. By allowing larvae to swim at their preferred depth, the depth gauge influences where they end up when they become adults.

      (3) Are there splicing isoforms of NOS in the Platynereis dumerilii? If so, do antibodies and probes for in situ distinguish them? In Drosophila, truncated isoforms can inhibit the function of full length isoform, and therefore it was important to use methods to distinguish spicing isoforms.

      We did not identify any alternatively spliced forms of Platynereis NOS in our published transcriptome resources.

      (4) “INRGWs” acronym appears without explanation in the first paragraph of the results.

      Corrected: “the INRGWa neurons (cholinergic interneurons that express an RGW neuropeptide)”

      (5) In Figure 1C, it is difficult to see the projection patterns of individual cell types. Figure supplement 3 can be combined with the current Figure 1.

      We have moved one panel from Figure supplement 3 to the main Figure 1 (panel D) to show the INNOS projections more clearly.

      (6) Add more explanation about NOSp::palmi-3xHA reporter. Is it membrane-targeted reporter with the upstream promotor sequence of NOS? Or is it endogenous NOS that was tagged with palmi-3xHA?

      It is a membrane-targeted reporter driven by the upstream promoter sequence of the NOS gene. The construct is delivered by plasmid injection. We have clarified the description in the text and the figure legend (“The four apical organ cells, but not the eyes, were also labelled with a transiently expressed NOS-reporter transgene. This transgene contains the upstream promoter sequence of the NOS gene that drives a membrane-targeted palmitoylated tdTomato reporter (Figure 1F).” and “Expression of a membrane-targeted reporter driven by the NOS regulatory region (NOSp::palmi-3xHA-Tomato; magenta”). More details are in the Methods section under Transient transgenesis.

      (7) Show lack of anti-NOS immunostaining in NOS mutant to warrant specificity of the antibody. The subcellular localization of NOS in the dendritic arbors of INNOS is essential for the proposed model. The immunostaining image in Figure 4 -figure supplement 3D can be in the main figure. Is NOS also in the axons of INNOS?

      We further optimised the immunostaining with the NOS antibodies in WT and NOS mutant and added the staining data to the main Figure 1G and Figure 3—supplement 1. NOS was observed to be clearly localised in a region in the neuropil corresponding to the dendritic site of the INNOS cells. The staining was completely absent in larvae of both NOS knockout alleles. We have also added these explanations to the text.

      (8) “that was defective in NOS mutant (Figure 5I)” should be corrected as “Figure 5H”.

      We have restructured this part of the text and figure, the data from the Ser-h1 cells are now in Figure 5 - figure supplement 2.

      (9) Related to Figure 7, how do the larvae respond to a sequence of UV light (e.g. ten times 0.5s ON and 0.5 OFF) or ramping up/down UV light? Can the model make any predictions?

      We carried out new experiments and generated new model predictions. In Figure 6 – figure supplement 7 we show how changing the amplitude and duration of a single UV stimulation influences the response and the model output. In Figure 6 – figure supplement 8 we show how the model behaves when we provide repeated stimulations.

      (10) What are the knock down efficiency of NOS and NIT-GC morpholnio?

      We have now quantified the fluorescence intensity by immunostaining in wild-type and morphant larvae. We show the data for NIT-GC1 and NIT-GC2 in Figure 4—supplement 3F. The knock downs are very efficient. For NOS, we did not do morpholino experiments since we have two null alleles (Figure 3 – figure supplement 1).

      Reviewer #2 (Recommendations For The Authors):

      Why is the behavior of the control animals so different in panels B and C of Figure 3? One group reaches only 15 mm vertical position and the other reaches ~60 mm. Is this just batch variation? Am I missing an experimental variable here? If this is indeed batch variation, then some additional text and statistical analyses might help the reader interpret the behavioral data.

      Thank you for pointing this out. This was a mistake in the original Figure 3C of the unit on the Y axis. We have now corrected this. Additionally we have added statistical tests.

      Really, that’s the only, relatively minor issue I could find. This was a pleasure to read, and I learned a lot. Congratulations on an excellent study.

      Thanks a lot for these comments.

      Reviewer #3 (Recommendations For The Authors):

      Page 3, Figure 1 A: The authors reconstructed the cPRC circuit in 3-day-old larvae and detected NOS gene expression in 2-day-old larvae in Figure 1D&E. Can the authors provide a cPRC circuit reconstruction for 2-day-old larvae?

      We do not have a full connectome of the 2-day-old larva. We also show NOS gene expression in three-day-old larvae (Figure 1 – figure supplement 2).

      Page 4, Figure 2: NO produced by UV/violet stimulation to cPRCs: Including a diagram of the larvae would enhance reader understanding.

      We have added a schematic diagram to Figure 2.

      Page 4: Two Platynereis NOS knockout lines (NOSΔ11/Δ11 and Δ23/Δ23) using the CRISPR/Cas9: Could you direct me to the knockout conformation results for NOS knockout lines NOSΔ11/Δ11 and Δ23/Δ23?

      This is shown in Figure 3 – figure supplement 1 (genetic deletion and loss of antibody signal).

      Page 4: “Three day-old but not two-day-old NOS-mutant larvae also showed reduced phototactic behaviour, suggesting a function for NOS in the visual eyes that mediate three-day-old phototaxis” This sentence is unclear. Why is it only three days old but not two days old?

      2-day-old and 3-day-old larvae have very different type of phototaxis. 2-day-old larvae use their eyspots and show non-visual helical phototaxis. 3-day-old larvae use their visual (‘adult’) eyes for visual phototaxis. We have clarified this sentence: “Given that phototaxis in 1 and 2-day-old trochophore larvae is mediated by their non-visual eyespots (Jékely et al., 2008) and in 3-day-old nectochaete larvae by the visual eyes (Randel et al., 2014), these data suggest a function for NOS in the visual eyes (Figure 3D and Figure 3—figure supplement 1E).”

      Page 5: “Figure 3. NOS is required for UV avoidance in Platynereis larvae.” Detailed experiment setup, if possible, schematic or real setup images would help to replicate the experiments.

      We have added a schematic diagram to Figure 3. 

      Page 5: “All trajectories start at 0 x and y position and time 0 corresponding to 10 sec after the onset of 395 nm stimulation from the side.” How did the authors determine the 10-second duration? Are there any reasons for this choice?

      In our experimental setup we had a limit of tracking individual larvae of approximately 40 seconds. We therefore restricted our analysis to a 10 sec pre and 30 sec post-stimulus interval.

      “(A) Swimming trajectories of wild type (WT, n=32) and NOS mutant (NOSΔ11/Δ11, n=26 and NOSΔ23/ Δ23, n=47) three-day-old larvae.” Authors keep shifting between 2-day and 3-day-old larval data in Figures 1, 2, and 3, causing inconsistency.

      Due to their elongated shape and active muscular contractions it is difficult to carry out calcium imaging experiments with 3-day-old larvae. These experiments were therefore done with 2-day-old larvae. However, we did our behavioural experiments with both 2- and 3-day-old larvae. We observed similar patterns of UV-avoidance behaviour, NOS gene expression, ciliary activity, and mutant phenotype. We are therefore confident that the cPRC responses and circuit activity are similar across these two stages.

      Page 5: “Figure 3. NOS is required for UV avoidance in Platynereis larvae.” Why didn’t the authors present any statistics on the plots? A statistical test is required to prove that the difference is significant.

      We have added statistical tests to the data in Figure 3.

      Page 5: “Analysis of sGCs in Platynereis indicated that these genes are not expressed in any of the cells of the cPRC circuit (not shown and (Verasztó et al., 2017)).” Hence the data is relevant; please provide this data in supplementary.

      We re-analysed previous single-cell data (Achim et al., 2018) and found no sGC homologues detected in cPRC and INNOS. However, we detected an sGCβ subunit in the INRGWa cells. We added these data to the source data and amended the figure and the text. “Analysis of sGCs in Platynereis showed that these genes were not expressed in cPRC or INNOS cells, we only detected expression of an sGCβ subunit in the INRGWa cells (Figure 6B).”

      Page 10: “Diagram of the mathematical model with the components, interactions, parameters and equations used to model Ca dynamics.” It would be helpful to include a legend explaining the meaning of each term (e.g. what is K, Co, Cp etc) and arrow color. Additionally, the main text should provide detailed descriptions to aid understanding.

      We have added a new diagram of the model and its parameters to Figure 6. In addition, in the Methods section we have an extensive description of the model and all its parameters.

      Page 12: “This activated state is maintained for several tens of second.” Can authors point out the data?

      We point to the relevant figure now. “The high-Ca2+-state is then maintained for several tens of second (Figure 4A, B).”

      Page 13: “In the Platynereis circuit, our mathematical model indicates that the magnitude of the NO-dependent signal depends on the intensity and duration of the UV/violet stimulus.” Did the authors conduct experiments of various intensity and duration apart from modelling? If not, such experiments should be carried out to validate the mathematical model.

      We have carried out new calcium imaging experiment with varying duration and intensity of the stimulation. These experiments were key to the revision of the model because they clearly pointed to the time-invariance of the NO-dependent peak in the cPRCs. These data and their model fits are summarised in Figure 6 – figure supplement 7.

      References

      Bauknecht P, Jékely G. 2017. Ancient coexistence of norepinephrine, tyramine, and octopamine signaling in bilaterians. BMC Biology 15. doi:10.1186/s12915-016-0341-7

      Jékely G, Colombelli J, Hausen H, Guy K, Stelzer E, Nédélec F, Arendt D. 2008. Mechanism of phototaxis in marine zooplankton. Nature 456:395–399. doi:10.1038/nature07590

      Jékely G, Yuste R. 2024. Nonsynaptic encoding of behavior by neuropeptides. Current Opinion in Behavioral Sciences 60:101456. doi:10.1016/j.cobeha.2024.101456

      Randel N, Asadulina A, Bezares-Calderón LA, Verasztó C, Williams EA, Conzelmann M, Shahidi R, Jékely G. 2014. Neuronal connectome of a sensory-motor circuit for visual navigation. eLife 3. doi:10.7554/elife.02730

      Verasztó C, Gühmann M, Jia H, Rajan VBV, Bezares-Calderón LA, Piñeiro-Lopez C, Randel N, Shahidi R, Michiels NK, Yokoyama S, Tessmar-Raible K, Jékely G. 2018. Ciliary and rhabdomeric photoreceptor-cell circuits form a spectral depth gauge in marine zooplankton. eLife 7. doi:10.7554/elife.36440

      Verasztó C, Jasek S, Gühmann M, Bezares-Calderón LA, Williams EA, Shahidi R, Jékely G. 2025. Whole-body connectome of a segmented annelid larva. eLife 13. doi:10.7554/elife.97964.3

      Williams EA, Verasztó C, Jasek S, Conzelmann M, Shahidi R, Bauknecht P, Mirabeau O, Jékely G. 2017. Synaptic and peptidergic connectome of a neurosecretory center in the annelid brain. eLife 6. doi:10.7554/elife.26349

    1. eLife Assessment

      This manuscript applies a theoretical analysis to two published datasets on yeast and bacterial evolution to compare different ways of quantifying fitness. It makes an important advance by clarifying how discrepancies can arise by using different approaches and provides recommendations for best practices. Overall, this is an impressive and highly beneficial study that is based on convincing evidence and has the potential of setting standards in this rapidly growing field.

    2. Reviewer #1 (Public review):

      [Editors' note: this version has been assessed by the Reviewing Editor without further input from the original reviewers. The authors have addressed the comments raised in the previous round of review.]

      The authors point out that the fitness estimates obtained from different experimental assays (monoculture, pairwise competition or bulk competition) are not generally equivalent, not even with regard to the fitness ranking of different genotypes. Using a computational model based on experimentally measured growth phenotypes for knockout strains in yeast, as well as data from Lenski's Long Term Evolution Experiment (LTEE), they derive a set of best practice rules aimed at extracting the optimal amount of information from such experiments.

      The study is very complete on a technical level, and the conceptual weaknesses raised in the first round of reviews have been fully addressed in the revision.

    3. Reviewer #2 (Public review):

      Summary:

      The manuscript "Quantifying microbial fitness in high-throughput experiments" provides a comprehensive analysis of the various approaches to quantifying fitness in microbial evolution, focusing on three primary factors: encoding of relative abundance, time scale of measurement, and the choice of reference subpopulation. The authors systematically explore how these choices impact fitness statistics and provide recommendations aimed at standardizing practices in the field. This manuscript aims to highlight the impact of differing fitness definitions and the methodologies utilized for analysis and how that can significantly alter interpretations of mutant fitness, affecting evolutionary predictions and the overall understanding of genetic interactions in the experiments.

      Strengths:

      The choices for quantifying fitness in evolution experiments are critical and highly relevant given the increasing prevalence of high-throughput experiments in evolutionary biology. The authors methodically categorize fitness statistics and their implications, providing clarity on a complex subject. This structured approach aids in understanding the nuances of fitness measurement. The manuscript effectively highlights how different choices in fitness measurement can influence fitness rankings and the understanding of epistasis, which is important for modeling evolutionary dynamics.

      Comments on revisions:

      The authors have comprehensively addressed all previous comments and suggestions. In particular, the addition of the new methods section: 'A guide to calculate pairwise relative fitness under the logit encoding from bulk competition data' - significantly improves the clarity of the implementation and helps in the overall interpretation of the framework.

    4. Reviewer #3 (Public review):

      Summary:

      The authors present analyses of different fitness measures derived from empirical data from yeast knock-out mutants and the long-term evolution experiment (LTEE) with Escherichia coli to explore discrepancies and identify preferred methods to estimate relative fitness in high-throughput experiments. Their work has three components. They first discuss the different "encodings" of relative abundance data and conclude that logit-transformations are preferred, because they transform nonlinear abundance trajectories into linear trajectories with greater predictive power. Next, they compare per-generation with per-growth cycle relative fitness estimates inferred from simulations of pairwise competitions based on published growth traits for the yeast strains and on published pairwise competition measurements for the LTEE data. Both data sets show quantitative and qualitative (i.e. rank order) discrepancies of estimates across different time scales, which are highlighted by considering possible underlying causes (i.e. trade-offs between growth traits) and consequences (i.e. epistasis among mutations affecting different growth traits). Finally, the authors compare simulated pairwise and bulk (i.e. where many mutants compete during a growth cycle in a single environment) competition assays based on the yeast knock-out mutants and demonstrate an optimal ratio of collective mutants to wild-type strains that minimizes both sampling error and overestimation of fitness estimates when compared with pairwise competitions.

      Strengths:

      The study deals with a highly relevant topic. Fitness is central to general evolutionary theory, but also poorly defined and implies different traits for different organisms and conditions. For microbes, which are often used in evolution experiments, high-throughput experiments may yield different measures to quantify abundance over time, from individual growth traits to bulk competition experiments. Hence, it is relevant to consider discrepancies among those measures and identify preferred measures with respect to predicting population dynamic and evolutionary processes. The present study contributes to this aim by (i) making readers aware of differences among commonly used fitness estimates, (ii) showing that simulated (yeast) and calculated (E. coli) competitive fitness may differ across time scales, and (iii) showing that bulk competitions may yield relative fitness estimates that are systematically higher than pairwise competitions. The study is rather thorough on the theory side, with extensive derivations and analyses of various fitness measures using their resource competition model in the Supplementary Information. The study ends with a few practical recommendations for preferred methods to infer relative fitness estimates, that may be useful for experimentalists and stimulate further investigations.

      Comments on revisions:

      I appreciate the thorough and effective response to all recommendations and have no further comments.

    5. Author response:

      The following is the authors’ response to the previous reviews

      We thank both editors and the three reviewers for their positive feedback on the revised manuscript. Following suggestions from Reviewer #1, we cite additional literature to clarify the use of ’fitness potential’ and we have revised the caption of Figure 3 to better explain how we generate the panel of mutants. We have also added a sentence to emphasize that the selection coefficient should match the time-scale of bottleneck effects in the evolution environment.

    1. eLife Assessment

      The authors ask whether a simple whole-head spectral power analysis of human magnetoencephalography data recorded at rest in a large cohort of adults shows robust effects of age, and their results provide compelling evidence that it does. The relative simplicity of the analysis is a major strength of the paper, and the authors are careful to control for many different confounds - although perhaps highly correlated factors like brain anatomy still pose a slight issue. The paper provides a valuable power analysis framework that should inform researchers across the broader neuroimaging community.

    2. Reviewer #1 (Public review):

      Summary:

      This is a careful, well-powered treatment of age effects in resting-state MEG. Rather than extracting (say) complex connectivity measures, the authors look at the 'simplest possible thing' : changes in the overall power spectrum across age.

      Strengths:

      They find significant age-related changes at different frequency bands: broadly: attenuation at low-frequency (alpha) and increased beta. These patterns are identified in a large dataset (CamCAN) and then verified in other public data.

      Weakness:

      Some secondary interpretations (what is "unique" to age vs global anatomy) maybe go beyond what the statistics strictly warrant in the current form, but these can be tightened with (I think pretty quick) additions already foreshadowed by the authors' own analyses.

      Aims:

      The authors set out to replace piecemeal, band-by-band ageing claims with t-maps, and Cohen's f2 over sensors×frequency ("GLM-Spectrum").

      On CamCAN, six spatio-spectral peaks survive relatively strict statistical controls. The larger effects are in low-frequency and upper-alpha/beta ranges (f2 approx. 0.2-0.3), while lower-alpha and gamma reach significance but with small practical impact (f2 < 0.075). A nice finding is that the same qualitative profile appears in three additional independent datasets.

      Two analyses are especially interesting. First, the authors show a difference between absolute and relative spectral magnitude (basically within-subject normalization). Relative scaling sharpens spectral specificity of the spatial maps while absolute magnitude is dominated by a broad spatial mode that correlates positively across frequencies, likely reflecting head-position/field-spread factors. The replication of the main age profile is robust to preprocessing decisions (e.g. SSS movement compensation choices) - the bigger determinant of the effect is whether they apply sensor normalization (relative vs absolute).

      Second, lots of brain-related things might be related to age and the authors spend some time trying to back out confounds / covariates. This section is handled transparently (in general I found the writing style very clear throughout) - they examine single covariates (sex, BP, GGMV, etc.) and compare simple vs partial age effects. For example, aging is correlated with reductions in global grey-matter volume (GGMV) but it would be nice to find a measure that is independent of this : Controlling for GGMV (via a linear model) reduces age-related effect sizes heterogeneously across space/frequency but does not eliminate them, a nuance the authors treat carefully.

      This is a nice paper and I have only a few concrete suggestions:

      (1) High-gamma<br /> There can be a lot of EMG / eye movement contamination (I know these were RS eyes closed data but still...) above 30-40 Hz and these effects are the weakest anyway. Could you add an analysis (e.g. ICA/label-based muscle component removal) and show the gamma band's sensitivity to that step. Or just note this point more clearly?

      (2) GGMV confound control<br /> Controlling for GGMV reduces, but does not eliminate, age effects. I have a few questions about this: a) Could we see the residuals as a function of age? I wonder if there are non-linear effects or something else that the regression is not accounting for. Also, b) GGMV and age are highly colinear - is this an issue? Can regression really split them apart robustly? I think by some cunning orthogonalisation you can compute the effect of age independent of GGVM. I don't think this is the same as the effect 'adjusted' for GGMV (which is what is shown here if I'm reading it correctly). Finally, of course, GGMV might actually be the thing you want to look at (because it might more accurately reflect clinical issues) - so strong correlations are not really a problem: I think really the focus might even be on using MEG to predict GGMV and controlling for age.

      Minor presentation edits:

      It would be handy to see a single table listing each tested "analysis family" (e.g., sensors×frequency, source parcels×frequency), the multiple control used, and the permutation count. I kept wanting to see this as I was reading to compare back and fore.

      I loved the power-planning content (section 3.2, the table with peak f2, CIs, contour plot). I think you could somehow make this even more explicit because people will use it a lot - both for this age/MEG domain and more generally as a template for other types of power planning in the field. Perhaps a "How to use this paper to plan N" guide in a paragraph? Power analysis is surely both "important and difficult" - but also not impossible. A flowchart?

      Comments on the latest version:

      The authors address all my initial points in their revisions and I have no further comments.

    3. Reviewer #2 (Public review):

      This paper describes application of the "GLM-Spectrum" mass univariate approach to examine the effects of age on M/EEG power spectra. Its strengths include promotion of the unbiased approach, suitable for future meta/mega-analyses, and the provision of effect sizes for powering future studies. These are useful contributions to the literature. What is perhaps lacking is discussion of limitations of this approach, in comparison to other methods.

      An analogy is the mass univariate approach to spatial localisation of effects in fMRI/PET images. This approach is unbiased by prior assumptions about the organisation of the brain, but potentially also less sensitive, by ignoring that prior knowledge. For example, a voxelwise univariate approach is less sensitive to detecting effects in functionally homogeneous brain regions, where SNR can be increased by averaging over voxels. In the context of power spectra, the authors' approach deliberately ignores knowledge about the dominant frequency bands / oscillations in human power spectra. This is in contrast to approaches like FOOOF and IRASA, that explicitly parametrise frequency components. I am not saying these methods are better; I just think that the authors should acknowledge that these approaches have advantages over their mass univariate approach (in sensitivity and interpretation; see below). I guess it is a type of bias-sensitivity trade-off: the authors want to avoid bias, but they should acknowledge the corresponding loss of sensitivity, as well as loss of interpretation compared to model-based approaches (i.e., models that parameterise frequency; I don't mean the statistical models for each frequency separately).

      An example of the interpretational loss can be seen in the authors' observation of opposite-signed effects of age around the alpha peak. While the authors acknowledge that this pattern can arise from a reduction in alpha frequency with age, this is an indirect inference, and a direct (and likely much more sensitive) approach would be to parametrise and estimate the peak alpha frequency directly for each participant, as done with FOOOF for example (possibly with group priors, as in Medrano et al, 2025, EJN). The authors emphasise the nonlinear effects of age in Fig 2A, but their approach cannot test this directly (e.g. in terms of plotting effects of age on frequency, magnitude, width for each participant), so for me, this figure illustrates a weakness of their approach, not a strength.

      Then I think the section "Two dissociable and opposite effects in the alpha range" in the Discussion section is confusing, because if there is a single reduction in alpha peak frequency and magnitude with age, then there is only one "effect", not "two dissociable" ones. If the authors do want to claim that there are two dissociable age effects within the alpha range, then they need to do a statistical test, e.g., that the topographies of low and high alpha are significantly different. This then reveals another limitation of the mass univariate approach - that space (channel) is not parametrised either - so one cannot test for significant channel x effect interactions within this framework, as necessary to really claim a dissociation (e.g., in underlying neural generators).

      While the authors show that normalisation of each person's power spectra by the sum across frequencies helps improve some statistics, they might want to say more about disadvantages of this approach, e.g., loss of sensitivity to any effects (e.g. of age) that are broadly distributed across majority of frequencies, loss of real SI units (absolute effect sizes) (as well as problems if normalisation were used for techniques like FOOOF, where the 1/f exponent would be affected).

      Please give more information how artifactual ICs were defined. This may be important for cardiac artefacts, since Schmidt et al (2004, eLife) have pointed out how "standard" ICA thresholds can fail to remove all cardiac effects. This is very important for effects of age, given that age affects cardiac dynamics (even though the focus of Schmidt et al is the 1/f exponent, could residual cardiac effects cause artifactual age effects in current results, even above ~1Hz?).

      Please could the authors clarify the precise maxfilter arguments, and explain what "reference" was used for the "trans" option - e.g., did the authors consider transforming the data to match a sphere at the centre of the helmet, which might not only remove some of the global power differences due to different head positions, but also be best for generalisation of the effect sizes they report to future studies (assuming the centre of the helmet is the most likely location on average)? And on that matter, did head positions actually differ by age at all?

      Comments on the latest version:

      I am happy with their revised version.

    4. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      The authors ask whether a simple whole-head spectral power analysis of human magnetoencephalography data recorded at rest in a large cohort of adults shows robust effects of age, and their results provide compelling evidence that it does. The relative simplicity of the analysis is a major strength of the paper, and the authors are careful to control for many different confounds - although perhaps highly correlated factors like brain anatomy still pose a slight issue. The paper provides a valuable power analysis framework that should inform researchers across the broader neuroimaging community

      Many thanks to the reviewers and editorial team. This is an insightful and engaging set of reviews with a range of productive suggestions. We’re pleased that the strengths of this approach show through and that this can be a positive contribution to the community.

      We have implemented the large majority of suggestions and believe that the paper is greatly improved with them in place.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This is a careful, well-powered treatment of age effects in resting-state MEG. Rather than extracting (say) complex connectivity measures, the authors look at the 'simplest possible thing': changes in the overall power spectrum across age.

      Strengths:

      They find significant age-related changes at different frequency bands: broadly, attenuation at low-frequency (alpha) and increased beta. These patterns are identified in a large dataset (CamCAN) and then verified in other public data.

      Weaknesses:

      Some secondary interpretations (what is "unique" to age vs global anatomy) may go beyond what the statistics strictly warrant in the current form, but these can be tightened with (I think, fairly quick) additions already foreshadowed by the authors' own analyses.

      Aims:

      The authors set out to replace piecemeal, band-by-band ageing claims with t-maps, and Cohen's f2 over sensors×frequency ("GLM-Spectrum").

      On CamCAN, six spatio-spectral peaks survive relatively strict statistical controls. The larger effects are in low-frequency and upper-alpha/beta ranges (f2 approx 0.2-0.3), while lower-alpha and gamma reach significance but with small practical impact (f2 < 0.075). A nice finding is that the same qualitative profile appears in three additional independent datasets.

      Two analyses are especially interesting. First, the authors show a difference between absolute and relative spectral magnitude (basically, within-subject normalization). Relative scaling sharpens the spectral specificity of the spatial maps, while absolute magnitude is dominated by a broad spatial mode that correlates positively across frequencies, likely reflecting head-position/field-spread factors. The replication of the main age profile is robust to preprocessing decisions (e.g., SSS movement compensation choices) - the bigger determinant of the effect is whether they apply sensor normalization (relative vs absolute).

      Second, lots of brain-related things might be related to age, and the authors spend some time trying to back out confounds/covariates. This section is handled transparently (in general, I found the writing style very clear throughout) - they examine single covariates (sex, BP, GGMV, etc.) and compare simple vs partial age effects. For example, aging is correlated with reductions in global grey-matter volume (GGMV), but it would be nice to find a measure that is independent of this: controlling for GGMV (via a linear model) reduces age-related effect sizes heterogeneously across space/frequency but does not eliminate them, a nuance the authors treat carefully.

      Thank you for this concise summary of the work. We’re glad that the strengths of the approach come through clearly.

      This is a nice paper, and I have only a few concrete suggestions:

      (1) High-gamma 

      There can be a lot of EMG / eye movement contamination (I know these were RS eyes closed data, but still..) above 30-40 Hz, and these effects are the weakest anyway. Could you add an analysis (e.g., ICA/label-based muscle component removal) and show the gamma band's sensitivity to that step? Or just note this point more clearly?

      Thanks for this suggestion. We agree that there is concern about eye movements for these gamma band analyses. The ICA preprocessing we conducted was relatively thorough, and we were able to remove EOG-related components from the majority of datasets, even though the experimental protocol involved eyes closed at rest. It is possible that some components were missed, but they would require more advanced labelling tools or manual intervention to identify.

      There are alternative automated ICA labelling tools that could be used, but these are either not suitable for MEG (ICLabel; https://labeling.ucsd.edu/tutorial, https://mne.tools/mne-icalabel/dev/api/iclabel.html) or optimised for data from CTF/4D systems (MEGNet; https://doi.org/10.1016/j.neuroimage.2021.118402). It is beyond our capacity to modify one of these tools for CamCAN for the current analyses.

      It was much more straightforward to rerun the analysis without the ICA step to see the impact of reintroducing all ocular artefacts removed from the v1 analysis. These results are shown in Supplemental section XXX and replicated below.

      We have added the following Figure 11 and text to the paper.

      Main Text

      “We have completed several control analyses to support these findings. Firstly, we have explored the correspondence between alpha peak frequency and the two effects we identified within the canonical alpha range (see supplemental section A.1). Secondly, the overall pattern of findings is consistent in an equivalent source space analysis using LCMV beamforming and parcellation (see supplemental section A.2). Finally, we have repeated the analyses with and without ICA denoising and find that the overall spectral profile is very similar. Effects in low-frequencies, low-alpha and high-gamma are increased with application of ICA, whilst high-alpha and beta remain unchanged and low-gamma effects are reduced (see supplemental section A.3). “

      Supplemental materials

      “The age results may be contaminated in some way by residual cardiac or ocular artefacts that are not removed during preprocessing. Though ICA denoising was applied, it is possible that some artefactual components were not identified and removed from the dataset. The results at low frequencies and in the gamma range are most likely to be directly impacted by this contamination.

      To explore the impact this has on our analysis, we reran the core GLM effect of age on the data with no ICA artefact rejection at all, allowing all eye movements and heart rate components to remain in the data (Figure 10).

      This no-ICA analysis has three differences to the original in the publication. The low-frequency decrease with age is stronger and more widespread in the analysis that removes ocular artefacts with ICA. A large negative effect is visible in both analyses, though without ICA several frontal and temporal sensors no longer show significant effects. Similarly, the effect size of the low-frequency effect of age is substantially larger when ICA denoising is applied.

      The low-alpha effect was strongly reduced in the analyses that do not remove artefacts with ICA. A large central-occipital group of sensors shows an effect between 7 and 8.5 Hz in the ICA analysis, but this is reduced to a single sensor at 8 Hz when ICA is not computed.

      At high frequencies, the age effect in frontal sensors is larger with ICA, and the age effect in posterior sensors is larger without ICA, though the position and frequencies of significant effects are largely unchanged. The remaining effects in the high alpha and beta ranges are unchanged by application of ICA.

      Overall, ICA either improves the estimation of age effects (low-frequency, low-alpha, high-gamma) or has negligible effects (high-alpha, beta). Only the posterior low-gamma effect is reduced by ICA. Together, we take this as evidence that our core results are robust to interference by eye movements and that the ICA denoising is working effectively to reduce noise in the analysis.”

      (2) GGMV confound control 

      Controlling for GGMV reduces, but does not eliminate, age effects. I have a few questions about this: a) Could we see the residuals as a function of age? I wonder if there are non-linear effects or something else that the regression is not accounting for. Also, b) GGMV and age are highly colinear - is this an issue? Can regression really split them apart robustly? I think by some cunning orthogonalisation, you can compute the effect of age independent of GGVM. I don't think this is the same as the effect 'adjusted' for GGMV (which is what is shown here if I'm reading it correctly). Finally, of course, GGMV might actually be the thing you want to look at (because it might more accurately reflect clinical issues) - so strong correlations are not really a problem: I think really the focus might even be on using MEG to predict GGMV and controlling for age.

      This is an interesting area with some tricky interpretation. Thanks for the nudge to help us clarify further.

      We have added the following text and Figures 17 & 18 to the paper to clarify these points.

      Main Text Section 2.8

      “It is important to note that the correlation between age and GGMV does impact the interpretation of the GLM results, but does not prevent the model fit. We explore the model validation and diagnostics in detail in Supplemental Section A.6. In brief, the model is able to separate the unique contributions of age and GGMV. However, the correlation between factors leads to an inflation in the standard error of the estimates. A hypothesis test on these estimates is valid. However, the inflated variance reduces our ability to detect statistically significant partial effects.”

      Supplemental Section A.6

      “There is a strong correlation between age and Global Grey Matter Volume (Pearson’s r=-0.75). The shared variance arising from this collinearity adds nuance to the interpretation of the results, which we explore in more detail in this section.

      Firstly, the sum-square residuals for the group-level model fit including both Age and GGMV are shown as a function of frequency in Figure 17. We see that the residuals broadly follow the overall distribution of variance in the data, peaking at low frequencies and in the alpha range. These frequency ranges are where the strongest signal is visible, but also the highest variability between participants. We would expect that the group model would not perform so well in the points of greatest variability. Importantly, though the residuals are relatively high in the alpha, this is still in the context of a very well-performing model with R2 values of around 80%. “

      “Secondly, the correlation between age and GGMV is not inherently problematic for the GLM, but it does add complexity and nuance to the interpretation of the results. Some additional model validation statistics are shown in Figure 18. “

      “The singular value spectrum of the design matrix indicates whether a design is low-rank, the smallest singular value in this case in 0.37 which indicates that there isn’t a rank deficiency which would prevent us from estimating the model. Though we can estimate the model, correlated regressors can reduce its efficiency. The variance inflation factors for the joint AGE-GGMV model are above 1 for both parametric regressors, indicating that the standard errors of their estimates are inflated. The VIF of 2.35 indicates that the standard errors of this joint model are around sqrt(2.35) = 1.533 times greater than they would be in a separate or uncorrelated model. Though there is no hard rule for this, the literature generally suggests that a VIF above 5 (or sometimes 10) indicates severe multicollinearity.”

      Including additional regressors in the model can change the age estimate by ‘partialling’ out the variance that can be attributed to the other variables and by inflating standard errors. In this specific case of GGMV, the partialled estimates are reduced heterogeneously across space and frequency, and the amount of inflation is at a tolerable level. Overall, the regression is able to separate the unique effects of age and GGMV, at the cost of this inflation in the associated standard errors.”

      Reviewer #2 (Public review):

      This paper describes the application of the "GLM-Spectrum" mass univariate approach to examine the effects of age on M/EEG power spectra. Its strengths include promotion of the unbiased approach, suitable for future meta/mega-analyses, and the provision of effect sizes for powering future studies. These are useful contributions to the literature. What is perhaps lacking is a discussion of the limitations of this approach, in comparison to other methods.

      Thank you for this summary and the thoughtful review. The emphasis for this paper is exactly on the points you highlighted, and we’re glad that this has come across well. We agree that a broader comparison to other methods would be a useful addition and have included a series of additional discussion points to address this.

      We will take the opportunity to reclarify that our intention for this method is not to replace other, more complex or targeted approaches, but to establish a more generalisable foundation for their development. We’re fully supportive of other approaches and are working on their application ourselves. On reflection, this was not clear enough in the first submission, and we have added the following text to the introduction to clarify.

      “Reporting of whole-head and full-frequency spectra of effect estimates would make it straightforward to aggregate across studies and eventually enable identification of sub-threshold effects that may be missed in single analyses but are consistent across studies. We argue that this approach provides a generalisable foundation that can support more complex analyses with a frequency component (such as aperiodic slopes, burst detection, and dynamic functional networks) that require more researcher degrees of freedom.”

      An analogy is the mass univariate approach to spatial localisation of effects in fMRI/PET images. This approach is unbiased by prior assumptions about the organisation of the brain, but potentially also less sensitive, by ignoring that prior knowledge. For example, a voxelwise univariate approach is less sensitive to detecting effects in functionally homogeneous brain regions, where SNR can be increased by averaging over voxels.

      In the context of power spectra, the authors' approach deliberately ignores knowledge about the dominant frequency bands/oscillations in human power spectra. This is in contrast to approaches like FOOOF and IRASA, which explicitly parametrise frequency components. I am not saying these methods are better; I just think that the authors should acknowledge that these approaches have advantages over their mass univariate approach (in sensitivity and interpretation; see below). I guess it is a type of bias-sensitivity trade-off: the authors want to avoid bias, but they should acknowledge the corresponding loss of sensitivity, as well as loss of interpretation compared to model-based approaches (i.e, models that parameterise frequency; I don't mean the statistical models for each frequency separately).

      This is an important point, and we are in complete agreement about the importance of giving a balanced description of how this approach fits within the broader literature. We have added the following paragraph to the discussion on limitations to lay this out more clearly.

      “Our approach promotes an exploratory and unbiased approach to quantifying the age effect on neuronal power spectra, which is intended to complement more focused analyses. This has the benefit of reducing researchers’ degrees of freedom and of being broadly generalisable. These come at the cost of a loss in specificity and in sensitivity. Our approach does not specifically quantify features derived from the power spectrum, such as alpha-peak frequency or the aperiodic component of the spectrum. These features are mixed into our full-spectrum estimates but not directly quantified. Thus, they can be challenging to interpret from our approach. Secondly, the mass-univariate approach suffers from a potential loss in sensitivity compared to results that aggregate across spatial or spectral regions that contain consistent results. Where a region or frequency band of interest can be supported from the literature, an approach focusing on a single region has the benefit of reduced noise by averaging estimates from a larger range of observations. Finally, models that consider the whole shape of the spectrum [Donoghue et al. 2021] would also be able to combine information across a range of frequencies rather than depending on a single frequency bin for each estimate. These models have the additional benefit that their parameters are often directly interpretable as features of interest, such as spectral slope or peak frequency.”

      An example of the interpretational loss can be seen in the authors' observation of opposite-signed effects of age around the alpha peak. While the authors acknowledge that this pattern can arise from a reduction in alpha frequency with age, this is an indirect inference, and a direct (and likely much more sensitive) approach would be to parametrise and estimate the peak alpha frequency directly for each participant, as done with FOOOF for example (possibly with group priors, as in Medrano et al, 2025, EJN). The authors emphasise the nonlinear effects of age in Figure 2A, but their approach cannot test this directly (e.g., in terms of plotting effects of age on frequency, magnitude, and width for each participant), so for me, this figure illustrates a weakness of their approach, not a strength.

      We agree that this point might be misleading within its own figure and have moved the result to the supplemental material with a more lightly phrased wording in the main text. Figure 3 on effect sizes has moved to Figure 2, and a new Figure 3 shows the quadratic effect of age as suggested later in the review.

      “We have completed several control analyses to support these findings. Firstly, we have explored the correspondence between alpha peak frequency and the two effects we identified within the canonical alpha range (see supplemental section A.1). Secondly, the overall pattern of findings is consistent in an equivalent source space analysis using LCMV beamforming and parcellation (see supplemental section A.2). Finally, we have repeated the analyses with and without ICA denoising and find that the overall spectral profile is very similar. Effects in low-frequencies, low-alpha and high-gamma are increased with application of ICA, whilst high-alpha and beta remain unchanged, and low-gamma effects are reduced (see supplemental section A.3).“

      This supplemental section contains some additional content relevant to the next comment.

      Then I think the section "Two dissociable and opposite effects in the alpha range" in the Discussion section is confusing, because if there is a single reduction in alpha peak frequency and magnitude with age, then there is only one "effect", not "two dissociable" ones. If the authors do want to claim that there are two dissociable age effects within the alpha range, then they need to do a statistical test, e.g., that the topographies of low and high alpha are significantly different. This then reveals another limitation of the mass univariate approach - that space (channel) is not parametrised either - so one cannot test for significant channel x effect interactions within this framework, as necessary to really claim a dissociation (e.g., in underlying neural generators).

      As above, we agree that this can be misleading and that our intention got muddled in the heading and writing of this subsection. We are not intending to argue that there are definitely two distinct and different effects. We clarify this in the main text of our first version, but not clearly enough:

      “The two effects identified in the present analysis may combine to represent a decrease in power and frequency of a single alpha peak”.

      We choose to present the first paragraph of this section as a discussion of two effects because this is what is already reported in the literature on changes in alpha power with age. Whilst the decrease in alpha frequency with age is well replicated, the decrease in alpha power is less consistently reported, and the literature shows a highly variable picture of the spatial pattern of this effect. We do not claim a strong dissociation based on these results, but both are reported within the literature.

      Our intention is to provide some clarity to this literature by taking a step back and looking at the ‘lay of the land’ in an unbiased way. With this approach, we see different effects of age on magnitude within a canonical alpha range that are completely separated in frequency. With this perspective, it is feasible that different publications taking different regions of interest, frequency band definitions, and processing options could lead to mixed reports of increases and/or decreases in alpha power with age.

      When single papers that take focused but inconsistent approaches report inconsistent results, we would argue that our approach leads to a substantial interpretational gain on the level of collections of papers in the literature.

      We have added the following content to supplemental section A.1 to clarify and Table 3 illustrates the issue. We have retained the paragraph discussing the possibility that these two effects combine to represent a shift in frequency of a single peak.

      Main text section 3.1, replacing the section on ‘Two dissociable effects…’

      “Reconciling conflicting reports of the ageing effect on alpha power.

      The literature exploring how ageing changes alpha power is heterogeneous. Papers that report results in resting alpha power, either from a canonical band or from an individual peak frequency, include reports of a variety of contrasting age effects. This includes positive correlations with age [Rempe et al., 2023, Stier et al., 2023], negative correlations [Thuwal et al., 2021, Lodder and van Putten, 2011, Medrano et al., 2025, Park et al., 2024], both positive and negative effects separated by space [Hoshi and Shigihara, 2020, Pathak et al., 2022], or null results when correcting for individual frequencies and aperiodic slopes [Scally et al., 2018, Merkin et al., 2023] (see supplemental section A.1 more detailed summary). There is broad variability in methodological approaches, which could account for the variety of results. Stier et al. [2023] suggest that analyses in sensor or source space may lead to different effects. Critically for our work, this variability also prevents formal aggregation of results and meta-analyses that could clarify the picture.

      We have proposed that, by taking a step back and tolerating a reduction in sensitivity, we can map out the whole-head whole-frequency structure of the age effect with minimal researcher degrees of freedom and bias. Investigating age effects as a complete spectrum shows that two contrasting age effects on alpha magnitude coexist in close proximity in space and frequency: an increase with a small effect size in central occipital sensors around 7-8.5 Hz and a decrease with a large effect size across a broad set of occipital, temporal and frontal sensors between 9.5-12.5 Hz. Different data samples and different data analysis choices, particularly the selection of regions of interest, source reconstruction, sensor normalisation, or correction of aperiodic components, might emphasise one effect or the other in each analysis. Whilst we have not conclusively explored all possible variants of these analyses, we have provided a framework that would allow future studies to perform formal comparisons and meta-analyses to resolve this bottleneck.”

      “Relationship between effects on alpha power and alpha individual frequency.

      The two effects identified in the present analysis may combine to represent a decrease in power and frequency of a single alpha peak (Seen qualitatively in Figure 1A). This change in alpha peak frequency is highly replicable [Cesnaite et al., 2023, Dustman et al., 1993, Sahoo et al., 2020, Scally et al., 2018, Pathak et al., 2022, Zibrandtsen and Kjaer, 2021] and is a highly predictive spectral marker of ageing [Stier et al., 2024]. Decreases in alpha peak frequency have been linked to a decline in cognitive performance in healthy ageing [Cesnaite et al., 2023, Finley et al., 2024] and MCI [Garc´es et al., 2013, L´opez-Sanz et al., 2016, Puttaert et al., 2021].

      This compelling possibility that the age effect on alpha is a shift in a single peak is complicated by strong evidence for presence of multiple alpha peaks within individuals [Lodder and van Putten, 2011, Chiang et al., 2011, 2008, Klimesch, 1999], with distinct generators and functional relevance [Sokoliuk et al., 2019]. A complex pattern of changes in power, frequency, and spatial distribution likely underlies age-related change in alpha oscillations. Future work will need to explore all three features at the individual level to clearly illuminate the change.”

      Supplemental section A.1

      “Part of our motivation for this method is that variability in the methodological choices in different publications makes it difficult to aggregate varying results across the literature. For example, though a decrease in alpha peak frequency with increasing age is reported highly consistently, the effect of age on alpha power is much more variable. Table 3 shows a representative sample of publications over the last 20 years that report a change in alpha power with age (note that this is intended to be a representative rather than an exhaustive list). Over half of publications (9/15) report a decrease in alpha power, whilst the remaining publications report an increase (2/15), both increases and decreases (1/15), a decrease but only without correcting for aperiodic slope (1/15), a quadratic effect (1/15), and no effect (1/15). Stier et al. [2023] suggest that the choice of analysis space (sensor space or source reconstruction) is likely a source of discrepancies between studies.

      Critically, it is difficult to reconcile these findings with the information reported in the publications. For example, even within the 9 publications that report a decrease, there is little correspondence in the spatial location of the effect. We argue that focused approaches cannot resolve this issue alone, as targeted analyses are more specific to each dataset and less generalisable.

      In the specific case of the mixed literature on change in alpha power with age, our results show that both effects are present and separated in frequency. It is feasible that the different methodological choices and datasets used by each study in our survey means that one or other of these two effects were emphasised. As a result, the literature may not be mixed in scientific terms, but that a rich pattern of results is obscured by methodological variability.”

      While the authors show that normalisation of each person's power spectra by the sum across frequencies helps improve some statistics, they might want to say more about disadvantages of this approach, e.g., loss of sensitivity to any effects (eg of age) that are broadly distributed across majority of frequencies, loss of real SI units (absolute effect sizes) (as well as problems if normalisation were used for techniques like FOOOF, where the 1/f exponent would be affected).

      This is an important point, and we have added the following text to clarify, with one exception. Firstly, the normalisation we applied scales the whole spectrum linearly and would change the 1/f intercept but not the 1/f^alpha exponent.

      Main text section 3.3

      “Both absolute and relative power measures are used throughout the literature, but there is little consensus about their interpretation [Sandre and Troller-Renfree, 2026]. We focus on relative power for most of our results. By normalising each participant’s power spectrum by the sum across frequencies, we found that the results gained specificity in frequency band and reduced concern about wide intersubject differences in overall variance. Though this improved some analyses, relative power has important drawbacks. In particular, it can reduce sensitivity to effects that are broadly distributed across the spectrum and uses arbitrary scaling rather than meaningful physical units. We support calls in the literature to report both relative and absolute power measures [Rempe et al., 2023, Sandre and Troller-Renfree, 2026].”

      The authors should give more information on how artifactual ICs were defined. This may be important for cardiac artefacts, since Schmidt et al (2004, eLife) have pointed out how "standard" ICA thresholds can fail to remove all cardiac effects. This is very important for the effects of age, given that age affects cardiac dynamics (even though the focus of Schmidt et al is the 1/f exponent, could residual cardiac effects cause artifactual age effects in current results, even above ~1Hz?).

      Artefactual components were estimated using standard tools in MNE python. Specifically:

      https://mne.tools/stable/generated/mne.preprocessing.ICA.html#mne.preprocessing.ICA.find_bads_ecg

      https://mne.tools/stable/generated/mne.preprocessing.ICA.html#mne.preprocessing.ICA.find_bads_eog

      We have clarified the text in the methods to make the overall process clearer.

      We believe that this process has been broadly effective, and we have rejected an average of 2.25 ECG components within each dataset. There remains a strong possibility that residual ECG artefact is present in the data.

      We have added the following text to methods section 4.2

      “Artefactual components relating to eye movements or the heart rate were automatically identified by correlation with the simultaneous EOG and ECG channels. ECG artefacts were identified using cross-trial phase statistics [Dammers et al., 2008] and an automatic threshold based on the sample rate of the data, as implemented in the mne.preprocessing.ICA.find_bads_ecg function in MNE Python. Between 0 and 3 EOG components were rejected in each dataset, with an average of 0.99 (standard deviation: 0.79) across all datasets. EOG artefacts were identified by correlation with the HEOG and VEOG channels, with a threshold set to r = 0.35, as implemented in the mne.preprocessing.ICA.find_bads_eog function in MNE Python. Between 0 and 5 ECG components were rejected in each dataset, with an average of 2.25 (standard deviation: 0.84) across all datasets. The continuous sensor data were then reconstructed without the influence of the components labelled as artefacts.”

      Similar to the response to Reviewer 1, we have not been able to rerun a more advanced ICA algorithm on the data but have repeated the analysis without any ICA to see if including all ocular and cardiac artefacts influences the results. This change does not introduce any new signal components to the results but does attenuate the low-frequency and low-alpha effects. With the assumption that our initial ICA analysis captured the majority of the largest ECG components, we are confident that our core findings are not compromised by cardiac artefacts.

      Please see the response to comments to Reviewer 1 for additional text in the manuscript on this point.

      The authors should clarify the precise maxfilter arguments, and explain what "reference" was used for the "trans" option - e.g., did the authors consider transforming the data to match a sphere at the centre of the helmet, which might not only remove some of the global power differences due to different head positions, but also be best for generalisation of the effect sizes they report to future studies (assuming the centre of the helmet is the most likely location on average)? And on that matter, did head positions actually differ by age at all?

      We have used the maxfilter files as provided by the CamCAN team and an equivalent implementation defined in OSL-ephys for the Oxford and Cambridge MEGUK data. We have clarified the text in section 4.2

      “All MEG data pre-processing was carried out using MNE-Python [Gramfort, 2013] and OSL-ephys [Quinn et al., 2022, van Es et al., 2025] using the OSL batch pre-processing tools. For the CamCAN data, we proceeded with analysis on post-maxfilter processed data provided by the CamCAN team. Briefly, the data were processed using AA [Cusack et al. 2015] with automatic bad channel detection (limited to 7 channels) with the origin set to the centre of a sphere fitted to the individual’s Polhemus head shape points. Maxfilter signal-space separation was performed with the temporal extension enabled (temporal window of 10 second and correlation threshold of r = 0.98). Head position was continuously estimated and compensated for during periods where the HPI coils were on. After maxfilter processing, head positions were translated into a default head position defined as a point relative to each individual’s origin in a head coordinate frame. Data with the full maxfilter processing and with the head position translation or movement compensation were extracted from the CamCAN database. An equivalent pipeline was implemented in OSL-Ephys and applied to the data from the MEG-UK datasets from Oxford and Cambridge. Files from the MEG-UK Nottingham dataset were processed with third-order gradiometry applied.”

      We found that head position in CamCAN does change with age. We have included the following in Supplemental section A.5

      “The head position of participants within the MEG sensor dewar is an important consideration that has the potential to change the signal-to-noise level of each individual data recording. There are significant differences in head position as a function of age in the CamCAN dataset in Y (front-back) direction indicating that older participants are seated further forward in the dewar than younger participants.”

      As a point of reference, this pattern is consistent with the largest change with age estimated using the un-normalised raw power spectra in Figure 5. A 1-95Hz region shows a change with age that is consistent with older participants sitting further forward in the dewar. This effect is absent in the relative power contrasts. 

      We have not fully explored this final point so have not included it in the main paper, but believe it is a useful addition to the discussion of the reviews.

      Recommendations for the authors:

      Reviewing Editor Comments:

      As you will see, both reviewers are enthusiastic about your paper, indicating that it provides compelling empirical support for its key claims and that it represents a valuable theoretical advance for the field. They provide a number of comments that you may want to consider prior to finalizing the manuscript for publication. All the best, Redmond O'Connell

      Reviewer #1 (Recommendations for the authors):

      (1) It would be handy to see a single table listing each tested "analysis family" (e.g., sensors×frequency, source parcels×frequency), the multiple controls used, and the permutation count. I kept wanting to see this as I was reading to compare back and forth.

      This is a helpful suggestion, we have included the table in a new supplemental section which is referenced from the main text at the start of the results

      Main text section

      “A summary of all GLM analyses carried out in this work is available in supplemental section A.7.”

      With the following table included in supplemental section A.7

      (2) I loved the power-planning content (section 3.2, the table with peak f2, CIs, contour plot). I think you could somehow make this even more explicit because people will use it a lot - both for this age/MEG domain and more generally as a template for other types of power planning in the field. Perhaps a "How to use this paper to plan N" guide in a paragraph? Power analysis is surely both "important and difficult" - but also not impossible. A flowchart?

      We’re very glad that this section is working well and agree that the practical planning content should have been more constructive! We have added the following text as a guide

      Main text section 3.2

      “We propose the following steps as a practical guide for researchers looking to plan a data sample with a reasonable chance of correctly identifying a particular effect.

      (1) Define research question and identify previous results: Your research question must be defined in advance and well specified. There should be relevant data or literature that can be used to guide your decision.

      (2) Define the smallest effect size of interest and the decision criterion for the sample decision: It is critical to define the parameters of how you will make your sample size decision ahead of time. We recommend considering what the ’smallest effect size of interest’ [Anvari and Lakens, 2021] would be for your question. This is specific to your question and is about more than statistical significance. What effect size would indicate that there is a practically meaningful effect for the future literature to consider?

      (3) Estimate effect sizes to inform your decision: Either by aggregating reported statistics from the literature, or by dedicated processing of previous data, compute an estimate of the effect size. Effect sizes are only estimates, so it is important to compute confidence intervals around your estimate to get a measure of variability.

      (4) Compute power/precision curves assuming these results: Using the estimated effect sizes, compute power curves [Baker et al., 2021] that visualise the relationship between effect size, statistical power, and sample size.

      (6) Select the sample size that meets your pre-defined criteria: Using your definitions from step 2, and when considering the whole power curve, make a decision about what sample size would give you a reasonable chance of replicating the effect of interest.

      (7) Make note of any differences or deviations from this plan during your data collection: There are many practical reasons why your planned sample might not match the data acquired in practice. Such deviations from a plan are ok but should be acknowledged, and any mitigating steps explained [Lakens, 2024].

      (8) Document your process for inclusion in a preregistration or publication: Include details on how a sample size decision was made in your research outputs including preregistrations, preprints, and publications. These details will be useful for future researchers to understand your process and to implement their own.”

      Reviewer #2 (Recommendations for the authors):

      (1) Though they mention in the Discussion, the authors could have noted earlier in Section 2.1 that one does not need to assume a linear effect of age - one could use a polynomial expansion or even local splines within the same GLM framework. Indeed, it seems unlikely a priori that effects of age across the wide range of ages in the CamCAN data are linear for all frequencies.

      This is an important point, and one that was straightforward to implement in our model. Based on feedback on other sections, we have removed the previous Figure 2, moved Figure 3 on effect sizes to the second position, and added a new Figure 3 containing a model of the quadratic age effects. We believe that this is a substantial improvement in the paper.

      Main text section 2.3

      “(2.3) Quadratic effects of age

      The linear effect of age is a convenient and simple regression model. However, it makes a strong assumption that change with age is uniform across the whole age range. A second group-level GLM was computed with an additional regressor to quantify quadratic effects of age, which have been reported in the ageing literature [G´omez et al., 2013, Rempe et al., 2023, Stier et al., 2023]. The spectrum of t-values for the quadratic age predictor (Figure 3) shows significant effects for a U-shaped change with increasing age in the low frequency (1-5 Hz) range in central sensors. Significant inverted-U shaped effects are present in two frequency ranges in the beta band. A low-frequency beta effect is present in occipital and temporal sensors between 16 Hz and 20 Hz, whilst a second high-beta effect is present in central sensors between 24 Hz and 30.5 Hz. The low-frequency and low-beta effects overlap in space and frequency with linear effects, suggesting that the low-frequency change with age has both a linear increase with age and a U-shaped component, whilst the low-beta effect has a decrease with age and an inverted-U shaped component. The high-beta effect does not overlap in frequency with any of the reported linear effects. The effect sizes for quadratic effects range between Cohen’s F 2 values of 0.033 for low frequency to 0.076 for high beta and are generally lower than the effect sizes for linear change.”

      Main text section 2.4

      “(2.4) Sample size planning for effects of age on the neuronal power spectrum

      We use the 95% confidence intervals around the effect sizes to make recommendations for future sample sizes for future samples that plan to replicate these results. The observed power calculations have no bearing on the interpretation of the present results. Instead, they should be used as a general guideline for planning future studies. Table 1 gives a full summary of the peak statistics and future sample size range for the six age effects identified in Figure 1B. These results have implications for future sample planning for resting-state electrophysiology studies of ageing. Using the upper bound of the sample size estimates as a conservative estimate, the linear changes with age that have relatively large effect sizes would have well-powered replications with sample sizes of around 50-60 participants. However, the smaller linear effects and all the quadratic effects would require samples of 200 or more participants to have the same probability of detecting the effect if it is indeed present (Table 1). This indicates that study samples should be planned with the smallest effect of interest in mind and that ageing effects in different frequency bands may not all be well powered within the same sample.”

      Discussion section 3.1

      “Linear increase and inverted-U effects in the beta band. The literature consistently reports an increase in low-beta power with age [Gomez et al., 2013, Heinrichs-Graham and Wilson, 2016, Heinrichs-Graham et al., 2018, Hubner et al., 2018, Koyama et al., 1997, Rempe et al., 2023, Stier et al., 2023, Veldhuizen et al., 1993, Xifra-Porxas et al., 2019]. We observed significant inverted-U-shaped quadratic effects of age in two frequency ranges in the beta band, a posterior low-beta component (centred around 18 Hz) and an anterior high-beta component (centred around 25 Hz). This is broadly consistent with reports of quadratic ageing effects in the beta band in the literature [Rempe et al., 2023], though, to our knowledge, our results are the first to suggest a separation of effects into different parts of the beta range. These spectrum changes may be associated with age-related changes in underlying bursting dynamics [Brady et al., 2020, Power et al., 2023].”

      “Methods section 4.7

      Age plus quadratic age models: To move beyond linear change and explore U and inverted-U shaped effects of age, we fitted a group model with three regressors. One constant term, one z-transformed age regressor, and one z-transformed quadratic age regressor (age − mean(age))<sup>2</sup>.”

      (2) The authors could point out an obvious extension of their approach to mixed-effects models, which could properly model within- and between-participant effects (repeated measures), e.g., for longitudinal effects of ageing, GAMMs, etc, including hierarchical linear models that combine trials and participants in the same model, and potentially model trial-specific/stimulus effects, etc.

      This is an important point; we have added the following paragraph to the discussion section.

      Discussion section 3.5

      “Future extensions

      A clear future extension for this work is to formally incorporate estimates of within-subject variability into a mixed-effect model. These powerful models would enable modelling of both fixed effects and random effects, allowing researchers to account for variation within individuals over time and between individuals. LMMs also provide improved approaches for handling missing observations and unbalanced designs, making them especially useful for longitudinal and hierarchical data. A second expansion of this work could use Generalised Additive Mixed Models (GAMMs) to model non-linear relationships using smooth functions while also accounting for random effects. This may allow for more realistic representations of complex patterns of change across age. Linear Mixed Modelling comes with a substantial increase in researcher degrees of freedom and can be challenging to implement and report accurately [Meteyard and Davies, 2020]. We have used fixed-effects modelling in this work in line with our objectives of maintaining generalisability and minimising researcher degrees of freedom.”

      (3) Do the authors want to comment on why effect sizes in Figure 4Ci are so much higher for the Oxford sample? I know the authors' main point is that smaller samples lead to more variable estimates of the true effect size, but there are also other interpretations for these sample differences, e.g., recruitment differences, scanner differences, etc. Have the authors considered implementing empirically Bayesian approaches (like COMBAT, a type of mixed-effects model) to adjust for site differences in both offset and scaling, which I think should be a fairly simple extension of GLM-spectrum?

      We agree that our explanation of here is somewhat lacking. To be transparent, we have thought long and hard about this difference and cannot identify a clear reason why this difference is so striking. There is no apparent difference in the other covariates, such as head position, age, or gender, which could explain why the Oxford dataset has larger effect sizes in this specific frequency range (the results in the low-frequency and alpha ranges are consistent).

      We have added the following to the main text to highlight dedicated data harmonisation strategies that could be applied in this situation.

      Main text section 2.5

      “This may arise from relatively poor estimates of the population level variability from smaller data samples. It is possible that more systematic differences in the participant sampling, recruitment, and data acquisition process could contribute to between-site differences. At present, our analyses can show that the core effects of ageing are replicable across several datasets, though we have not formally combined these datasets into a single analysis. Formal methods for data harmonisation, such as ComBat [Johnson et al., 2006], could correct for additive and multiplicative differences in data scaling across sites to improve site comparisons.

      (4) There are a number of formatting problems - at least in the PDF provided to reviewers - for some symbols, e.g., "below ¡7 Hz" (I presume "j" was originally "~" or something?). Also, sometimes a figure is cited by a single, hyperlinked number, without the word "Figure".

      (5) Line 402: "influence" = "influenced"

      (6) Line 431: date for Gelman & Loken?

      Thank you for highlighting these, we have done a thorough proofread and fixed a large number of spelling and formatting issues.

    1. eLife Assessment

      This valuable study demonstrates how rhythmic inputs can shape sequential working memory. The study puts forward plausible dynamic mechanisms, but some of the central behavioral effects are small, so the evidence - at present - remains incomplete.

    2. Reviewer #1 (Public review):

      Summary:

      This manuscript investigates how rhythmically presented stimuli support working memory by using task-trained recurrent neural networks (RNNs) endowed with short-term synaptic plasticity. RNNs trained with rhythmic sequences have a marginal performance increase (0.4%) over models trained with jittered input and show increased phase-locking and oscillatory organisation during the sample period. While the question addressed in this paper is highly relevant, the core conclusion that regular temporal structures provide a functional scaffold for sequence working memory lacks evidence. The extensive post-hoc filtering pipeline obscures whether there is phase coding or not, and whether or not the found oscillatory phenomena are truly emergent or a mathematical artefact of the analytical selection criteria.

      Strengths:

      (1) The manuscript addresses a highly relevant question.

      (2) The introduction is nicely written and presents relevant background work.

      (3) The authors' results are robust in the sense that they analysed and trained an ensemble of models instead of single networks.

      Weaknesses:

      (1) Misalignment between analysis epoch and core claims. The manuscript argues that temporal regularity supports sequence working memory. However, the majority of analyses focus on the sample/encoding period rather than the delay period during which memory maintenance occurs.

      (2) Ambiguity in the neural code (rate vs. phase). The decoding accuracies suggest that the memory can be well decoded from the instantaneous activity, implying a rate (not a phase) code. This raises two questions:<br /> a) Can memory-related information be decoded directly from the oscillatory phase, particularly during the delay period?<br /> b) What would be the mechanism with which the increase in phase organisation improves a representation that seems otherwise decoded/represented from activity levels?

      (3) Absence of any RNN activity plots. The manuscript would benefit from showing, e.g., single neuron response plots, raster plots, phase histograms of units, etc. Are there actually spontaneous oscillatory dynamics as the paper writes (line 243)? Can the authors show baseline activity (which is also supposed to be oscillatory, line 219)?

      (4) Potential concerns in the analysis pipeline: The data undergo an intensive, selective pipeline that might be susceptible to introducing circularity and selection bias. I highlighted some points here:<br /> a) Many analyses are performed on (summed) data projected on demixed PCs (extracted from time-warped data). Crucially, dPCAs are not unsupervised; they already explicitly maximise the variance of interest.<br /> b) For the phase extraction during sample encoding: after dPCA percentile clipping, z-scoring, and z-score clipping are applied (lines 762-764), low-amplitude trials are excluded (lines 779-781), and there is further selection based on a valid-point criterion and r2 thresholding (lines 804-805). Do all of these selection criteria risk introducing bias?<br /> c) Some statistical assumptions are not explicitly evaluated. E.g., the sign-flip permutation test (lines 735-742) relies on sign-exchangeability.<br /> d) For selectivity analysis of oscillatory organisation (Figure 4C, lines 893-896): Units are first selected by ANOVA, and then on the selected units further stats (Power and PLV) are computed. Unless the further stats are completely independent of the ANOVA, this may introduce selection bias.<br /> e) The finding of stronger power around f0 given rhythmic inputs of that exact frequency seems somewhat circular (Figure 3A)?<br /> f) The dPCA description seems a little odd, e.g., line 674, for the ordinal component you would normally actually average (i.e., marginalise out) everything except the ordinal axis.

      (5) Conflation of RNN learning dynamics with working memory mechanisms. The authors show that rhythmic input makes learning marginally easier, but in principle, both RNNs reach full performance (so working memory can be done as well with either case). To avoid the findings depending on learning dynamics, it could be of interest to test the RNNs trained with jittered input on fixed input (or retrain RNNs with both jittered and non-jittered input). It is also unclear if the small increase in performance (0.4%) can be expected to hold across different initialisations and/or learning rate /regularisation strengths.

      (6) STSP. It is unclear if the findings rely on STSP being present or not (or what the role of STSP is in the model at all currently). Note that in Liebe et al. 2025, RNNs were trained on an almost identical task without STSP, and phase-coding was demonstrated in the models.

      (7) Writing redundancy. The methods subsections "Population signal construction for oscillatory analysis" and "Oscillatory phase organization during sample encoding" seem to define exactly the same quantity with different characters (activity projected in dPCA space), which leads to confusion (in one section, z is the PC component, in another, it's the complex signal). There are also slightly different definitions of the wavelets in either section, for which the reasoning is unclear.

    3. Reviewer #2 (Public review):

      Summary:

      The authors train E-I recurrent networks with short-term synaptic plasticity on a sequential delayed match-to-sample task, comparing regular versus jittered sample timing. They report a small accuracy gain under rhythmic input, a more separable population geometry during encoding, organization of internal oscillations around the dominant input frequency, a preference for temporal order over feature encoding, and improved decodability and persistence of stimulus information in both activity and synaptic efficacy. A delay-period perturbation shows synaptic efficacy contributes more than activity to maintenance.

      Strengths:

      The model is well-specified. Dale's law, the STSP formulation, the training objective, and the hyperparameters are all reported clearly enough to reproduce, and code is shared. The statistical machinery is appropriate, with cluster-permutation tests for the spectral analyses and across-network sign-flip tests rather than naive pooling. The temporal-order versus stimulus-direction dissociation in Figure 4C is the most interesting result. The negative association between phase locking and direction selectivity is non-trivial and argues against a simple global-gain reading, and it connects to Liebe et al. 2025. The serial-position decoding curves and the synaptic-versus-neuronal perturbation are well-motivated tests of the maintenance claim.

      Weaknesses:

      The behavioral effect is very small. Match accuracy is 0.991 versus 0.987, and non-match is 0.973 versus 0.969, on networks already at the ceiling. The entire mechanistic analysis is built to explain a roughly 0.4 percentage point difference, and the paper does not establish that this difference is functionally meaningful rather than a marginal byproduct of the timing manipulation. The IOI-dependence result meant to support it is weak, with an R-squared of 0.071 at a p-value of 0.029 on n of 67.

      The core spectral and phase results are close to definitional and should be framed that way. The regularity index R is computed from IOI variability, the dominant frequency f0 is computed from the same IOIs, and the oscillatory metrics in Figures 3 and 4 are then measured relative to f0 and correlated against R. This shows that more regular input produces internal phase progression closer to the input-derived reference frequency, partly restating the input statistics rather than uncovering an independent network mechanism. The phase-locking-increases-with-regularity finding is the clearest case. This does not invalidate the analyses, but the manuscript currently reads them as a mechanism when much of the signal is built into the measurement.

      The only genuinely causal manipulation is the delay-period shuffle, and it is underpowered at n of 15. Its main conclusion, that synaptic efficacy matters more than activity for maintenance, largely recovers prior STSP results (Mongillo et al. 2008, Masse et al. 2019) rather than establishing something specific to rhythm. The result the authors most want, that disrupting synaptic state removes the rhythmic advantage, is predicted in the Discussion but not tested.

      The authors should add a control that breaks the circularity (a held-out f0/phase reference, or shuffling R against the metric) and run the causal STSP-disruption test that is mentioned in the Discussion.

      The oscillatory framing is stronger than the model supports. Phase locking to a periodic input can reflect temporal predictability or repeated preparation without self-sustained entrainment, and the authors acknowledge this once but then use entrainment-style language throughout. The signals are extracted from firing-rate units and should not be read as LFP or EEG oscillations.

      The authors should show raw single-unit and population activity so readers can verify the oscillations before the filtered pipeline. The delay perturbation largely recovers Mongillo 2008 / Masse 2019 rather than anything rhythm-specific, and the relationship to Liebe et al. 2025 should be addressed in the Results.

      Appraisal and impact:

      The authors largely achieve their stated aim of describing how temporal regularity constrains recurrent dynamics in this model, and the temporal-order preference is a useful prediction. The reach of the conclusions exceeds the evidence in two places: the functional importance of the behavioral effect and the degree to which the phase results are independent of the input construction. With the framing corrected and one causal test added, this would be a useful contribution to the modeling literature on timing and working memory rather than a definitive account.

    4. Author response:

      We thank the editors and the two reviewers for their careful evaluation of our study, as well as for their positive assessment of the research question, model reproducibility, and the results concerning temporal-order representations. We also agree with the core issues raised in the reviews: the current manuscript has not yet sufficiently distinguished descriptive changes in recurrent dynamics from the functional mechanisms underlying the behavioral advantage; the behavioral effect itself is small and close to the performance ceiling; the phase analyses require more stringent controls; and the specific role of short-term synaptic plasticity in the rhythmic advantage has not yet been directly tested. We plan to add the corresponding analyses in the revised manuscript and to temper several mechanistic claims.

      First, we would like to clarify three aspects of the study design and analysis pipeline. First, the rhythmic and arrhythmic conditions were not performed by two separately trained groups of networks. Each network was jointly trained using balanced batches containing rhythmic-match, rhythmic-non-match, arrhythmic-match, and arrhythmic-non-match trials. Therefore, the behavioral differences were compared within the same independently initialized network. We will revise the relevant descriptions in the Abstract, Results, and Methods to make this joint-training procedure more explicit. Second, Figures 2 and 3–4 used different dimensionality-reduction approaches because they addressed different analytical questions. In Figure 2, dPCA was applied to time-aligned population activity to separate task-related variance into temporal, ordinal-position, and stimulus-related components, allowing us to examine how rhythmicity affected each representational component. In contrast, Figures 3 and 4 used standard PCA to construct a low-dimensional population signal for spectral and phase analyses without explicitly demixing task variables. We will clarify this distinction and the rationale for the two analysis pipelines in the revised Methods. Third, STSP was not introduced as an auxiliary module. Previous theoretical and computational studies have suggested that STSP can contribute to working-memory maintenance (Mongillo et al., 2008; Masse et al., 2019). Based on previous studies, we incorporated STSP alongside persistent neural activity to examine how synaptic and consistent neuronal activities jointly support sequential working memory. Under the current architecture and training settings, our ablation experiments showed that RNNs without STSP had difficulty successfully learning the task. We will emphasize this result and further distinguish the roles of STSP in task learning, delay-period maintenance, and the rhythmic advantage. We will also explore whether vanilla RNNs can successfully learn the same task under alternative hyperparameter settings.

      Our core hypothesis is that temporal regularity improves the encoding of sequential information by organizing recurrent population dynamics and phase structure during the encoding period, and that this organization subsequently influences information maintenance during the delay period and behavioral performance. The current results establish a relationship between temporal regularity and encoding-period dynamics, but direct validation of how this organization influences subsequent working-memory maintenance remains insufficient. In the revision, we will focus on strengthening the link between the encoding and delay periods and test whether population dynamics and phase organization during encoding are associated with subsequent information maintenance and behavioral performance.

      To provide a more direct view of the network dynamics, we plan to add intuitive visualizations of network activity and phase structure, allowing readers to evaluate the reported temporal organization with less dependence on dimensionality reduction, filtering, and complex statistical processing.

      The current behavioral advantage of approximately 0.4 percentage points is small, and network performance in both conditions is close to ceiling. Following the reviewers’ suggestions, we plan to evaluate the stability of this effect across learning, random initializations, and representative hyperparameter settings, and to examine whether the rhythmic advantage becomes more pronounced under higher memory load when ceiling effects are reduced. We will also avoid equating statistical significance directly with functional importance.

      Although the manuscript already includes decoding and perturbation analyses of neural activity and synaptic efficacy during the delay period, we will perform additional delay-period analyses to more directly examine whether the temporal organization established during encoding is associated with subsequent working-memory maintenance. These analyses will help distinguish the contributions of rhythmic input to stimulus encoding and memory maintenance.

      We will also test the role of STSP more directly. The current delay-period shuffle results show that synaptic efficacy makes a functional contribution to delay-period maintenance, but this does not yet demonstrate that STSP specifically supports the rhythmic advantage. We will further illustrate the roles of STSP during encoding and delay-period maintenance. In the revision, we plan to increase the number of independent networks in the perturbation analysis and test the interaction between temporal regularity and STSP disruption.

      Finally, we will further tighten the conceptual framing of the manuscript by more clearly distinguishing stimulus-driven phase organization from self-sustained oscillations, and by clarifying that the analyzed signals are model population signals derived from firing-rate population activity rather than LFP or EEG field potentials. Where the current evidence is insufficient to support interpretations in terms of entrainment or self-sustained oscillations, we will adopt more cautious terminology and revise the corresponding conclusions accordingly. We will also clarify the relationship between our phase-organization results and the findings of Liebe et al. (2025) in the revised Results.

      Overall, the revised manuscript will more precisely frame the study around how temporal regularity improves sequential working memory and how this behavioral advantage is associated with the organization of recurrent population dynamics and synaptic-state representations during encoding. After completing the additional training and analyses, we will also make the corresponding model weights, training configurations, and necessary analysis files publicly available to improve reproducibility.

      We again thank the editors and the two reviewers for their detailed and constructive comments. We will revise the manuscript carefully on the basis of these suggestions.

      References

      Mongillo, G., Barak, O., & Tsodyks, M. (2008). Synaptic theory of working memory. Science, 319(5869), 1543–1546. https://doi.org/10.1126/science.1150769

      Masse, N. Y., Yang, G. R., Song, H. F., Wang, X.-J., & Freedman, D. J. (2019). Circuit mechanisms for the maintenance and manipulation of information in working memory. Nature Neuroscience, 22(7), 1159–1167. https://doi.org/10.1038/s41593-019-0414-3

      Liebe, S., Niediek, J., Pals, M., Reber, T. P., Faber, J., Boström, J., Elger, C. E., Macke, J. H., & Mormann, F. (2025). Phase of firing does not reflect temporal order in sequence memory of humans and recurrent neural networks. Nature Neuroscience, 28, 873–882. https://doi.org/10.1038/s41593-025-01893-7

    1. eLife Assessment

      This important study provides new insight into how frontostriatal circuits encode elapsed time and exhibit decision-related dynamics during an auditory change-detection task. Analyses of simultaneously recorded neurons from the frontal orienting field (FOF) and anterior dorsal striatum (ADS) provide solid evidence that FOF and ADS show similar dynamics during evidence accumulation, but that FOF shows stronger movement-aligned reorganization near the time of the decision report. However, the mechanistic interpretation would be strengthened by clearer links between the population-level analyses, single-neuron activity, and relevant circuit anatomy. The work will be of broad interest to systems neuroscientists studying timing, decision-making, and frontostriatal dynamics.

    2. Reviewer #1 (Public review):

      Summary:

      The authors explore how temporal information and decision-related dynamics are represented across FOF and ADS in rats. The authors used Neuropixels to record neurons simultaneously from FOF and ADS during a free-response auditory change-detection task. They then applied single-trial temporal decoding to estimate both the time elapsed since stimulus onset and the time remaining until movement initiation. When neurons in both FOF and ADS were sorted based on decoder weights, they showed ramping and transient bump-like dynamics aligned to stimulus onset. However, around the decision report, FOF showed a clearer ramping signal and stronger movement-aligned population reorganization than ADS. These results suggest that FOF and ADS share similar temporal dynamics during evidence evaluation, but that FOF undergoes a stronger reorganization near decision commitment.

      Strengths:

      (1) The authors recorded large-scale neural populations simultaneously from FOF and ADS, allowing direct and fair comparison between them in the same sessions.

      (2) The free-response auditory change-detection task, which requires rats to evaluate sensory evidence over time and initiate a decision report, is suited to address the question. The behavioral results support that rats used sensory evidence to guide their choices.

      (3) The authors used multiple approaches, including single-trial temporal decoding, decoder-weight PCA, PC loading trajectory, and population-geometry analyses, to explore the FOF and ADS dynamics. These methods provide converging evidence supporting that FOF and ADS share similar temporal dynamics during evidence evaluation but diverge around movement/decision commitment.

      (4) The population-geometry analysis is quite strong and interesting because it compares epoch-specific neural subspaces and quantifies dimensionality and subspace alignment, showing stable subspaces during evidence evaluation and stronger subspace reorganization in FOF near movement initiation.

      Weaknesses:

      (1) The manuscript failed to include histological confirmation of probe placement.

      (2) The direct FOF-ADS decoding comparison in fig 3f and 4f includes only 16 of 61 sessions because of imbalanced unit counts. While controlling for unit number is important, excluding most sessions may waste data. Restricting analyses to only 16 sessions questions the generalizability of the result.

      (3) Fitted regression curves, and ideally confidence intervals, were missing from Figures 3c and 4c. Also, the confusion matrices in Figures 3a/b and 4a/b show a strong preference for predictions in the first and last time bins. The authors did not explain whether this reflects meaningful event-aligned neural activity or an endpoint artifact from decoding time as bounded discrete classes.

      (4) The interpretation of the neuron groups defined by PCA on the decoder-weight matrix was confusing. The authors perform PCA on an N units by T time-bin matrix of LDA decoder weights, then group neurons according to their PC1 and PC2 scores. This is an interesting approach, but the current wording could make readers think that neurons at the extremes of PC1 or PC2 are necessarily the most important neurons for temporal decoding. In fact, these groups appear to represent neurons whose decoder-weight profiles project strongly onto the dominant weight-space patterns. They are not necessarily the neurons that contribute most strongly to decoding accuracy, nor are they necessarily the most common firing-rate dynamics in the raw neural population.

      (5) Discussion is missing some needed context. First, given the causal role of ADS in evidence-accumulation-based choices (Yartsev et al., 2018), and its position as a key node that may integrate input from FOF (Brody & Hanks, 2016), the weaker decision-aligned transition in ADS compared with FOF should have been further discussed. If ADS contributes causally to the decision process, why does it show a much weaker population-state transition near decision commitment in the present data? Second, in DePasquale et al. (2024), more extensive choice vacillation was found in ADS, while greater choice certainty was found in FOF. Does this follow the same principle as the current manuscript, where FOF shows stronger reorganization near decision commitment compared to ADS?

      (6) Current analyses do not fully exploit the simultaneous nature of the recordings. Apart from the comparison of decoding accuracy, most analyses could have been performed and compared based on the data collected independently from two regions.

      (7) Figures 3-10 are hard to read and unpolished. Fonts are too small, and legends/labels are redundant.

    3. Reviewer #2 (Public review):

      Summary:

      This work investigated differences in the temporal dynamics of neural populations in frontal orienting fields (FOF) and anterior dorsal striatum (ADS) in rodents during an auditory change detection task. The relative roles of these two regions have been studied previously and have been shown to play a role in the accumulation of evidence, with FOF converting this evidence into a categorical decision. By focusing on the temporal dynamics of neurons in these regions, the authors identified a subpopulation of neurons within FOF that displayed an abrupt ramping of activity near the time of decision commitment. Both FOF and ADS contained subpopulations exhibiting ramping activity aligned to stimulus onset. This is an interesting finding, suggesting that FOF contains a subpopulation of neurons that transforms accumulating evidence from other subpopulations in ADS and FOF into an action.

      Strengths:

      The conclusions of this paper are mostly well supported by data.

      Weaknesses:

      (1) In the neural analysis, the authors use a technique in which the weights of a linear decoder are used to define a feature vector for each neuron. These weights are used to measure the overall contribution of a neuron in decoding time (from stimulus or decision commitment). Interpreting decoding weights in this way is technically not correct (Kriegeskorte and Douglas, "Interpreting encoding and decoding models"), as a large weight in a decoder is not necessarily indicative of a large effect. Weights in decoding models can become large in order to cancel noise. Alternative analyses, for instance, treating the time series of each neuron as a feature vector, could have supported the conclusions from this technique.

      (2) In this same analysis, it appears that the abrupt change in response in FOF at the time of decision commitment is coming from a single subpopulation of about 130 neurons. In the example session (Figure 8J), there is a clear outlier (the neuron in the top right corner). A closer inspection of the single neuron responses in this group would strengthen the results to confirm the abrupt change in mean population response is not coming from a relatively small number of neurons and sessions.

      (3) The significance of the dynamical motif corresponding to transient bumps was unclear. For example, when looking at Figure 6K-L, I do not see any neuron groups that exhibit a clear transient bump. I would characterize all groups as ramping, with some groups showing steeper ramps. It would be helpful if the figure displayed the fraction of variance explained by PC2 so that it would be clear how much variance the bump motif is contributing. Given that there was no discussion of the functional relevance of this second motif, interpretation of this result is unclear.

      (4) The finding that FOF contains subpopulations which slowly ramp during the trial as well as a subpopulation which acts like a switch that abruptly turns on at the time of decision commitment is interesting and significant and presents several computational questions. For example, is this subpopulation a non-linear readout of the more slowly ramping populations? The approach based on constructing a feature vector for each neuron, projecting these vectors into a low-dimensional subspace, and partitioning into subpopulations is insightful and allowed distinguishing these different computational functions within a single region (FOF). However, I found this particular result to not be clearly stated and obscured by other seemingly less significant results (e.g., existence of the transient bump motif) and other less interpretable analyses (e.g., subspace re-alignment).

    4. Reviewer #3 (Public review):

      Summary:

      This study investigates how frontostriatal circuits encode elapsed time and exhibit decision-related dynamics during an auditory change-detection task. Using population-level temporal decoding and analyses of low-dimensional neural dynamics, the authors compare activity in the frontal orienting field (FOF) and anterior dorsal striatum (ADS). The manuscript addresses an important question in systems neuroscience: how cortical and striatal circuits represent elapsed time and signal action initiation during decision-making.

      The results suggest that FOF and ADS differ in how they represent decision-related information near decision commitment or behavioral report. In particular, FOF shows greater movement-aligned changes in temporal decoding and population geometry than ADS. These findings are potentially important because they may help clarify how cortical and striatal circuits contribute to timing, decision formation, and action initiation.

      Strengths:

      A major strength of the study is its use of population-level analyses to identify temporal structure and movement-aligned changes in neural dynamics. The analyses provide evidence that neural dynamics and low-dimensional population geometry change around the time of behavioral report, especially in FOF. This provides a useful population-level description of decision-related dynamics beyond what could be inferred from average firing rates alone.

      Another strength is that FOF and ADS activity were recorded simultaneously during the same auditory change-detection task. This design strengthens the regional comparison by minimizing confounds related to session-to-session variability, including differences in task engagement, decision accuracy, or other behavioral variables across recordings. The simultaneous recordings therefore provide a strong basis for comparing temporal decoding and population dynamics between cortical and striatal circuits.

      Weaknesses:

      One limitation is that the physiological interpretation of the population-geometry analyses remains somewhat abstract. Concepts such as low-dimensional subspaces, subspace alignment, and subspace rotation are potentially powerful, but it is not always clear what specific changes in neural activity give rise to these effects. For example, it is difficult to tell whether changes in population geometry primarily reflect recruitment of different neurons, or changes in the dominant temporal profiles of the same neurons. This limits the physiological interpretability of the population-level findings.

      A second limitation is that the mechanistic interpretation of the FOF-ADS difference remains underdeveloped. The observed differences could reflect an internally generated transition in frontostriatal dynamics, similar to the dynamical-regime and neural-mode transition described by Luo et al. (2025). Alternatively, they could reflect a circuit-readout process, analogous to the framework proposed by Stine et al. (2023), in which cortical activity drives threshold crossing in a downstream circuit, triggering orienting or motor signals that terminate the decision process. The current manuscript describes the regional differences clearly, but it does not fully discuss these mechanistic interpretations.

      Finally, the strength of the evidence would be easier to evaluate if the manuscript more clearly reported the number of animals contributing to each major analysis and the consistency of the main effects across animals. Because many analyses are performed across sessions, the absence of this information makes it difficult to assess whether the key findings are robust across animals or could be influenced by one or a small number of animals.

    1. eLife Assessment

      This important study combines behavioral testing, fiber photometry recordings, and optogenetic manipulations to understand the role of CRH+ neurons in the PVN of the hypothalamus in social and non-social settings, while varying the degree of familiarity. The approaches used and the finding that these neurons respond to various social (and non-social) stimuli are strong; however, the conclusion that the activity of these neurons is driven solely by unfamiliar situations and the lack of dynamic analysis of fiber photometry data leaves the manuscript incomplete. This elegant study will likely be of interest to the field of social neuroscience and, if outstanding issues are addressed, would expand our understanding of the role of the PVN in behavior.

    2. Reviewer #1 (Public review):

      Summary:

      Here, the authors examine how CRH neurons in the PVN track social behaviours. They use fiber photometry to record the bulk activity of PVN CRH neurons during the resident-intruder test. They find that PVN CRH activity increases when the intruder enters, and also when mice make movements to approach the intruder. They further show that the magnitude of this response differs depending on the familiarity of the mouse. Specifically, if the intruding mouse is unfamiliar, there is a greater PVN CRH response relative to a familiar mouse. The authors argue that this is specific to social familiarity, as they do not see the same differentiation in the PVN CRH response when mice approach a familiar or unfamiliar object. Finally, the authors conduct optogenetic experiments and show that inhibition of PVN CRH neurons reduces social investigative behaviour. The authors then conclude that PVN CRH neurons are a part of a decision-making circuit to influence behaviour in ambiguous settings, specifically that they are a "key component of the neural circuitry underlying rapid social appraisal, linking endocrine regulation to real-time behavioural decision making".

      The data are interesting and novel. They help us understand the dynamics and range of situations in which PVN CRH neurons are activated. There is some overinterpretation of the data and restriction of what this signal means (i.e., specifically driven by unfamiliar social situations), which doesn't seem to be supported by the data. Indeed, PVN CRH neurons are robustly activated by scenarios outside unfamiliar social ones.

      Strengths:

      The experiments are run and presented very beautifully in a sophisticated way. The data are novel and interesting. They help us understand the time course of PVN CRH responding in social and object settings, and how this differs with the familiarity of a social stimulus.

      The optogenetic manipulation is also very nice. The authors optically inhibit during just the first 20 seconds of the resident-intruder test. They find that this inhibition results in a long-term reduction in social behaviours. To me, this supports an idea that the PVN CRH signal triggers a cascade of behaviours, but is not necessarily driving these behaviours per se.

      Weaknesses:

      It would be great to see more sophisticated analysis of the fiber photometry data, which may reveal interesting effects that are currently being occluded by static AUC analysis. One pipeline that is freely available that could be used is found in Jean-Richard-dit-Bressel, Clifford, and McNally (2020) Frontiers in Molecular Neuroscience. Referred to as waveform analysis, this would allow the authors to examine the significance of their data across time. There are multiple points at which this would be interesting. For example, in Figure 3F, it is possible that differences between the familiar and unfamiliar objects emerge. Also, there seems to be one outlier in this figure in the familiar object group. What happens if it is removed (Figure 3H)?

      Similarly, what do these signals look like when aligned with making contact with the social or object stimuli? It is possible that the objects do not elicit a difference depending on familiarity when approaching because: (1) they are not moving, and (2) it is unclear whether they are familiar or not until contact is made, consistent with the object recognition literature. What would inhibition of the PVN CRH signal do to investigative behaviours directed towards objects?

      Finally, given the robust nature of the response to the approach to the objects and familiar mouse, why is this signal being argued to predominantly act in unfamiliar social settings? The lack of difference between the familiar and unfamiliar objects doesn't negate the importance of this signal. To me, this is the most interesting finding: PVN CRH neurons that are usually activated in stressful situations can also be robustly activated by familiar objects. Relatedly, while the authors argue that inhibition of PVN CRH neurons only reduces social behaviours in the unfamiliar case, there is likely a floor effect in the behaviours that they are looking at, which occludes observation of a reduction via optical inhibition.

    3. Reviewer #2 (Public review):

      Summary:

      The authors investigated the role of hypothalamic CRH neurons in social behavior. They performed fiber photometry recordings in mice from CRH neurons and showed that novel conspecifics trigger stronger and more prolonged responses compared to familiar conspecifics and objects. The activity of CRH neurons appears to be related to risk assessment, as interactions with juvenile unfamiliar mice (lower-risk conspecifics) trigger responses similar to those of familiar adult mice. Behaviorally, CRH neurons were linked to increased anogenital investigation of unfamiliar compared to familiar mice. Optogenetic suppression of CRH neurons decreased anogenital sniffing of unfamiliar conspecifics.

      Strengths:

      The manuscript is elegant, and the results are compelling. The approaches are well justified, and the methods are validated (eg: Arch inhibition).

      The findings substantiate the role of CRH neurons in responses to stress and uncover the involvement of these neurons in the assessment of social risk.

      Weaknesses:

      These are not weaknesses, just some observations: It is somewhat surprising that CRH neurons respond similarly to familiar and unfamiliar objects; it would be good to have more insights into that aspect.

      Similarly, the novel context by itself is expected to lead to increased activity of CRH neurons (based on data from the last author's lab as well as other labs in the field). It is somewhat surprising (and interesting) that the novel environment did not affect the magnitude of CRH responses to unfamiliar conspecifics.

    1. eLife Assessment

      This important work provides new insights into the role of lysine acetylation of alpha-synuclein, the protein involved in Parkinson's Disease. The evidence is convincing and the work will be of interest to researchers in the fields of protein biophysics and post-translational modifications.

    2. Reviewer #1 (Public review):

      [Editors' note: the authors have revised the work in response to the original reviews.]

      Summary:

      This paper describes experiments with alpha-synuclein (aS) with acetylated lysines (acK) at various positions. Their findings on how to use non-canonical amino acid (ncAA) mutagenesis to generate aS with acetylated lysines are valuable. The paper then continues with a range of experiments to characterise the acetylated alpha-synuclein constructs at different positions, with the aim of providing insights into which sites are relevant to disease or their function inside cells. The paper concludes these experiments with the suggestion that inhibiting the Zn2+-dependent histone deacetylase HDAC8 to potentially increase acetylation at lysine 80 may have therapeutic benefit. However, the relevance of most of these experiments is unclear, mainly as the filaments that form from these constructs are different from those observed in human disease (but see below for more details). Moreover, using the recombinantly produced acetylated versions of alpha-synuclein to normalise mass-spectrometry data, the authors themselves report that acetylation of alpha-synuclein does not differ between individuals with Parkinson's disease or healthy controls.

      Strengths:

      The authors report difficulties with chemical synthesis and then decide to make these constructs using non-canonical amino acid (ncAA) mutagenesis, which seems to work reasonably well (yields vary somewhat). In the Conclusion section, the authors report that they used these recombinant proteins to obtain quantitative insights into the levels of acetylation of lysines in individuals with PD versus healthy controls, for which they find no significant differences. This part of the work is valuable.

      Weaknesses:

      The authors then use circular dichroism to show that aSyn with acK at position 43 has less alpha-helical content. From this result, they deduce that "only this site could potentially perturb aS function in neurotransmitter trafficking", but no experiments on neurotransmitter trafficking were performed.

    3. Reviewer #2 (Public review):

      Summary:

      Shimogawa et al. studied the effect of lysine acetylation at different sites in the alpha-synuclein (aS) sequence on the protein-membrane affinity, seeding capacity in the test tube and in cells, and on the structure of fibrils, using a range of biophysical methods. They use non-canonical amino acid (ncAA) mutagenesis to prepare aS lysine acetylated variant at different sites.

      Strengths:

      The major strength of this paper is the approach used for the production of site-specific lysine acetylated variants of aS using ncAA mutagenesis, as well as the combination of a range of biophysical methods together with cellular assays and structure biology to decipher the effect of lysine acetylation on aS-membrane binding, seeding propensity, and fibril structure. This approach allowed the author to find that lysine acetylation at positions 12, 43, and 80 led to lower seeding capacity of aS in the test tube and in cells, but only acetylation at lysine 80 did not affect aS-membrane interaction. These results suggest that lysine acetylation at position 80 may be protective against aggregation without perturbing the proposed functional role of aS in synaptic plasticity.

      Weaknesses:

      SDS is not a good membrane model to investigate the effect of lysine acetylation on aS membrane-binding because it is a harsh detergent and solubilizes membranes. Negatively charged vesicles or vesicles made of a mixture of lipids mimicking the lipid composition of synaptic vesicles are more accepted in the field to study aS-membrane interactions. The authors used such vesicles for the FCS experiments, and they could be used for the initial screening of the 12 lysine acetylated variants of aS.

    4. Reviewer #3 (Public review):

      Shimogawa et al. describe the generation of acetylated aSyn variants by genetic code expansion to elucidate effects on vesicle binding, aggregation, and seeding effects. The authors compared a semi-synthetic approach to obtain acetylated aSyn variants with genetic code expansion and concluded that the latter was more efficient in generating all 12 variants studied here, despite the low yields for some of them. Selected acetylated variants were used in advanced NMR, FCS, and cryo-EM experiments to elucidate structural and functional changes caused by acetylation of aSyn. Finally, site-specific differences in deacetylation by HDAC 8 were identified.

      The study is of high scientific quality, and the results are convincingly supported by the experimental data provided. The challenges the authors report regarding semi-synthetic access to aSyn are somewhat surprising, as this protein has been made by a variety of different semi-synthesis strategies in satisfactory yields and without similar problems being reported.

      The role of PTMs such as acetylation in neurodegenerative diseases is of high relevance for the field, and a particular strength of this study is the use of authentic acetylated aSyn instead of acetylation-mimicking mutations. The finding that certain lysine acetylations can slow down aggregation even when present only at 10-25% of total aSyn is exciting and bears some potential for diagnostics and therapeutic intervention.

    5. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Weaknesses:

      The authors then use circular dichroism to show that aSyn with acK at position 43 has less alpha-helical content. From this result, they deduce that "only this site could potentially perturb aS function in neurotransmitter trafficking", but no experiments on neurotransmitter trafficking were performed.

      We agree with the reviewer that neurotransmitter trafficking studies would be interesting, but they would presumably require the use of acetylation mimic mutants (Lys-to-Gln mutations), which we would want to validate by comparison to our semi-synthetic proteins with authentic AcK. Such experiments are planned for a follow-up manuscript, and we will investigate the reviewer’s suggested experiment at that time. Thus, we have not modified the manuscript to address this issue.

      Subsequently, they measure the aggregation speed of the variants in seeded aggregation experiments with preformed fibrils (PFFs) from WT aSyn, and conclude that acK at positions 12, 43, and 80 yields slower aggregation. They reach similar conclusions when measuring seeded aggregation in primary cultures. As far as I understand it, the seeding experiments in cells use seeds that are assembled from partially acetylated alpha-synuclein, but that are made of non-acetylated wildtype alpha-synuclein, and the alpha-synuclein that is endogenous in the cells is also non-acetylated (or at least not beyond what happens in these cells at endogenous levels). It is therefore unclear how the cellular seeding experiments relate to the in vitro aggregation assays with (partially) acetylated substrates.

      We understand the reviewer’s concerns and have modified the manuscript to clarify that the method of in vitro seeding really reports on the impact of acetylation on the elongation phase of aggregation. We have also clarified that this is different than the role that acetylation plays in seeding cellular aggregation with pre-acetylated fibrils. We note that having the monomer population acetylated in cells presents technical challenges that might also be addressed with Gln mutant mimics, and we plan to pursue such experiments in the follow-up manuscript described above.

      Anyway, both aggregation experiments ignore that the structures of aSyn filaments in Parkinson's disease (PD) or multiple system atrophy (MSA) are different from those formed in these experiments, and that, therefore, the observed aggregation kinetics are likely irrelevant for the speed with which disease-relevant filaments form in the brain.

      Finally, the authors describe the cryo-EM structure of mixtures of acK80:WT aSyn filaments, which are predominantly made of WT aSyn, with a previously described structure. Filaments made of only acK80 aSyn have a modified arrangement of this structure, where the now neutral side chain of residue 80 packs inside a hydrophobic pocket. The authors discuss differences between the acK80 structures and those of other structures from in vitro assembled aSyn filaments, none of which are the same as those observed from PD or MSA brains, nor are any attempts made to transfer observations from the in vitro experiments to the structures of disease. The relevance of the cryo-EM structures for human disease, therefore, remains unclear.

      The Conclusion on p.20 mentions an interesting and valuable result: the authors used the acetylated recombinant proteins to determine the extent of acetylation within human protein samples by quantitative liquid chromatography MS (SI, Figures S41-S49). Their conclusion is that "The level of acetylation was variable - no clear trend was observed between healthy control and patients - nor between patients of different diseases (SI, Table S4, Supplementary Data 1)" This result implies that acetylation of aS is not directly related to its pathogenicity, which again adds doubts on the disease-relevance of the results described in the rest of the paper.

      We acknowledge the concerns raised in the above paragraphs and believe that they can all be addressed by clarifying our purpose. The different fibril polymorphs adopted in PD and MSA are likely the result of an interplay of many PTMs and non-proteinaceous cofactors. Therefore, we are not necessarily trying to claim that our AcK80 fold is populated in health or disease, but that by driving Lys80 acetylation, one could push fibrils to adopt this conformation, which is less aggregation-prone. A similar argument has been made in investigations of alpha-synuclein glycosylation and phosphorylation. Our results in Figure 9 imply that Lys80 acetylation could be increased with HDAC8 inhibition. We have revised the manuscript to make these ideas clearer, while being sure to acknowledge the limitations noted by Reviewer #1.

      Reviewer #2 (Public review):

      Weaknesses:

      SDS is not a good membrane model to investigate the effect of lysine acetylation on aS membrane binding because it is a harsh detergent and solubilizes membranes. Negatively charged vesicles or vesicles made of a mixture of lipids mimicking the lipid composition of synaptic vesicles are more accepted in the field to study aS-membrane interactions. The authors used such vesicles for the FCS experiments, and they could be used for the initial screening of the 12 lysine acetylated variants of aS.

      We have noted this shortcoming in revisions of our manuscript, but have not performed new experiments as we do not believe that using vesicles instead would change the conclusions of these experiments (that only AcK43 produces an effect, and a modest one at that).

      It would help the reader to have the experimental details (e.g., buffer, protein/lipid concentrations) for the different assays written in the figure legend.

      We have added additional detail to the figure captions.

      The authors use an assay consisting of mixing 10% fibrils + 90% monomer to investigate the effect of lysine acetylation on aS. However, the assay only probes fibril elongation and/or secondary processes. The current wording can be misleading, and the term aggregation could be replaced by seeding capacity for clarity. For example, the authors state that lysine acetylation at sites 12, 43, and 80 each inhibits aggregation, but this statement is not supported by the data. Instead, the data show that the acetylation at these sites slows down the fibril elongation and thus decreases the seeding capacity of aS fibrils. In order to state that lysine acetylation has an effect on aS aggregation, fibril formation, the author should use an assay where the de novo formation of fibrils is assessed, such as in the presence of lipid vesicles or under shaking conditions.

      As noted in our response to Reviewer #1, we have clarified which phase of aggregation we were investigating in our in vitro experiments.

      It is not clear from the EM data that the structures of the different lysine acetylated variants are different, unlike what is stated in the text.

      We feel that it is clear from structures in Figure 8 and the EM density maps in Figure S38 that the AcK80 fold is indeed different. Although the overall polymorphs are somewhat similar to WT, the position of K80 clearly changes upon acetylation, altering the local fold significantly and the global fold more moderately. We have added backbone RMSD calculations to quantify the differences in WT and AcK80 folds in Figure SX. They differ by ~5 Å in the fibril core region.

      Reviewer #3 (Public review):

      Weaknesses:

      The challenges the authors report regarding semi-synthetic access to aSyn are somewhat surprising, as this protein has been made by a variety of different semi-synthesis strategies in satisfactory yields and without similar problems being reported.

      We understand the reviewer’s surprise and have edited the manuscript to clarify that the NCL yields were not unusually low, but were comparable to ncAA yields, and since it is significantly easier to scan AcK positions using ncAAs, we felt that ncAAs are the method of choice in this case.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) Cryo-EM data processing: particles were pre-processed in cryosparc and ChatGPT-generated scripts. Are priors on tilt and psi angles defined through this procedure? And are segments from individual filaments kept strictly in the same half-sets for gold-standard estimation of resolution? Absence of psi and tilt priors will lead to worse refinements than otherwise possible. Worse, a mixture of segments from different filaments into the two half-sets may lead to overestimated resolution estimates.

      We thank the reviewer for their thoughtful analysis of our cryo-EM data processing approach. In response, we have modified our approach and added additional commentary on processing to the Materials and Methods section.

      “CryoSPARC (.cs) files were then converted to RELION STAR files using the PyEM csparc2star.py script. Following this data conversion, a custom Python script developed with ChatGPT precisely determined the start and end coordinates of each fibril. This script processed the csparc2star.py star file output to create coordinate pairs based on the cryoSPARC Fibril ID, notably without transferring the original tilt and psi angles from CryoSPARC to RELION. The output of this custom script served as the input for the autopick RELION extraction step.

      To handle curved fibrils, a special segmentation strategy was implemented: the script traced the coordinates along the fibril, generating a new start and end coordinate pair for individal segments, with the segment's end coordinate assigned either after spanning 10 particles (around 50 nm in length) or when the end of the fibril was reached. This resulted in shorter, straighter segments for processing. Finally, the createAutopick function from cryolo_boxmanager_tools.py in crYOLO was utilized to generate a STAR file linking the final particle coordinates to their movie files, which was then used to perform particle extraction in RELION. The standard RELION image processing pipeline then followed, including 2D classification, refinement, and 3D classification.

      This approach resolves the issue with curved fibrils, but also results in all picked fibrils being 50 nm or shorter in RELION. Consequently, segments from the same fibril receive unique fibril IDs in RELION and may be split into different half-maps. Although this avoids random splitting of individual particles without regard to their origin from the same fibril, it could still cause an overestimation of resolution.

      To assess any potential overestimation of resolution, another script was created to reassign the fibril ID of each particle in the RELION STAR file to that of the closest particles in the cryoSPRAC data, effectively ensuring that all particles from an individual fibril have the same fibril ID and are not assigned to different half-maps. This revised STAR file was then used as the input images STAR file for 3D auto-refinement in RELION. A comparison of maps before and after reassigning the fibril IDs shows negligible effects on the resulting structures and on their estimated resolution, indicating that the original procedure did not results in any significant overestimation of the resolution.”

      Author response image 1.

      RELION Cryo-EM maps before (gray) and after (yellow) fibril ID reassignment for (A) WT-A (B) WT-B (C) <sup>Ac</sup>K<sub>80</sub>-A (D) <sup>Ac</sup>K<sub>80</sub>-B, showing that the new processing approaches did not significantly change the fibril structures.

      (2) The 25% acK80 structure in S52 is understood to be wildtype-only. The authors mention in the main text that this is because of the strong density for the K80 side chain. An additional argument would be that an acK80 would leave an unshielded negative charge on the neighbouring E46, as K80 and E46 form a salt bridge in this structure.

      We appreciate the reviewer’s idea and have included a comment on the salt bridge impact.

      (3) It would be valuable to include side views of the density for all reported reconstructions to assess to what extent the beta-rungs are separated.

      The requested side views have been included in Figure SX.

      (4) Methods sections should be moved into the main text of the paper.

      This Materials and Methods portion of Supporting Information has been moved to the main text.

      Reviewer #3 (Recommendations for the authors):

      (1) We suggest removing "all" from the manuscript title as this claim might not hold up in the future.

      We understand the reviewer’s concern. Our title was meant to imply “all currently known” rather than “all” forever, but as this wording is awkward, we have deleted “all” as suggested.

      (2) The white font in Figure 1A is sometimes hard to read, especially on the yellow-green background between amino acids 70-80.

      We have changed this to black font.

      Figure 1 panels C-E are not referenced in the manuscript text?

      References to the Figure 1 panels have been added.

      In panel C, a structure is predicted, but based on what data? Why is there both a small and a big structure in panel C?

      Explanations of the Figure 1C images have been added to the caption.

      (3) Typo ε-acetyllysine -> Nε-acetyllysine

      This has been corrected throughout.

      (4) You report solubility issues during NCL, and these are typically alleviated by the use of chaotropes during ligation. Please specify what you mean by "standard NCL conditions" and include parameters such as guanidine concentration, pH, concentrations, volumes, and temperature.

      These details have now been added to the methods section.

      (5) According to Figure 2D, the desulfurization was incomplete. Please explain.

      The small peak observed next to the product peak is an adduct with sinapic acid (+206 Da), the matrix mixed in for MALDI acquisition. We do not see a sign of +32/64Da peak which would correspond to incomplete desulfurization

      Author response image 2.

      (6) Please include sequences of your constructs for ncAA mutagenesis. Especially, which intein was attached at what position. This is important for other groups to fully understand the production of acetylated aSyn variants.

      The full DNA sequence has been added to Supporting Information. This plasmid has also been reported previously in the referenced publications.

      (7) You state ncAA mutagenesis "yielded 0.11-1.5 mg" aSyn. I guess this is per-liter culture expression medium?

      This has been corrected.

      (8) The resolution/DPI of Figures 4, 5, S14-16, and S18 should be increased. They look blurred compared to Figures 6 or S19.

      These figures have been updated.

      (9) You provided only summaries in Figures 4 and 5 because of space restrictions. However, it would be nice to have at least some selected individual experiments (WT, acetylation K12/43/80) right next to it without the need to switch to S14, S15, or S16.

      Select experimental data has been added to main text Figures 4 and 5 as requested.

      (10) Please increase the size of the microscopy images in Figure 6 (in Figure S19, it looks much better).

      This size of the images in Figure 6 has been increased.

      (11) The labelling A1-A3 in Figure S33 was somehow unclear to me.

      These three panels show different sections of the TEM grid, illustrating heterogeneity in the 25% <sup>Ac</sup>K<sub>12</sub> fibrils that was not observed for fibrils of the other acetylation variants. We have clarified this in the figure caption.

      (12) In Figure 9, you present preliminary results regarding site-specific deacetylation by HDAC8. These results are interesting, but considering the exploratory nature of this in vitro experiment, I suggest toning down the highly enthusiastic discussion of potential in vivo effects.

      We understand the reviewer’s concern and have mitigated the claims of potential impact from these in vitro results.

    1. eLife Assessment

      The authors propose the existence of an AQP4-positive tanycyte-associated canal system in the hippocampus based on histological, ultrastructural, immunohistochemical, and RNA-based approaches. There are serious concerns regarding the evidence, and the identification of the tanycyte-related structures can be questioned. At this stage, the evidence for the central claims of the manuscript has to be regarded as inadequate.

    2. Reviewer #1 (Public review):

      In the manuscript by Fabian-Fine et al., the authors employ neuroanatomy to investigate aquaporin-4 expression in cells they consider tanycytes and their supposed involvement in tau tangles and amyloid-beta plaques in the hippocampus. This study includes samples from three mice and two Alzheimer's disease (AD) patients.

      My key concern and question is whether the cells presented in the manuscript are tanycytes. Tanycytes are specialized ependymoglial cells located in the circumventricular organs and are known to express specific markers. Importantly, they are not myelinated cells, which is a crucial distinction that the authors do not address.

      Additionally, the methodologies described in the manuscript lack clarity and controls. For instance, the use of Cdh5-GCaMP882 mice is not adequately justified. It is unclear what these mice contribute to the study's objectives, particularly concerning the aim of investigating waste removal processes in the brain. Moreover, the rationale behind the purported "fluorophore uptake experiments" is unclear and appears to involve the uptake of fluorophore-labeled goat anti-rabbit secondary antibody, which seems implausible to me.

      The hypotheses and claims presented in this manuscript are not sufficiently substantiated and are conceptually unclear. The notion that amyloid beta and tau proteins play structural roles in a hypothesized "tanycytes"-derived canal network is not sufficiently supported by the evidence. Furthermore, the study lacks rigorous data to convincingly establish the proposed interactions between these proteins and the processes of waste internalization.

      In conclusion, due to conceptual and methodological issues, I consider the current evidence as inadequate to support the primary claims.

    3. Reviewer #2 (Public review):

      Summary:

      In this study, the authors propose the existence of an AQP4-positive tanycyte-associated canal system in the hippocampus and suggest that this system participates in waste clearance and contributes to Alzheimer's disease pathology. Using histological, ultrastructural, immunohistochemical, and RNA-based approaches, the manuscript attempts to reinterpret amyloid-β plaques and tau-associated structures as components of a tanycyte-derived waste-internalization system. The work is conceptually ambitious and raises observations that may stimulate discussion regarding glial organization and waste clearance in the diseased brain.

      Strengths:

      A strength of the manuscript is the combination of imaging modalities and anatomical observations across mouse and human tissue. Some of the reported morphological features are intriguing and may warrant additional investigation. The study also attempts to integrate structural observations with broader hypotheses regarding neurodegeneration and Alzheimer's disease.

      Weaknesses:

      The central interpretation depends almost entirely on identifying the observed hippocampal structures as tanycytes, and the evidence supporting this conclusion remains insufficient. Tanycytes are classically associated with ventricular regions in circumventricular organs, particularly in the third ventricle and median eminence region, yet the manuscript does not provide sufficiently specific anatomical or molecular evidence to convincingly distinguish the described structures from astrocytic, ependymal, radial glial-like, oligodendroglial, myelin-associated, vascular-associated, or degenerative elements. The marker profile used throughout the study, particularly the reliance on AQP4 labeling and Luxol-positive structures, is not sufficiently selective to establish tanycyte identity, especially in pathological tissue where reactive glial changes may occur.

      This becomes particularly important because the manuscript repeatedly interprets Luxol-positive and myelin-associated structures as tanycytic processes or "myelin-derived tanycyte protrusions," despite tanycytes not being known to produce myelin. Alternative explanations are not sufficiently explored. Some of the canal-like structures shown in Figure 4 also resemble vascular profiles, and additional vessel markers would be necessary to exclude this possibility.

      Several of the proposed structures and mechanisms are also difficult to reconcile with established cell biology and neuroanatomy. The introduction of new terminology such as "tanysomes," "waste receptacles," and "toroids" further extends the interpretation beyond what is currently demonstrated experimentally.

      The discussion and integration of the existing literature on tanycytes are also insufficient. Tanycytes themselves are not clearly introduced; the manuscript does not adequately discuss what is currently established regarding tanycyte anatomy, ventricular localization, morphology, and function. Foundational literature defining tanycyte biology, including work from the Prévot group or others, is largely absent despite its central importance to the field. Because the manuscript proposes a substantial departure from established neurobiological concepts, it is particularly important that previous literature be discussed comprehensively and critically. The current version does not sufficiently contextualize the proposed model within the existing literature on tanycyte, AQP4, glymphatic, and Alzheimer's disease, making it difficult to evaluate what is genuinely novel versus what is merely being reinterpreted. It is also not entirely clear what is genuinely new here compared with the authors' previous work, particularly reference 11, which appears to present a highly similar conceptual framework.

      More broadly, several of the manuscript's mechanistic conclusions extend well beyond the available evidence. The proposal that amyloid-β plaques and tau pathology represent hypertrophic tanycyte-derived waste structures is provocative and potentially interesting, but currently remains largely correlative and speculative. At several points, it becomes difficult to distinguish direct observations from broader mechanistic interpretation. The manuscript itself acknowledges that the proposed glial-canal hypothesis contradicts the current understanding of nervous system organization and states that ultrastructural serial-section analysis would be required to unambiguously determine the origin of the myelinated profiles described. This point is critical because the study's central conclusions depend on the assumption that these structures are tanycyte-derived. At present, this interpretation remains insufficiently demonstrated, which substantially limits the strength of the broader pathological and mechanistic conclusions proposed throughout the manuscript.

      Although access to human material is understandably limited, the study appears to include only one male and one female AD patient, making it difficult to assess the reproducibility or frequent these structures are across individuals and pathological conditions. The manuscript would benefit from clearer characterization of prevalence, reproducibility, and variability across samples.

      Overall, the manuscript presents an unconventional and thought-provoking model that may stimulate discussion. However, the evidence currently provided does not convincingly establish tanycyte identity for the described hippocampal structures, and several of the broader disease-related interpretations would require substantially stronger anatomical and molecular evidence before the proposed model can be convincingly supported.

    4. Author response:

      Reviewer #1 (Public review):

      My key concern and question is whether the cells presented in the manuscript are tanycytes. Tanycytes are specialized ependymoglial cells located in the circumventricular organs and are known to express specific markers. Importantly, they are not myelinated cells, which is a crucial distinction that the authors do not address.   

      We agree with the reviewer that tanycytes that have been described in the third ventricle have not been reported to be myelinated. However, our study was conducted on the hippocampal formation that borders the ventral horn of the lateral ventricles. We will include images of the myelin-forming ependymal cells that we refer to as tanycytes  

      Additionally, the methodologies described in the manuscript lack clarity and controls.

      We will expand on our methods section and include controls.

      For instance, the use of Cdh5-GCaMP882 mice is not adequately justified. It is unclear what these mice contribute to the study's objectives, particularly concerning the aim of investigating waste removal processes in the brain. Moreover, the rationale behind the purported "fluorophore uptake experiments" is unclear and appears to involve the uptake of fluorophore-labeled goat anti-rabbit secondary antibody, which seems implausible to me.

      Most experiments described in this manuscript were carried out on human brain. However, functional studies will have to be carried out on rodent brain. We will thus process rodent tissue as well to test whether our observed findings are consistent between human and rodent.  The animals used for these experiments were raised for bladder research and are wild type regarding neuronal and glial cells, particularly using the Cy3 channel. Utilizing these brains for our experiments has allowed us to test rodent tissue at both light- and electron-microscopic levels without having to sacrifice additional animals. The consistency of our findings between human and rodent brain further supports that the calcium indicator in the vascular system of these mice did not affect neurons or glial cells.

      Regarding the uptake experiment: When we initially discovered that myelin-forming macroglia form waste-internalizing glial canals within neuronal in spider brain it was unclear where the AQP4-immunoreactive cells were located. The somata of the myelin-forming cells lacked AQP4 immunoreactivity. Suspecting a synergistic interaction between the myelin-forming and AQP4 expressing cells whereby the myelinforming cells create the canal structure that sequesters waste from the neuron and the AQP4-expressing cells create a convective flow toward the waste-internalizing structures. However, unable to locate the somata of these cells, we submerged a freshly dissected spider brain with the attached surrounding tissue intact in physiological spider saline and slowly added blue vital dye solution to test which cells would internalize the dye. We then identified the (blue) cells in the lining of the dorsally located tubular system that we routinely detached from our brain preparations explaining why we were unable to locate these cells.  Immunolabeling of this tubular system revealed the cells that reside in the lining of this tubular system (see Author response images 1 and 2) the original (Figure 8) shows their long slender processes. Interestingly, this system is continuous with the stomatogastric system. 

      Author response image 1.

      Shows the proposed canal system in spiders

      Author response image 2.

      AQP4-immunoreactive cells in the spider primitive ventricular system that we localized due to similar uptake experiments we have conducted in mouse brain.

      We have utilized this method in mouse brain to test the validity of our postulation, that ependymal tanycytes internalize the presented fluorochrome from extracellular spaces and test which areas and structures may be involved in this uptake. As demonstrated in this experiment, the alveus, and a fine network of cell processes within the brain parenchyma show fluorescence, indicative that they internalize the fluorochrome from extracellular spaces. We used goat-coupled fluorochrome to further test with a FITCcoupled secondary antibody that the observed fluorescence is indeed due to uptake of the goat-coupled secondary antibody and not due to intrinsic autofluorescence. Control preparations lacked this fluorescence.  To further test our postulation that the uptake is indeed AQP4-mediated we have applied an AQP4-blocker, which showed a significantly reduced fluorochrome uptake compared to the controls without this blocker.

      As we state in the text, we are aware of the limitations of this experimental design, however, like in our spider experiments we consider these findings helpful as they likely show an overview of the cellular network in the hippocampus that governs waste-uptake and may help identify suitable target areas for similar studies on organotypic tissue cultures utilizing two-photon microscopy.

      The hypotheses and claims presented in this manuscript are not sufficiently substantiated and are conceptually unclear. The notion that amyloid beta and tau proteins play structural roles in a hypothesized "tanycytes"-derived canal network is not sufficiently supported by the evidence. Furthermore, the study lacks rigorous data to convincingly establish the proposed interactions between these proteins and the processes of waste internalization. In conclusion, due to conceptual and methodological issues, I consider the current evidence as inadequate to support the primary claims.

      We respectfully disagree with this comment and hope that the inclusion of additional evidence together with the clear visibility of this canal system in the spider brain will encourage the reviewer to investigate this possibility themselves. We cannot ignore large amounts of amyloid beta-immunolabeled receptacles emanating from tanysomes in swell-bodies and declare them fixation artifacts, particularly when we demonstrate the expression of Presenilin 1 and APP in swell bodies. We furthermore encourage the reviewer to revisit myelinated cells in the brain in both depictions in the available literature and actual brain preparations. We have not been able to locate actual electronmicrographs of longitudinal sections through neurons that show myelination past the axon hillock at the EM-level consistent with our current understanding of myelination. The only depiction of this form of myelination we found were schematic drawings. We will include several new images that show such longitudinal sections through neurons that are easily obtained and we have numerous additional images that we are happy to share. In all our preparations (mouse, rat and human) the myelination pattern is consistent with the images we will depict in new figures 1 and 2.  Not to bring attention to this inconsistency would be dishonest scientific conduct.

      Reviewer #2 (Public review):

      Summary:

      In this study, the authors propose the existence of an AQP4-positive tanycyte-associated canal system in the hippocampus and suggest that this system participates in waste clearance and contributes to Alzheimer's disease pathology. Using histological, ultrastructural, immunohistochemical, and RNA-based approaches, the manuscript attempts to reinterpret amyloid-β plaques and tau-associated structures as components of a tanycyte-derived waste-internalization system. The work is conceptually ambitious and raises observations that may stimulate discussion regarding glial organization and waste clearance in the diseased brain.

      Strengths:

      A strength of the manuscript is the combination of imaging modalities and anatomical observations across mouse and human tissue. Some of the reported morphological features are intriguing and may warrant additional investigation. The study also attempts to integrate structural observations with broader hypotheses regarding neurodegeneration and Alzheimer's disease.

      Weaknesses:

      The central interpretation depends almost entirely on identifying the observed hippocampal structures as tanycytes, and the evidence supporting this conclusion remains insufficient. Tanycytes are classically associated with ventricular regions in circumventricular organs, particularly in the third ventricle and median eminence region, yet the manuscript does not provide sufficiently specific anatomical or molecular evidence to convincingly distinguish the described structures from astrocytic, ependymal, radial glial-like, oligodendroglial, myelin-associated, vascular-associated, or degenerative elements. The marker profile used throughout the study, particularly the reliance on AQP4 labeling and Luxol-positive structures, is not sufficiently selective to establish tanycyte identity, especially in pathological tissue where reactive glial changes may occur.

      As mentioned in our response to reviewer 1 we have now included additional experimental evidence that demonstrates the myelinated ependymal cells and additional gene expression experiments. 

      This becomes particularly important because the manuscript repeatedly interprets Luxolpositive and myelin-associated structures as tanycytic processes or "myelin-derived tanycyte protrusions," despite tanycytes not being known to produce myelin. Alternative explanations are not sufficiently explored. Some of the canal-like structures shown in Figure 4 also resemble vascular profiles, and additional vessel markers would be necessary to exclude this possibility.

      Several of the proposed structures and mechanisms are also difficult to reconcile with established cell biology and neuroanatomy. The introduction of new terminology such as "tanysomes," "waste receptacles," and "toroids" further extends the interpretation beyond what is currently demonstrated experimentally.

      The discussion and integration of the existing literature on tanycytes are also insufficient. Tanycytes themselves are not clearly introduced; the manuscript does not adequately discuss what is currently established regarding tanycyte anatomy, ventricular localization, morphology, and function. Foundational literature defining tanycyte biology, including work from the Prévot group or others, is largely absent despite its central importance to the field. Because the manuscript proposes a substantial departure from established neurobiological concepts, it is particularly important that previous literature be discussed comprehensively and critically. The current version does not sufficiently contextualize the proposed model within the existing literature on tanycyte, AQP4, glymphatic, and Alzheimer's disease, making it difficult to evaluate what is genuinely novel versus what is merely being reinterpreted. It is also not entirely clear what is genuinely new here compared with the authors' previous work, particularly reference 11, which appears to present a highly similar conceptual framework.

      More broadly, several of the manuscript's mechanistic conclusions extend well beyond the available evidence. The proposal that amyloid-β plaques and tau pathology represent hypertrophic tanycyte-derived waste structures is provocative and potentially interesting, but currently remains largely correlative and speculative. At several points, it becomes difficult to distinguish direct observations from broader mechanistic interpretation. The manuscript itself acknowledges that the proposed glial-canal hypothesis contradicts the current understanding of nervous system organization and states that ultrastructural serialsection analysis would be required to unambiguously determine the origin of the myelinated profiles described. This point is critical because the study's central conclusions depend on the assumption that these structures are tanycyte-derived. At present, this interpretation remains insufficiently demonstrated, which substantially limits the strength of the broader pathological and mechanistic conclusions proposed throughout the manuscript.

      Although access to human material is understandably limited, the study appears to include only one male and one female AD patient, making it difficult to assess the reproducibility or frequent these structures are across individuals and pathological conditions. The manuscript would benefit from clearer characterization of prevalence, reproducibility, and variability across samples.

      Overall, the manuscript presents an unconventional and thought-provoking model that may stimulate discussion. However, the evidence currently provided does not convincingly establish tanycyte identity for the described hippocampal structures, and several of the broader disease-related interpretations would require substantially stronger anatomical and molecular evidence before the proposed model can be convincingly supported.

      We agree with the reviewer that it is important to correctly investigate and describe cellular structure. The first author of this manuscript is a 30-year veteran of published cellular ultrastructure and the three-dimensional reconstruction of cells and entire cell networks. We have spent the last five years to try and confirm our current understanding of cellular structure, and we are unable to reproduce our current identification of cell structure. One example is mentioned in response to reviewer 1. We are unable to find electonmicrographs in publications that show myelination of neurons consistent with the countless schematic depictions available online and in the literature. This includes publications about myelination. Our findings are all consistent with the new images we will include in our revised manuscript. Longitudinal sections through neurons are easily obtained and neurons can be followed well beyond the axon hillock.

      A second example that is inconsistent with our current understanding of cells and biochemical processes in cells are ‘astrocytes’ and ‘reactive astrocytes’. We will show in the revised manuscript swell-bodies have no defined cytoplasm that every cell requires to fulfil basic cellular functions required for survival. It is very apparent to a structural expert that swell-bodies lack cytoplasm. The ‘consistency’ of this lacking cytoplasm is ‘inconsistent’ with fixation artifacts. In Author response image 3 we demonstrate that immunolabeling for AQP4 shows two types of structures. (1) Immunolabeled structures that are void of immunolabeling in their lumina and show receptacle-like immunoreactivity along the outside (panels I,J,L,M image below) consistent with swellbodies as indicated by the provided amyloid beta-stained and Luxol H&E-stained swellbodies that are associated with immunolabeled receptacle-shaped structures on the outside (panels K,N). A structurally trained eye quickly recognizes that these are not immunolabeled cells but are consistent with swell-bodies and referred to as ‘reactive astrocytes’ in the literature. We show in panel O what an immunolabeled cell looks like, with slender processes and immunolabeled cytoplasm. In this image in panels G1-4 we demonstrate that swell-bodies contain AQP4 mRNA explaining why they are immunoreactive for this protein. It is important that we bring attention to these details and do not randomly describe structure just based on a signal. This is exactly the point we make and I hope that the reviewer recognizes our expertise in cellular neuroscience and in particular in recognizing cellular structure. 

      Again, I urge the scientific community not to dismiss our findings but to actually study the validity of our findings. 

      Author response image 3.

      In the revised manuscript we will include additional experiments that we have carried out, that have also increased the sample number of tested human brain. We have in total so far investigated 13 different human brain samples, six of which are AD-affected. 

      We will revise our discussion to explain our observations in context with the literature better and include so much evidence in support of our hypothesis that it would be unreasonable to dismiss all this compelling and logical evidence.

      This research was not planned; it resulted from our accidental discovery in the spider system when our animals struggled with early onset neurodegeneration that we needed to address. This is when we recognized the waste-internalizing role of myelin in giant spider neurons. As we will discuss in the revised manuscript, such systems are highly conserved throughout evolution, and this is what made us realize that the only images of myelinated neurons we could find were either schematic drawings, single cross sections through myelinated cell profiles or very high magnification insets that also did not show the actual neuron that is myelinated.

      One can argue that there are both myelinated and unmyelinated axons. However, the varicose projections clearly originate in the ependymal lining and double-label for AQP4. This is consistent with our postulation and inconsistent with our current understanding of myelination. 

      I sincerely urge the neuroscience community to re-visit myelination in the brain, we have done this for the past five years with extensive experience in this field and the only hypothesis that is supported by our findings is presented in this manuscript. Please do not dismiss these findings, the spiders have uncovered a waste canal system in the brain and putting this system in place will help us to gain a better understanding regarding neurodegenerative diseases. Lastly, and maybe most importantly, understanding how this system works in spiders and how the myelin is structurally anchored to microtubule that were missing in our degenerating spiders has allowed us to identify the cause for this sudden neurodegeneration and rescue our tropical, cold-blooded spider colony by installing a new heating system and raising the room temperature so that the coldsensitive microtubules no not dissociate anymore.  

      We would like to thank the reviewers to strengthen the content of this manuscript with their critical comments, we hope that our revision will help clarify some doubts.

      (1) Pasquettaz R, Kolotuev I, Rohrbach A, Gouelle C, Pellerin L, Langlet F. Peculiar protrusions along tanycyte processes face diverse neural and nonneural cell types in the hypothalamic parenchyma. Journal of Comparative Neurology. 2021;529(3):553. doi: 10.1002/cne.24965. PubMed PMID: edsgcl.646848213.

    1. eLife Assessment

      This work describes a valuable method for monitoring pathogens using an innovative, field-deployable device that attracts animals and collects saliva on filter paper in a non-invasive manner. While the strategy is effective at collecting samples and ensuring their preservation, support for some claims remains incomplete. This study will interest scientists in various fields, ranging from biodiversity, ecology, and conservation to infectious diseases and public health.

    2. Reviewer #1 (Public review):

      Summary:

      This manuscript describes the development and validation of a low-cost device to identify viruses from saliva samples of animals non-invasively. This device was tested under laboratory conditions to assess whether viruses could be recovered in different environmental conditions and after different durations of time. The devices were then used to sample mice and cats in shelters to assess utility.

      Strengths:

      Sampling animals is cost-effective and highly labour-intensive, and this device has the potential to substantially improve surveillance. The device is relatively low-cost, and the authors demonstrate that the virus can be obtained from these filter papers after different durations of time and in different environmental conditions.

      Weaknesses:

      The authors do not discuss if different volumes were obtained from different animals (for example, due to different behaviours or attractiveness of the odour baits). Additionally, it appears the virus results were cross-validated using the serological status of the animals. While I am not an expert on FIV, there seems that there could be potential for different levels of viral shedding, and it would be more prudent to cross-validate against blood or another gold standard sample. Finally, the statistical analysis could be improved as there appear to be relatively few replicates and limited analysis conducted.

    3. Reviewer #2 (Public review):

      Summary:

      The study introduces an innovative device designed to collect non-invasive saliva samples from animals using disposable cassettes with odor attractants and filter paper. The authors aimed to validate this tool for pathogen monitoring, specifically by detecting pathogen RNA in animal models. While the concept is compelling and the problem statement well-framed, the validation of the device for pathogen detection was not achieved. For example, the rabies virus was not detected in the chosen model, and results were limited primarily to FeLV. The work highlights the potential of saliva-based sampling for microbiota analysis, but the rationale for virus selection and the experimental design require further clarification. Overall, the study presents a novel approach with promise, though its current scope is better suited to microbiota monitoring rather than pathogen surveillance.

      Strengths:

      The innovative design of the device, which enables non-invasive saliva collection through disposable cassettes with odor attractants, represents a creative and practical advance in sampling methodology. The authors undertook an extensive experimental effort, generating a substantial amount of data that highlights the feasibility of saliva-based monitoring. The rationale for exploring saliva as a medium is valid, and the work successfully shows that the device can be applied to microbiota profiling, where the strongest results were obtained. This methodological innovation could be valuable for expanding non-invasive approaches to animal health monitoring.

      The authors acknowledge that metabarcoding sequencing has limitations; however, the study could be refocused on the microbiota in general rather than on pathogen detection. They could give greater prominence to the taxonomic composition of microorganisms in saliva using high-throughput sequencing. That is where they obtained the most results.

      Weaknesses:

      Despite the enormous experimental effort undertaken, the results fall short of the expected success of the proposed test. The rationale and criteria for virus selection are not clearly explained, leaving the experimental design insufficiently justified.

      The central aim of validating the device for pathogen detection was not achieved, particularly in the case of the rabies virus. The mouse infection model used for the rabies virus does not seem to adequately replicate the natural course of the disease. This could explain, at least in part, the negative results obtained.

      Of the three viruses evaluated, satisfactory results were obtained only for FeLV, and the sample size remains limited. According to the literature reviewed, this virus is not common in wild cats, so the applicability of the results would appear to be limited primarily to domestic cats.

      The collected samples were stored at −80 {degree sign}C for later analysis, which likely contributed to the high Ct values observed with the device. The need to store samples at low temperatures may be a limitation to applying this technique in wildlife sampling scenarios where access to dry ice or liquid nitrogen tanks may be difficult.

      Stating that the device can be used for pathogen monitoring in wild animals is not desirable, since the viruses for which results were obtained are not relevant in wild animals. On the other hand, claiming that this is a tool for monitoring diseases in endangered species is also misleading. Endangered species are typically scarce and therefore would not be the reservoirs that these surveillance efforts should target. In fact, groups such as wild rodents would be a better target for monitoring zoonotic pathogens.

    1. eLife Assessment

      This study provides an important and biologically plausible account of how human perceptual judgments of heading direction are influenced by a specific pattern of motion in optic flow fields known as retinal curl. By combining psychophysical experiments and neural modeling, the authors demonstrate that what was previously considered an incidental, 'nuisance' signal actually serves as a functional control signal for estimating heading and steering toward a fixated target. The evidence supporting the role of curl signals is convincing and advances our understanding of vision-based navigation at the behavioral level. A particular strength of the work is the direct manipulation of curl within flow fields, demonstrating that it effectively cancels or reverses heading biases. This work provides an invaluable framework for future exploration into the neural mechanisms of steering control based on retinal curl.

    2. Reviewer #1 (Public review):

      Summary:

      This carefully executed study uncovers the functional relevance of curl signals that impinge on the retina every time an observer's gaze direction and movement direction are not aligned. This finding is important, highlighting the functional role of an abundant incidental signal (curl in retinal motion) that has thus far believed to be a nuisance that needs to be filtered out of the retinal motion stream. As such, the study forms an important contribution to the emerging recognition that incidental sensory signals are not a challenge to the sensorimotor system, but contain functionally relevant and effectively used visual signals. The study's evidence is compelling: A combination of psychophysical experiments and critical manipulations, control theory and neural modeling makes an internally consistent and biologically plausible case for the role of curl signals in estimating heading direction. The experimental and modeling results clearly go beyond previous studies and significantly advance our understanding of vision-based navigation.

      Strengths:

      The study has its strengths in the combination of psychophysical experiments and critical manipulations, control theory and neural modeling, which together make an internally consistent and biologically plausible case for the role of curl signals in estimating heading direction.

      This study uncovers the functional relevance of curl signals that occur on the retina when an observer is moving and gaze is not straight ahead. The experimental and modeling results clearly go beyond previous studies and significantly advance our understanding of vision-based navigation.

      Another clear strength is that the study uses tightly controlled experimental manipulation to provide strong test cases for the hypothesis that curl is used for visual navigation. These conditions are important to constrain the proposed model (and future models) of heading control.

      The modeling is very clearly described and the modeling and analysis code is published and freely available. The authors go beyond a back-of-the-envelope control model and show how it might be implemented at the neural-circuit level. The model is biologically plausible.

      Weaknesses:

      I see no major weaknesses of the study. I expect it to inspire future research that extends these findings to a wider range of visual environments (including walking in natural scenes), motion speeds and kinds of movements.

      Comments on revised version.

      I have no additional comments for the authors.

    3. Reviewer #2 (Public review):

      This study examines how curl in the retinal flow field can be used as a control variable for estimating and controlling the heading of a moving observer. The basic idea (which is not entirely new, see Matthis et al. 2022) is that translation along a path with eccentric gaze (meaning that the subject is not heading toward the point they are looking at) produces a pattern of optic flow on the retina with a rotational component around the point of fixation (which can be captured by the mathematical "curl" operator). The sign and magnitude of retinal curl varies with heading relative to the point of fixation, such that curl can be used as a control variable to steer rightward or leftward to move toward the fixated target. The authors perform behavioral experiments and show that there are biases in perceived heading that seem to be largely governed by retinal curl. They also show that a simple controller model can use curl to steer toward a target, and they provide a neural network model that provides a biologically-plausible implementation of the controller (although there are some questions about that).

      There is a core of interesting work here that I think can be important to the field. However, there is a lack of clarity on several important fronts, including design of the behavioral experiments, presentation of the behavioral data, conceptual framing of what curl can and cannot do, etc. Equally importantly, the manuscript is not written in a manner that will make it accessible to most vision scientists. I consider myself to be pretty knowledgeable about optic flow, and I had to read most of the manuscript 3 or 4 times to be able to understand the bulk of it. And my experience is that most vision scientists do not understand optic flow well, so I fear that most of the readers that the authors should want to reach would struggle to understand the work. As written, this is mainly going to make an impact on a handful of optic flow gurus. Thus, this manuscript is going to need a major overhaul to clarify important issues and make this more accessible.

      Major issues:

      (1) The manuscript contains inconsistent, if not misleading, messaging about what information retinal curl does, and does not, provide regarding heading estimation. In the Abstract, the authors state: "We propose an alternative: the visual system utilizes retinal curl directly to estimate heading, rendering the explicit recovery of the FOE unnecessary." Based on my understanding of the rest of the manuscript, I find this statement to be a misrepresentation for two main reasons:<br /> a. To "directly estimate heading" relative to what? When not qualified, most people interpret "heading" to mean an observer's heading relative to the world (or some allocentric reference frame). But retinal curl only gives information about an observer's heading relative to the point on which their eyes are fixated. Moreover, that point of fixation will change every few hundred milliseconds in natural viewing, so the retinal curl will change with each new fixation even as heading relative to the world remains unchanged. So, I think most readers would grossly misinterpret the claim that retinal curl can be used "directly to estimate heading". Indeed, in the authors' controller model, the initial heading needs to be given and then the controller can work. But from where does the visual system get the initial heading, since it does not come from curl? These issues are left hanging. Thus, while curl can provide a very useful input for steering toward a fixated target, other signals are needed to estimate heading relative to the world. This has to be made much clearer early on, and a conceptual schematic diagram might help. Also, the authors generally do not specify the reference frame of the variables they are talking about, leaving lots of room for misinterpretations. It should be clear each time they are talking about a variable, such as heading, whether it is relative to the fixation target, body, world, etc.<br /> b. It seems to me that retinal curl will depend on other variables, in addition to heading relative to the fixation target. For example, it seems to me that the magnitude of retinal curl will depend on self-motion speed, the depth structure of the scene, the angle of elevation of the fixated target, and perhaps others. This is not discussed at all, and many readers would get the misguided impression that there is a 1:1 mapping from curl to heading (relative to fixation). If I am right that this is not correct, it means that retinal curl can tell the observer whether to steer right or left to move toward the fixated target, but it cannot tell them how much to steer. Indeed, in the authors' controller model, there is a free parameter that calibrates curl to angle. It makes sense that this works to fit trajectory data that are given from a fixed environment, but it is unclear how the brain would use retinal curl to control steering when these other variables are uncertain or changing unpredictably. Moreover, how does the system change the mapping from curl to steering command as the location of fixation changes relative to the current heading? These are issues that need to be brought up in framing the problem and discussed at some length. If the authors can show mathematically that retinal curl is only dependent on heading (relative to fixation) and not any of these other variables, it would be very valuable to show the equations for this relationship.

      (2) The description of the behavioral experiment and presentation of behavioral data leaves a lot to be desired.<br /> a. First, it is stated (line 158) that "Participants continuously reported their perceived direction of self-motion while maintaining fixation on the yellow dot." Again, reference frame is completely unspecified. Participants were reporting their perceived heading relative to what? The fixation target? The world? What exactly were the instructions given to the subjects to perform the task? Based on the description of how perceived paths are computed (line 166-), it seems to be presumed that subjects are reporting their heading relative to the world because those angles are then converted into x and z coordinates in what I presume is a world-centered reference frame. But how do we know that subjects are accurately reporting their heading relative to the world? What if they are biased in their reports by the location of the fixation target relative to the scene, or by some other reference signal? Is it possible for the authors to rule out the possibility that perceptual biases seen in the unaltered curl condition result from observers not fully adopting the assumed reference frame of the task? If this cannot be firmly excluded, it seems to create problems for the rest of the study.<br /> b. I also feel that there is a mismatch between what the behavioral task requires and what the controller model does. Subjects are apparently asked to report their heading relative to the world, but the controller model only controls their heading relative to the point that they are fixating. I understand how this is resolved in the model, but I think this type of distinction is buried and will not be apparent to most readers. Again, the reference frames of what is being measured and controlled need to be specified explicitly in all parts of the paper, and the authors needs to explain how the system would combine curl-based control with some other measures of (at least initial) heading for world-centered heading to be computed. All of the assumptions need to be clearly specified.<br /> c. Second, I found it frustrating that the authors never present raw perceptual data from the observers. Rather, in Figure 2, we see reconstructed trajectories that are perfectly smooth with no indications of noise whatsoever. Since these paths are computed from the perceptual reports, there must be some noise inherent in them. The figures should represent this uncertainty somehow, and it should be explained how these perfectly smooth trajectories are obtained.

      (3) "...the magnitude of retinal curl in the fovea can specify the body trajectory relative to gaze (Matthis et al., 2022)." The main idea put forward by the authors here seems to overlap heavily with this statement that they attribute to Matthis et al. 2022. While I think this paper still adds importantly to the topic, the authors do not discuss how their findings are different from those of Matthis et al. 2022, why they are an important extension, etc. Readers should not have to go read this other paper to have any idea how the present findings are placed in importance relative to the literature.

      (4) The analysis and treatment of eye movements is extremely weak. The authors discarded trials for which gaze deviated from the fixation point by more than 3 degrees (which is a LOT given that the eye speeds are generally in the neighborhood of 0.5 deg/sec), and they provide basic stats on the distribution of positions. But this largely misses the point: it is not small position errors that are likely to matter, but rather velocity errors. Even a small amount of retinal slip of the target while it is being pursued will cause image motion that is going to alter the optic flow field around the fixation target. So, for example, the retinal curl field may no longer be centered on the fixation target. How do we know that some of the perceptual biases are not influenced by image motion resulting from imperfect tracking of the fixation target? This needs to be analyzed and discussed.

      (5) I found the sections of text comparing the separate and joined fits (starting line 287) to be a bit too rosy. The authors show the separate fits in the main text, and it is not very surprising that these fits are good given that the model has 30 parameters, and these data are pretty low dimensional. The authors only show the joined fits in the supplement, and they say that they are almost as good as the separate fits (indeed they are better in a model comparison sense, but this is 30 parameters vs. 2 parameters). However, when I look at the fits of the joined model in the supplement, I don't find them to be very impressive. In particular, the model grossly misses the data for the straight paths for several subjects (e.g., id5, id6, id8, id10). And fitting the straight paths would presumably be easiest. This implies that the joined model is really missing something and that fitting the curved paths interacts strongly with fitting the data for different fixation target locations on the straight path. I think that the authors should discuss the results a bit more soberly and tone down their conclusions here.

      (6) The section of the paper on neural simulations (starting line 387) has a few weaknesses. First, why are only straight paths simulated here? This does not seem to provide a very rigorous test of the model. Second, it is awkward that the simulation results are presented in units of pixels, rather than degrees. Third, the authors seem to downplay the fact that the neural estimates of heading seem to oscillate rather wildly (over a range of hundreds of pixels, whatever that means, see especially Fig. S16). It was far from clear to me how an estimate of heading with these large oscillations is useful. It would seem to require that heading estimates are integrated over substantial lengths of time to be reliable. It was therefore unclear how the model produces such smooth paths from these oscillating estimates.

      Comments on revised version.

      Overall, the authors have done a responsible job of responding to the comments of my previous review, and the manuscript is substantially improved. There are a few points on which I still do not completely agree with the authors, and I think these are important to document for the record:

      (1) Introduction: "Pure visual decomposition should function regardless of 3D depth or whether the rotation stems from an active eccentric fixation." Perhaps in a world of noiseless perfect computation, this might be true. But I generally disagree. When there is more depth structure in an environment, then translation of the observer is generally going to create greater motion parallax. That is a fact that I don't think can be disputed. And greater motion parallax should help to decompose optic flow into components related to translation and rotation (the latter of which is not depth dependent), especially when there is noise in estimating location motion vectors.

      (2) Related to point #9 of my previous review: I had asked why the authors believed that retinal curl was computed in area MSTd. Their response is that previous studies (i.e., Graziano et al. 1994) show selectivity to spiral motion stimuli in MSTd. That is true, but those studies typically placed the spiral stimulus centered on the MSTd receptive field, hence they were not presenting something like retinal curl as defined here. So, I think it is still an open question as to where in the brain retinal curl is encoded, and from which areas it would be possible to decode retinal curl from population responses.

      (3) Related to point #10 of my previous review: I had asked about biological plausibility of the gaze-centered inhibition signal in the model. The authors' response is that parietal neurons show gain fields in which response depends (usually monotonically) on eye position. This is true, but it is not a trivial jump from gain fields in individual neural responses to a gaze-centered inhibition signal, and I think the authors should have been more forthcoming about the lack of an established neural signal that directly signals what they want in their model.

      (4) The authors point out that the perceptual biases they measure take a few seconds to emerge and they attribute this to temporal integration. But in their curl manipulations, they temporally average over a 2.4 second window in computing the curl signals that they use to cancel or over-cancel curl. So, it is not clear whether some of the delay in the behavioral effects might result from their computations.

      (5) Related to point #13 of my previous review: I had asked about empirical evidence for the assumption of a relationship between the heading preferences of MSTd neurons and their receptive field locations. In response, the authors state that such a relationship is built into the Layton and Browning (2014) model. While that is a precedent, citing another model as a response to a question about empirical evidence is not a convincing response. If there is no empirical evidence to support such a relationship, it would have been better for the authors to acknowledge this.<br /> Given the way that the eLife review model works, it is not necessary for the authors to address these comments, but I think they should be included in the public review record.

    4. Reviewer #3 (Public review):

      Major strengths include the use of realistic retinal motion recorded during virtual walking, an elegant manipulation of curl, converging behavioral and modeling evidence, and grounding in control theory. This provides a novel and important contribution to our understanding of how the brain processes motion information and intuition about how that information might be used to guide steering. In addition, they provide a computational mechanism by which retinal flow curl can be used as a control signal.

      The revised ms has been strengthened by more explicit discussion of the literature where there has been mixed evidence for the use of the Focus of Expansion. Since the ms is a strong test of the use of curl as a heading signal, this allows a deeper understanding of the importance of the finding and historical context. The ms has also been strengthened by a more explicit discussion of integration of the time-varying signal over periods of several seconds, which is an important demonstration. The implications of the ms are still a little unclear, as the results involve visual judgements in seated subjects. The use of different sources of information when humans walk from one place to another in real life may be complex and involve a variety of different sources of information.

    5. Author response:

      The following is the authors’ response to the current reviews.

      We thank the editors for their positive assessment of our manuscript, and all the referees for their constructive comments, which have substantially improved this work. We welcome the opportunity to address referee #2's points for the public record, as they highlight key theoretical nuances and valuable future research directions.

      (1) We appreciate the reviewer's point that, in a noisy biological system, the increased motion parallax provided by a rich 3D depth structure naturally aids in separating translation from rotation. We fully agree on this point. Our argument aimed at highlighting a fundamental theoretical distinction. Pure algebraic decomposition algorithms are mathematically capable of solving heading on flat planes. The fact that human perception often shows biases in these zero-depth conditions, unless extra-retinal cues are present, suggests that the visual system does not rely on a generalized, global de-rotation algorithm. Instead, it relies on heuristic, depth-dependent structural signals (like motion parallax and retinal curl). We maintain that while depth certainly reduces noise, its strict necessity points toward an ecologically grounded control strategy rather than a noisy global decomposition process.

      (2) We think the reviewer raises a valid point regarding the exact neural locus of retinal curl encoding. It is true that Graziano et al. (1994) utilized centred spiral stimuli rather than the spatially offset curl geometries defined in our task. We view the spiral tuning of MSTd not as a direct, one-to-one mapping of full-field retinal curl, but rather as the foundational computational building block required to extract such a signal. We fully agree with the reviewer that identifying exactly where and how this population response is decoded into a unified, gaze-relative retinal curl signal remains an exciting and open empirical question for future neurophysiological research.

      (3) We acknowledge the reviewer's call for transparency here. The transition from well-documented multiplicative gain fields (which modulate response amplitude based on eye position) to a direct, localized gaze-centered inhibitory drive is indeed a theoretical abstraction in our model. We utilized this localized inhibition as a functional mechanism to demonstrate how sensory evidence and spatial priors might competitively interact within a standard Mexican-hat recurrent architecture. While gain fields clearly establish that parietal networks integrate gaze position, we agree that the exact local-circuit implementation mapping these gain fields to the specific inhibitory dynamics we modeled has yet to be empirically established.

      (4) The reviewer smartly questions whether the 3-5 second behavioral biases delay emerges from the 2.4-second smoothing window used in our computational flow manipulation. It is important to clarify that this 2.4-second window was used solely to stabilize the computed curl signal against high-frequency gait oscillations. This smoothing was restricted strictly to the modeling phase of the controller and neural model and was not applied to the participants' responses. The gradual build-up of their perceptual bias over 3-5 seconds represents their own intrinsic temporal integration of this trajectory, independent of the smoothing parameters used to smooth the curl used in the fitting of the controller and neural modelling.

      (5) We concede the reviewer's point that citing a computational model (Layton & Browning, 2014) does not constitute direct empirical evidence for a relationship between MSTd heading preferences and their receptive field locations. Our intention was to highlight a successful theoretical framework that elegantly organizes known properties of MSTd into a system capable of bypassing global de-rotation. We readily acknowledge that direct, single-cell empirical validation of this specific topographic relationship is currently lacking in the literature, and we appreciate the reviewer ensuring this distinction is clearly noted for the record.


      The following is the authors’ response to the original reviews.

      eLife Assessment

      This study provides an important and biologically plausible account of how human perceptual judgments of heading direction are influenced by a specific pattern of motion in optic flow fields known as retinal curl. By combining psychophysical experiments and neural modeling, the authors demonstrate that what was previously considered an incidental "nuisance" signal actually serves as a functional control signal for estimating heading and steering toward a fixated target. While the evidence for the role of curl signals is convincing and advances our understanding of vision-based navigation, the work's impact would be strengthened by situating these findings among other cues that contribute to heading estimation, and by clarifying both the time course of these computations and their generalizability across different navigational contexts.

      We thank the editors and reviewers for their insightful feedback and positive assessment of our study. In this revised version, we have made substantial modifications to better situate our findings within the broader landscape of cues contributing to heading estimation, while also clarifying the time course of these computations and their generalizability across different navigational contexts. These points are included in new sections in the revised discussion.

      In addition, and following eLife guidelines, we have moved the methods to the end and make sure that the manuscript reads well without needing to go through methods first.

      Next, we address all the concerns raised by the reviewers.

      Reviewer #1 (Public review):

      We appreciate Reviewer #1’s very positive feedback. Incorporating the perspective of ‘incidental’ sensory signals is a valuable suggestion that aligns perfectly with our findings. We agree that this perspective significantly strengthens the impact of our paper.

      In the revised version we have added a last section in the Discussion (Generalizability and Testable Predictions) to comment on the functional utility of 'incidental' signals and incorporated the suggested references. In addition, in the same heading, we briefly elaborate on the predictions and generalizability of the model and possible manipulations that might affect the integration between sensory evidence (curl signal) and straight-ahead prior.

      Reviewer #1 (Recommendations for the authors):

      It would be great if the authors could discuss the implications and predictions of their model.

      First, from a broader perspective, the study forms an important piece in the emerging recognition that incidental sensory signals are not a nuisance to the sensorimotor system, but contain functionally relevant and effectively used visual signals (Rolfs & Schweitzer, 2022). The authors may want to appreciate their contribution to this perspective in the discussion of the impact of their results. Indeed, a similar shift in perspective has been realized in the recognition that saccade-induced motion signals are not entirely suppressed but play a functional role for gaze correction (Schweitzer & Rolfs, 2021).

      Second, the authors could spell out additional predictions of their proposal: What are experimental manipulations that could shift the balance between relying on a straight-ahead prior and sensory estimation of curl? What would happen in extreme cases of such sensory evidence? When would priors become overwhelmingly influential?

      As commented in the public review we have now included these two aspects in the discussion.

      Minor point: After equation 11, the authors may want to specify that, like position, gaze g is also coded as {x,y}, just like position p.

      While the neural model details have been moved to Appendix 2 (including this equation), we added text before Equation 22 (previously Eq. 11) clarifying that gaze is encoded in image coordinates. We omitted point index i because the equation applies to all image points relative to a given gaze g.

      Reviewer #2 (Public review):

      We appreciate the reviewer’s feedback regarding the formalization of our reference frames. We agree that certain definitions were implicitly assumed rather than explicitly stated. We have revised the manuscript to provide all necessary self-contained information, ensuring that the geometry of the task response and the definition of heading are unambiguous. In the last paragraph of the revised introduction, we make clear the response frame of reference which is also included in the caption of fig 1. Also, we have addressed the gap between the task response (in world coordinates) and the functional role of the controller. This is particularly discussed in the discussion (section: reference frames) in which we also provide (and rule out) potential alternatives to our response biases. We also address all the other points raised by the reviewer.

      Major issues:

      (1) The manuscript contains inconsistent, if not misleading, messaging about what information retinal curl does, and does not, provide regarding heading estimation. In the Abstract, the authors state: "We propose an alternative: the visual system utilizes retinal curl directly to estimate heading, rendering the explicit recovery of the FOE unnecessary." Based on my understanding of the rest of the manuscript, I find this statement to be a misrepresentation for two main reasons:

      (a) To "directly estimate heading" relative to what? When not qualified, most people interpret "heading" to mean an observer's heading relative to the world (or some allocentric reference frame). But retinal curl only gives information about an observer's heading relative to the point on which their eyes are fixated. Moreover, that point of fixation will change every few hundred milliseconds in natural viewing, so the retinal curl will change with each new fixation even as heading relative to the world remains unchanged. So I think most readers would grossly misinterpret the claim that retinal curl can be used "directly to estimate heading". Indeed, in the authors' controller model, the initial heading needs to be given, and then the controller can work. But from where does the visual system get the initial heading, since it does not come from curl? These issues are left hanging. Thus, while curl can provide a very useful input for steering toward a fixated target, other signals are needed to estimate heading relative to the world. This has to be made much clearer early on, and a conceptual schematic diagram might help. Also, the authors generally do not specify the reference frame of the variables they are talking about, leaving lots of room for misinterpretations. It should be clear each time they are talking about a variable, such as heading, whether it is relative to the fixation target, body, world, etc.

      In our study, participants were instructed to report their “perceived direction of self-motion” by aligning a rotational encoder (steering wheel) with the direction they felt they were moving within the 3D simulated scene. Consequently, participants reported their instantaneous heading in a world-centered reference frame, from which the 3D trajectories were reconstructed. Since the reviewer had to infer this information, we have clarified this point at the end of the introduction, legend of figure 1, methods (now at the end of the ms.) and discussion to ensure it is immediately evident.

      Participants were informed that the initial heading (i.e. θ<sub>0</sub> in our controller nomenclature) was oriented “straight ahead” relative to their body which was aligned longitudinally with the experimental room. We have modified Figure 1B and revised the Methods section to explicitly clarify this initial alignment and the instructions provided to participants.

      In the revised manuscript, we have clarified that while the participant’s report is world-centered, the retinal curl provides a gaze-relative heading signal. Although this was already mentioned, we emphasize this point. In natural navigation toward a fixated target, a world-centered vector is often unnecessary; an error signal indicating heading relative to fixation is sufficient (as the reviewer also notes). However, the initial alignment of the heading within the 3D scene allows the brain to “calibrate” this internal controller, mapping the retinal curl signal onto the 3D world coordinates required for the task. Ad commented above, a new section in the discussion addresses and hopefully clarifies the relation with the controller.

      The reviewer also asks how we can be certain that participants were reporting in world coordinates rather than an alternative frame, such as “heading relative to the fixation target.” We believe our “Cancelled Curl” (and over-cancelled) conditions provide the most compelling evidence to rule out this alternative. In these conditions, the physical position of the fixation target in the scene remained identical to the unaltered flow condition. If participants were simply reporting heading relative to the fixation target’s spatial location, the observed biases should have persisted regardless of the flow manipulation. Instead, the bias vanished when the curl was removed. This causal evidence proves that the bias is driven by the retinal motion signal (curl) rather than the spatial orientation of the eyes or the target’s position in the scene. Furthermore, the temporal evolution of the response supports a world-centered integration (in agreement with Warren et 2001 Nat Neuro.). For simulated straight paths, the perceived heading remains straight for the first few seconds (consistent with the initial world-centred alignment), with biases only emerging after approximately 3 seconds of integration (a point we elaborate on in our response to Reviewer #3). Had participants been responding based on a simple gaze-relative reference frame from the onset, these biases would have manifested significantly earlier. We have incorporated these points into the revised Discussion to better frame our findings alongside other cues, such as the Focus of Expansion (FOE) and egocentric visual direction that contribute to heading estimation.

      Finally, we have rephrased the abstract sentence for clarity. However, we maintain that the original premise remains valid once the world-centered initial heading is aligned with the gaze-centered reference frame.

      (b) It seems to me that retinal curl will depend on other variables, in addition to heading relative to the fixation target. For example, it seems to me that the magnitude of retinal curl will depend on self-motion speed, the depth structure of the scene, the angle of elevation of the fixated target, and perhaps others. This is not discussed at all, and many readers would get the misguided impression that there is a 1:1 mapping from curl to heading (relative to fixation). If I am right that this is not correct, it means that retinal curl can tell the observer whether to steer right or left to move toward the fixated target, but it cannot tell them how much to steer. Indeed, in the authors' controller model, there is a free parameter that calibrates curl to angle. It makes sense that this works to fit trajectory data that are given from a fixed environment, but it is unclear how the brain would use retinal curl to control steering when these other variables are uncertain or changing unpredictably. Moreover, how does the system change the mapping from curl to steering command as the location of fixation changes relative to the current heading? These are issues that need to be brought up in framing the problem and discussed at some length. If the authors can show mathematically that retinal curl is only dependent on heading (relative to fixation) and not any of these other variables, it would be very valuable to show the equations for this relationship.

      The reviewer notes that we must be clear about the relationship between curl and heading (relative to fixation) and the variables that affect curl. We also thank the reviewer for encouraging to add the equations that show the relation of curl with additional variables. We have now included these equations in appendix 1.

      Beyond the discrepancy between heading (θ) and gaze (ψ), curl is geometrically determined by translational self-motion speed (v), eye height (h), and pitch (α). More specifically, curl = (v.sinψcosα)/h. The derivation is now included in appendix 1. Since h = dsinα, where d is the 3D distance to the fixation point, we could express cos α as a function of distance. Certainly, there is not a 1:1 map from curl signal to heading relative to gaze (e.g. θ-ψ). Participant would need to know v and eye height plus extra-retinal information. Frenz et al (2003, Vis Res.) showed that people can estimate self-motion directly from optic flow, across different simulated eye height and gaze angle; extra-retinal information can, in addition, provide knowledge to ψ and α. It is then plausible that the visual system can use and transform the curl signal from a qualitative directional cue (i.e. steering left or right of fixation) into a quantitative steering command. By combining curl with knowledge of gaze orientation and eye height, the visual system can resolve ambiguities in the flow field and utilize curl as a more precise error signal for locomotor control. These aspects are now included in the new version of the discussion.

      (2B) I also feel that there is a mismatch between what the behavioral task requires and what the controller model does. Subjects are apparently asked to report their heading relative to the world, but the controller model only controls their heading relative to the point that they are fixating. I understand how this is resolved in the model, but I think this type of distinction is buried and will not be apparent to most readers. Again, the reference frames of what is being measured and controlled need to be specified explicitly in all parts of the paper, and the authors need to explain how the system would combine curl-based control with some other measures of (at least initial) heading for world-centered heading to be computed. All of the assumptions need to be clearly specified.

      We thank the reviewer for this point. We have addressed the alignment of the reference frames in our response to Issues 1a and 2a. Once the initial orientation (θ<sub>0</sub>) is established in the world frame, the controller model generates steering adjustments that directly translate into heading predictions within that same world reference frame. By treating the perceptual report as an output of the locomotor controller, we resolve the discrepancy between the steering task and the reported heading.

      (2c) In addition, I found it frustrating that the authors never present raw perceptual data from the observers. Rather, in Figure 2, we see reconstructed trajectories that are perfectly smooth with no indications of noise whatsoever. Since these paths are computed from the perceptual reports, there must be some noise inherent in them. The figures should represent this uncertainty somehow, and it should be explained how these perfectly smooth trajectories are obtained.

      We respectfully disagree with the reviewer’s interpretation regarding data smoothing. The thin lines in Figure 2 represent the mean 3D paths derived directly from the response variable (θ<sub>t</sub>) across trials of identical conditions for each participant (as detailed in the ‘Computation of Perceived Path’ section). No smoothing or filtering has been applied to these plotted trajectories other than computing the mean across trials. We also wish to remind the reviewer that the raw data and analysis code remain publicly accessible for further inspection. Having said that, we include now a supplementary figure showing an example of raw data responses as a function of time. This figure will be a supplemental figure of main Figure 2 (now provisionally included in the Suppl Information).

      Regarding the visual representation: in earlier versions of the manuscript, we included shaded 95% Confidence Intervals (CIs) in Figure 2. However, this addition rendered the plot overly cluttered and obscured the individual trajectories. We therefore chose to present individual participant means (thin lines) alongside group averages (thick lines) to emphasize inter-subject variability. For clarity, the 95% CIs are explicitly displayed in Figure 3, where the data density is more conducive to shaded areas.

      (3) “...the magnitude of retinal curl in the fovea can specify the body trajectory relative to gaze (Matthis et al., 2022)." The main idea put forward by the authors here seems to overlap heavily with this statement that they attribute to Matthis et al. 2022. While I think this paper still adds importantly to the topic, the authors do not discuss how their findings are different from those of Matthis et al. 2022, why they are an important extension, etc. Readers should not have to go read this other paper to have any idea how the present findings are placed in importance relative to the literature.

      We have updated the Discussion to more specifically align our findings with Matthis et al. (2022). We emphasize that our study provides the perceptual validation for their ecological observation that the FOE is often too unstable for reliable use, whereas foveal curl remains a robust signal for path estimation. Our paper provides the causal link, since we manipulate curl in real-time (the ‘cancelled & over cancelled curl’ condition) providing the critical evidence that perceived heading is affected by this signal. The relation with this previous study is made clear in the revised discussion.

      (4) The analysis and treatment of eye movements is extremely weak. The authors discarded trials for which gaze deviated from the fixation point by more than 3 degrees (which is a LOT given that the eye speeds are generally in the neighborhood of 0.5 deg/sec), and they provide basic stats on the distribution of positions. But this largely misses the point: it is not small position errors that are likely to matter, but rather velocity errors. Even a small amount of retinal slip of the target while it is being pursued will cause image motion that is going to alter the optic flow field around the fixation target. So, for example, the retinal curl field may no longer be centered on the fixation target. How do we know that some of the perceptual biases are not influenced by image motion resulting from imperfect tracking of the fixation target? This needs to be analyzed and discussed.

      We thank the reviewer for noting that retinal slip (velocity error) is a more critical metric than positional gaze error. We agree that tracking inaccuracies can introduce translational noise into the flow field. The 3° threshold was established based on the eye tracker’s specifications and the naturalistic setup (1-meter viewing distance without head stabilization). Across all participants, the mean positional error ranged from 1.016° to 1.5° (1 deg is 2.08 cm in our setup). We also calculated retinal slip values, which ranged from 0.12 to 0.27 deg/s (X dimension) and 0.12 to 0.23 deg/s (Y dimension). These values are comparable to natural oculomotor drift (Kowler et al., 1979) and are understandably small given the low velocity of the fixation target. We have added this information about retinal sleep at the beginning of the results section.

      Consequently, it is highly unlikely that retinal slip influenced the results. Furthermore, assuming that tracking error remained consistent across fixation conditions, any present retinal slip cannot explain why the bias followed the retinal curl manipulation as predicted by the controller. We therefore consider retinal slip to be an unlikely confounding factor.

      (5) I found the sections of text comparing the separate and joined fits (starting line 287) to be a bit too rosy. The authors show the separate fits in the main text, and it is not very surprising that these fits are good, given that the model has 30 parameters, and these data are pretty low-dimensional. The authors only show the joined fits in the supplement, and they say that they are almost as good as the separate fits (indeed, they are better in a model comparison sense, but this is 30 parameters vs. 2 parameters). However, when I look at the fits of the joined model in the supplement, I don't find them to be very impressive. In particular, the model grossly misses the data for the straight paths for several subjects (e.g., id5, id6, id8, id10). And fitting the straight paths would presumably be easiest. This implies that the joined model is really missing something and that fitting the curved paths interacts strongly with fitting the data for different fixation target locations on the straight path. I think that the authors should discuss the results a bit more soberly and tone down their conclusions here.

      We thank the reviewer for the opportunity to clarify the logic behind our modeling choices. We acknowledge that the “separate fits” are inherently less informative due to the high number of free parameters relative to the data. Our primary scientific goal was not to achieve perfect descriptive accuracy via 30 parameters, but to test a specific functional hypothesis through the “joint fit.”

      The Logic of the Joint Fit:

      We agree with the reviewer that the joint fit misses some paths in some conditions. Of course, the joint fit reflects a significant compromise. The “Gain” (the weighting of the curl signal) is likely not a static constant but is dynamically tuned based on task demands, confidence in the visual signal, simulated speed, and so on. By using a single Gain parameter, we intentionally ignore this contextual variability to see how much of the behavior can be explained by a “minimalist” controller. In this sense, the 2-parameter joint model is a deliberate attempt to test this limit. By forcing a single Gain parameter to account for all conditions across both straight and curved paths within one flow manipulation (e.g. unaltered flow) we are asking if a single, fixed linear relationship between retinal curl and steering effort/gain can explain the results. We view the joint fit not as a “perfect” model, but as a stronger test of the curl-based control theory. The fact that a 2-parameter model can capture the direction and scale of biases across such a diverse set of conditions (straight/curved paths, five fixation eccentricities) suggests that retinal curl is a robust signal. Upon closer analysis, these discrepancies between the joint model and the data are most pronounced in the over-cancelled condition which is the one when sensory evidence becomes more ecologically inconsistent with the extra-retinal information (gaze direction). While the joint fit successfully demonstrates that a single parameter can capture the general functional role of curl, it fails to account for the complex sensory re-weighting that occurs in ecologically inconsistent conditions (like ‘over-cancelled’ flow). We have updated the manuscript to discuss these limitations in the “fitting the controller” section, framing the model as a parsimonious first-order approximation rather than a complete description of human heading perception based on a minimal set of parameters.

      (6) The section of the paper on neural simulations (starting line 387) has a few weaknesses. First, why are only straight paths simulated here? This does not seem to provide a very rigorous test of the model. Second, it is awkward that the simulation results are presented in units of pixels, rather than degrees. Third, the authors seem to downplay the fact that the neural estimates of heading seem to oscillate rather wildly (over a range of hundreds of pixels, whatever that means, see especially Figure S16). It was far from clear to me how an estimate of heading with these large oscillations is useful. It would seem to require that heading estimates are integrated over substantial lengths of time to be reliable. It was therefore unclear how the model produces such smooth paths from these oscillating estimates.

      We acknowledge that the presentation of the neural model requires more clarity regarding its objectives and its relationship to the behavioral data.

      We first wish to clarify the intended scope of the neural ring-attractor model. Our primary goal was not to provide a comprehensive account of behavioral performance across all conditions (which is the role of the controller model), but rather to demonstrate a biologically plausible mechanism that explains the emergence of the “Opposite-to-Gaze” bias. While the controller demonstrates that the bias follows a specific control law, the neural model shows how such a law can emerge from known primate neurophysiology, specifically, spiral-tuned MSTd neurons, gaze-contingent inhibition, and an egocentric “straight-ahead” prior.

      Why Straight Paths are Sufficient for this Objective. The reviewer asks why only straight paths were simulated. In our study, the straight-path condition with eccentric gaze is the purest test of the bias mechanism. Simulating the straight paths allowed us to isolate the interaction between foveal inhibition and the straight-ahead prior without the confounding variable of path-curvature flow. Given the complexity of the neural network’s parameter space, we focused on these conditions to provide a clear neuro-plausible explanation. We have added text when introducing the model (Neural simulations in the Results section) to make clear why we model straight paths only.

      Units: Pixels vs. Degrees. We acknowledge that the use of “pixels” in the plots of internal neural dynamics may appear awkward. The neural network operates on input stimuli that are defined by the pixel resolution of the videos used in the simulations, we used pixels as the native coordinate system to describe the movement of activity peaks within the network’s internal “map.” We have decided to keep the pixel units in these figures.

      Behavioral Output (Meters): Importantly, the final heading estimates produced by the network are not left in pixels. We use a pinhole camera model to reconstruct the 3D trajectories from the neural activity. These results are expressed in meters, allowing for a direct comparison with the human behavioral data.

      Addressing Wild Oscillations and Smooth Paths. The oscillations observed in the instantaneous heading estimates reflect the stochastic nature of the population peak when tracking high-frequency sensory inputs. In our model, the synaptic time constant (τ) was kept relatively small to ensure a fast, low-latency response to changes in self-motion. While increasing τ would have produced smoother internal dynamics, it would also have introduced delays into the control loop. Instead, we chose to maintain this high sensory responsiveness and applied a temporal moving average later to the network’s decoding to reconstruct the 3D trajectories. This is explicitly stated in the section “Heading Estimation and 3D path reconstruction” in the new appendix 2.

      In addition, the neural activity over time is shown in two ways: the heatmap shows the neuron with preferred heading (one can see more oscillations, specially when the fixation point is closer to the centre (eccentricities -2 and 2), due to larger competition between the sensory evidence and the straight-ahead prior. The other way is the decoded heading. In the ring-attractor model, the decoded heading (φ̂) is not determined by a single neuron but is calculated using a population vector average (equation 19). By summing across the entire population, the decoder effectively integrates sensory evidence from many neurons simultaneously. One can appreciate (see e.g. Fig. 5B) that averaged decoding, leads to a smoother resulting estimate (the white dashed line, whose visibility had been improved in the revised version). Behavioral work by Burr and Santoro (2001) suggests that global motion signals (divergence and rotation in optic flow) are integrated over much longer timescales—roughly 1000ms to 3000ms—compared to local motion units (~200 ms).

      In the previous manuscript, we discussed this aspect in lines 424-426. In the new version, we have added text in the Heading estimation and 3D path reconstruction section (now in appendix 2) stating that we smoothed the decoded signal in agreement with this psychophysical evidence before applying the camera model.

      See also our comment on temporal integration in the responses to reviewer #3

      Reviewer #2 (Recommendations for the authors):

      (7) Line 51: "...a functional role of rotational flow components has been largely neglected in both theoretical and experimental work on heading perception." I feel like this statement is too strong and that the authors try too hard to "sell" their findings by underrepresenting previous work. There are numerous studies (many not cited), both behavioral and electrophysiological, that have examined how heading perception depends on pursuit eye movements, either physical movements or visually simulated ones. These studies directly involve rotational flow components, and several of them have concluded that rotational flow components contribute to estimating heading in the presence of eye movements (just one example is Grigo and Lappe 1999). Because these studies generally involved horizontal pursuit of a target on the horizon, rather than tracking a point in the ground plane (like the authors' work), these studies generally did not involve flow fields with retinal curl around the fixation point. But I consider these older studies just a special case of the more general geometry, and they still involve rotational flow components. Moreover, various previous studies have used stimuli for which there was no FOE present in the visible display (either due to simulated rotation or masking out the FOE), and the authors do not seem to give credit to these works either. In addition, several studies have implicated a role of extraretinal signals in perceiving heading during eye movements, so retinal curl cannot explain everything. Rather than overemphasizing the limitations of previous work, the authors would be better served to explain how their findings extend and generalize from these previous studies.

      We thank the reviewer for pointing out this oversight; it was not our intention to overlook previous work. While our original version cited studies considering rotation-related cue, we have now substantially revised the introduction to include previous work and better acknowledge the informative role of rotation. Our central aim remains to distinguish between models that compensate for rotation to recover a heading vector and our proposal that the visual system exploits retinal curl directly as a primary, functional signal for locomotor control.

      We have now updated the Introduction and Discussion to better situate our work within the context of studies (including Grigo & Lappe, 1999 and some additional ones we have included in the new version) that have investigated the informative role of rotational flow. We now clarify that our study extends these findings by investigating the non-uniform rotational patterns (curl) that emerge during ground-plane fixation, representing a more general and biologically ubiquitous case of locomotor control, while acknowledging previous studies that have also considered the potential role of curl generated by gaze fixation.

      (8) Figure 3: I did not understand why there are two purple and two blue curves in the graphs of the middle column. And the caption does not explain this.

      This a very good observation. This was explained in lines 268-273 (previous version). When the gaze is straight-ahead (same direction as heading), there is no curl. However, we introduced positive or negative curl in the altered conditions. These purple and blue lines refer to these trials and show that the bias re-appears in the expected direction when curl is (unexpectedly) added. We think this adds additional evidence to the curl contributing to heading. Even though this was extensively explained we have added text in the caption of figure 3.

      (9) Line 331: What makes the authors think that retinal curl is computed in area MSTd? They should cite studies to support this idea if it has been shown in physiology.

      Evidence was cited in the introduction (Graziano et al 1994) of sensitivity to spiral motion in addition to neuro-computational models that implement this activity also cited (e.g. work of Leyton et al.)

      (10) The neural network model for computing heading from curl requires a "gaze-centered inhibitory drive" that inhibits activity around where the eyes are looking. This is probably a biologically plausible thing, but is there any evidence to support the idea that this signal exists in the parts of the brain where the authors believe these computations to be happening? They simply posit the existence of this gaze-centered inhibition as though it is common knowledge, but they provide no citations nor discuss any previous evidence for its existence.

      While neurophysiological evidence primarily describes this as gain-field modulation, this process frequently involves localized suppression of neural activity to facilitate coordinate transformations. In parietal areas such as LIP and 7a, eye-position signals do not just enhance responses but can also suppress them, effectively shifting the 'center of gravity' of a population response (Read et al 1997; Born et al 2005, cited in the discussion in the revised section re-evaluating the Focus of Expansion). In the context of our ring-attractor model, this functional modulation is most parsimoniously implemented as a gaze-centered inhibitory drive.

      (11) Lines 482-483: Why should perceptual biases related to retinal curl take seconds to show up?? The curl information itself must be present very quickly, perhaps requiring just a few video frames. So what does this imply about mechanisms? The authors throw out this assertion, but it is left hanging without any further analysis or support.

      The time course reflects the integration requirements of complex motion processing. While local flow is processed rapidly, global patterns like retinal curl require longer temporal windows to reach a stable estimate (Burr et al 2001). In our study, this integration is functionally necessary to filter the higher-frequency 'wobble' induced by gait-cycle oscillations. We now discuss the temporal integration aspects under a new heading in the discussion.

      (12) Line 508: "This suggests that the "bias" observed in our perceived headings may reflect the operation of a control law optimized for action rather than a failure of a perceptual system designed for passive estimation." The authors make this statement to justify why perceptual biases are present with unaltered curl. But I don't fully understand the logic. Are they saying that it is not possible to have a set of computations that can do both things accurately? Is it possible to show this theoretically? Moreover, if it is not possible to rule out other possible sources of the biases, such as those described above (reference frame of judgments, eye movements, etc), then is it necessary to invoke this logic?

      Our logic is that the observed 'bias' is not a representational failure, but a functional byproduct of a control law optimized for active steering. In a closed-loop system, the objective is to null the error signal (retinal curl) to maintain a stable path. When observers are asked to make an open-loop heading report, they likely utilize this same control signal, which manifests as a systematic bias toward the 'null' point of the controller as a result of a sustained fixation in discrepancy with the simulated translation/heading.

      We do not suggest that accurate perception and control are theoretically incompatible; rather, we suggest that perception and action rely in the same underlying information (e.g. work of Brenner & Smeets). While other factors, such as coordinate transformations between reference frames, certainly can contribute to the reporting process, our interpretation provides a parsimonious link between the psychophysical data and the underlying steering mechanism. By framing the bias as a consequence of a 'nulling' strategy, we explain not just the existence of the error, but its specific direction and magnitude relative to the fixated target.

      (13) Line 546: "...MSTd would simultaneously code curvature for trajectory estimation and heading across the neural population, with curvature encoded through the spirality of the most active cell and heading through the visuotopic location of its receptive field center." The latter part of this argument seems to imply a relationship between the heading preferences of MSTd neurons and the locations of their receptive fields. I am not aware of any evidence for such a relationship, so the authors should indicate whether this is based on some experimental data or just a speculation.

      We thank the reviewer for this observation. The proposal that heading is signaled by the visuotopic location of active MSTd populations is a core architectural feature of our model and is supported by several lines of evidence.In the Layton and Browning (2014) framework, MSTd is modeled as a visuotopic map of functional 'hypercolumns'. Each hypercolumn contains neurons tuned to a continuum of spiral patterns, but all neurons in a given hypercolumn share a receptive field center at a specific location in visual space. Consequently, the visuotopic location ($x, y$ coordinates) of the maximally active hypercolumn represents the center of motion (heading), while the spirality (the tuning dimension within that hypercolumn) represents path curvature. We have clarified this in the revised discussion (re-evaluating the FoE) to emphasize that this dual-coding scheme arises from the simultaneous representation of 'where' (population map location) and 'what' (spiral tuning) in MSTd.

      Reviewer #3 (Public review):

      The primary limitation of the paper is that it avoids discussion of some of the inevitable complexities of heading perception. The main issue is what exactly is meant by heading. Different behaviors evolve over different timescales. The geometry of retinal motion defines instantaneous heading, which varies widely through the gait cycle. Time-varying information like this is known to be important in the momentary control of balance. Heading can also be thought of as steering the body toward a distant goal, which evolves over longer timescales. The current manuscript appears to be concerned with heading information integrated over a few seconds and seems to provide evidence that heading is indeed integrated over the gait cycle. The issue of the time scale of the computation is touched on, but it is not related to how it might be used in normal walking or what situations it might apply to. Steering toward a distant goal during walking is not a very difficult problem and may not require evaluation of retinal motion, but control of balance is more challenging and may depend critically on curl. Consequently, the timescale of the computation needs to be considered in order to understand what is meant by heading.

      We thank Reviewer #3 the comments regarding the definition of heading at different time scales, the role of the gait cycle, and the temporal integration of the curl signal. These comments have helped us refine the manuscript’s core arguments.

      We agree that “heading” must be precisely defined within the context of the differing temporal demands of balance and steering. While instantaneous heading provides the high-frequency feedback necessary for momentary postural adjustments and balance, our study is concerned with heading as a gaze-relative signal used for the continuous control of a locomotor trajectory. As such, we have revised the manuscript to specify that the perceived heading measured in our task reflects a signal integrated over the gait cycle to filter out the oscillatory noise induced by head bob and sway (mainly in the Discussion section).

      The reviewer correctly notes that gait-induced head bob and sway produce high-frequency oscillations in the curl signal, yet our behavioral results show smooth, slowly evolving biases. The visual system does not react to “instantaneous” curl, which would lead to jittery, unstable heading estimates. Instead, it integrates flow over a timescale roughly commensurate with a full gait cycle (~500–1000ms). This implies a significant temporal integration process. This temporal integration is consistent with evidence (Burr and Santoro,2001, Vis Res) indicating that optic flow signals (radial and rotational components) are integrated over windows of approximately up to 3 seconds to ensure perceptual stability. Neurally, this likely involves the projection from area MSTd to the Ventral Intraparietal area (VIP), a pathway where fast, eye-centered sensory inputs are transformed into stable, body-centered representations suitable for guiding long-term steering behavior (Chen et al. 2011, JNeurosci.). By grounding our definition of heading in these specific temporal and neural constraints, we tried to clarify how the visual system exploits retinal curl for goal-directed action in natural, dynamic environments and relate our findings to recent studies addressing the role of retinal motion on balance (Powell et al. 2026 Bioarx).

      In our implementation, we explicitly address the high-frequency noise introduced by gait dynamics by smoothing the retinal curl signals computed from the stimulus videos before they are fed into the controller. This temporal filtering allows the fit of the controller’s prediction to the response data while remaining robust to the rapid fluctuations of head bob and sway. In contrast, the neural ring-attractor model would not require an external smoothing step; instead, the integration is an emergent property of the system’s architecture that can be controlled with different parameters, as commented above in a response to Reviewer #2. The dynamics of the synaptic weights and the characteristic “leak” in the population activity naturally implement a leaky integration of sensory evidence, ensuring that the decoded heading reflects a sustained estimate rather than an instantaneous response to visual noise.

      We also agree that we avoided discussing some complexities of the heading perception. In the new version, we also include and integrate the distinction between instant heading and future path in different parts of the ms (introduction) and mainly discussion (temporal integration and steering section) which have been revised substantially.

      Reviewer #3 (Recommendations for the authors):

      There are a number of points that require clarification.

      (1) Head bob and sway were included in the stimulus and need to be addressed in both the analysis of the data and the interpretation. The curl signal in the stimulus varied over time, commensurate with normal gait. However, the results don't reflect the same level of variability that would be produced from curl over the gait cycle. This means that the information must be integrated over some longer timescale. It is not clear from the data analysis what this integration is. Is there an implicit integration with the manipulation of the steering wheel? If subjects indeed appear to be able to use curl to evaluate heading over timescales of seconds, this needs to be explicitly addressed, as it is a novel result. This would require parts of the discussion to be changed/expanded to maintain consistency. For example, line 482 talks about the buildup of biases over time.

      As commented above in the public response, the curl estimated from the optic flow algorithm was smoothed before being input into the controller (path fitting and predictions). The smoothing was only applied to the curl signal, not to the participants responses. Also, as mentioned before, the time course of the bias is consistent with integration times of optic flow reported in the literature. All these aspects are now explicitly included in the new display and conditions section (Flow manipulation conditions).

      (2) Since the experiment included curl variability resulting from the gait cycle, some discussion is needed about the role of retinal motion in the control of balance and posture. There is a large literature about the role of flow in controlling gait and momentary adjustments of the body while walking. Additionally, it should be noted that in the task, head bob and sway from 1 prerecorded individual was shown to all subjects. It is known that gait varies significantly between different individuals, and it should be acknowledged that this may lead to differences at the individual subject level for perceiving heading.

      We have included a point in the discussion addressing the different time scales for different use of optic flow signals (postural control vs locomotion).

      We agree with the reviewer that utilizing a single gait profile for all participants may introduce individual differences in perceived heading, as the simulated head motion might not perfectly match each participant’s unique biological gait signature. However, we prioritized stimulus consistency over idiosyncratic accuracy. By ensuring that every participant viewed the exact same motion profile, we could be certain that the systematic 'opposite-gaze' biases observed across the population were driven by our experimental manipulations of gaze and retinal curl, rather than being confounded by variability in head-motion kinematics. We have added an acknowledgement of this point at the first paragraph of the displays and conditions section in the Methods.

      (3) More information is required about the use of the rotating wheel for the measurement of heading. How easy was it to use? What about time delay, and how does this deal with the bob and sway? Does the wheel impose a de facto integration on the perceptual measurement?

      The rotary encoder provided an intuitive steering-wheel interface that participants found easy to operate. To ensure minimal latency (1–5 ms), the device was interfaced via an Arduino Uno and sampled by a dedicated background Python thread, isolated from the visual rendering loop. We have incorporated these technical details into the Methods (Procedure) section.

      (4) Restructuring the description of the models It remains unclear why the dynamics of the neural network are a necessary inclusion in this paper. It seems interesting, but there is no comparison to actual neural data or other related work. Instead, this appears to be a description of what the network is doing, which is already defined by the equations. This needs to be clarified for its exact interpretation with respect to real neural data, and its importance here for understanding the biases that emerge in heading judgments. The paper would flow better if this section were included as supplementary material or omitted from the paper entirely, as it seems to detract from the other points. If this is a description of why the biases are seen in the controller, then the supplementary material is a good place for it.

      We thank the reviewer for this suggestion. We clarify that the neural model is not intended to simulate specific empirical neural data, but rather to provide a biologically plausible implementation of the controller. This allows us to demonstrate how the observed biases emerge from the dynamics of standard cortical architectures, such as ring attractors. This is now mentioned when introducing the neural simulation results.

      Following the reviewer's suggestion, we have moved the neural model equations to Appendix 2 while retaining the simulation results in the main text (Results). We believe it is essential to present not just the abstract controller, but also its functional implementation, as this provides a mechanistic bridge between retinal signals and locomotor behavior.

      (5) For the modeling approaches, the math would be more appropriate for supplementary materials.

      To ensure a better flow of the paper, we have moved the neural model equations to Appendix 2, while Appendix 1 now details the relationship between measured curl and other variables (speed, yaw, pitch, etc.). We have retained the controller model in the Methods section, consistent with the eLife layout where Methods follows the Discussion.

      (6) How do the models and data analysis deal with the influence of the gait cycle in the input? Do they integrate the information over that timescale? If so, the integration time needs to be specified.

      As commented above, to mitigate gait-cycle fluctuations, we smoothed the computed curl signal before inputting it into the controller and applied a similar smoothing process to the neural model’s readout. Using a LOESS filter, the effective integration window was 2.4 seconds. Close to the integration time reported in Burr et al. 2001. These parameters have now been explicitly specified in the Methods section. For the empirical data, we just utilized trial-averaging.

      (7) What does biologically plausible mean in terms of the neural network model? Especially when control wasn't explicitly a variable in the measured behavior of the subjects.

      By biologically plausible, we mean that our model is constrained by neural architectures documented in the primate brain—specifically ring-attractor dynamics, population coding, and gaze-centered gain-fields. Crucially, the network utilizes recurrent connectivity with a 'Mexican-hat' profile (local excitation combined with lateral inhibition). This is a standard and widely accepted motif in computational neuroscience, representing the consensus on how cortical circuits maintain a stable "bump" of activity to represent spatial variables. Rather than introducing ad-hoc mechanisms, we demonstrate that the observed behavioral biases emerge naturally from these established neural components when they are tasked with maintaining locomotor stability. Since we think this aspect was already emphasized, we haven’t added any additional detail.

      (8) There should be more extensive acknowledgement of the body of literature that has challenged the use of the focus of expansion. That section should also include references to work that has investigated extraretinal signals, as they may also be important.

      We have expanded the Introduction and Discussion to more thoroughly acknowledge research challenging FOE-based models and the critical role of extraretinal signals. These updates, which also align with our response to Reviewer #2, provide a more comprehensive context for our model within the existing body of heading and self-motion literature.

      Minor points:

      (1) Line 69 - In self-generated motion, spiral patterns are almost always centered on the fovea, but many physiological experiments present spirals in the peripheral retina. This is incompatible with the motion generated during self-motion. Therefore, clarify whether the type of spiral motion Graziano investigated was centered on the fovea.

      In the experiments conducted by Graziano et al. (1994), spiral stimuli were centered on the receptive field (RF) of the individual neuron being recorded to accurately characterize its tuning. While the reviewer correctly notes that spiral centers often align with the fovea during active steering (due to fixation on a goal), MSTd neurons possess large RFs that provide a comprehensive 'template' system across the visual field. This population-level representation allows the brain to recover trajectory information even when the focus of motion shifts relative to the fovea—for example, during pursuit eye movements or when fixating on landmarks off the direct path of travel. To keep this part of the text brief, we haven’t add more details concerning this study.

      (2) Line 72 - "Magnitude" instead of "amount".

      This has been changed.

      (3) Line 115 - State explicitly whether the scale of the visual stimulus was matched to the scale of the actual natural images shown in VR.

      We have updated the Methods (Displays and conditions) section to explicitly state that the visual scale was veridical. The virtual camera’s parameters were calibrated such that its field of view (91°) matched the physical dimensions of the projection screen (2.03 m × 1.16 m) at the 1.0 m viewing distance. This ensures that the angular size of the objects and motion gradients in the stimulus were 1:1 with the scale of the simulated natural environment.

      (4) Line 144 - While Farneback is a good dense flow estimation algorithm, it is noisy and may impose biases/variability in the calculation of curl. This should be acknowledged.

      We acknowledge that the Farneback algorithm can introduce variability in curl estimation. To mitigate this, we utilized 10 independent renderings of each experimental trial to compute the flow fields. Although this methodology was reflected in the data previously uploaded to our OSF repository, we have now explicitly added this detail to the manuscript (Flow (curl) manipulation conditions). The computed curl used for the modeling was derived from the aggregate of these different runs, ensuring a robust and stable signal that accounts for potential algorithmic noise.

      (5) Line 206 - "Join fits". Is this a technical term? It sounds awkward. Would "Joint fits" make more sense?

      The referee is right. We have corrected this.

      (6) Line 280 - "Consistent with" (typo).

      This has been corrected.

      (7) Line 286 - Typo in title.

      Also corrected to Fitting the controller.

      (8) Lines 482-493 - There should be more discussion on the time course of integrating the stimulus, and the relationship/generalizability to more natural stimuli.

      This part of the discussion (related to integration time) has been changed considerably to include discussion of postural control in addition to locomotion.

      (9) Lines 516-517 - It is mentioned that retinal flow dynamics override the visual direction cue. This may not be generally true, as the reweighting of the cues might depend on things like task demands or actual stimulus context. In the present experiment, the subject only has access to a large moving textured ground plane, and the body is stationary.

      We agree with the reviewer that cue reweighting is highly context-dependent. However, as noted in the original manuscript (Lines 516-517), we specifically stated that retinal flow dynamics 'can' override the visual direction cue, rather than asserting a universal rule. This phrasing was intentional to acknowledge that while flow is a potent signal—especially in the presence of a large, textured ground plane as used in our paradigm—the relative weighting of these cues remains contingent on the specific sensory and task conditions. We believe this remains a fair and cautious interpretation of our findings.

      (10) Lines 524-256 - It is unclear why Matthis et. al. 2022 is cited for this point. Some of the steering literature, like Wilkie Wann & Allison 2006 or Lappi & Mole 2018 (and some of their other work), would be more relevant for the definition of a control law under these circumstances.

      We agree and this part has been changed substantially.

      (11) Lines 528-532 - Warren et. al. 2001 should be cited in this section because their results were interpreted in terms of focus of expansion, but may result from the curl signal (Powell et. al., 2026. The Role of Retinal Flow in Walking. bioRxiv, 2026-02.).

      We agree with this suggestion and the citation has been added.

      (12) Lines 544-546 - Layton and Browning are focused more on path perception from the implemented spiral tuned cells. Because of this, it wouldn't be appropriate to say it is a shift away from FOE-based heading based on this citation alone. There are more models/psychophysical results that would strengthen this claim.

      We agree with the reviewer that Layton and Browning focus specifically on path perception. Our original intention in citing this work was to emphasize the neurophysiological continuum from radial to circular motion (spiral tuning) rather than discrete expansion/rotation channels. However, we have revised this section to clarify that the reliance on retinal curl represents a mechanism for determining the future path (locomotor trajectory) rather than merely instantaneous heading. This distinction acknowledges that while heading is a momentary vector, the integration of curl signals allows the system to anticipate and control the intended path over time—a framing that better aligns with both the cited literature and our proposed controller model. As commented above this is now extensively discussed in the revised version.

      (13) Figures - There are minor visibility issues for some of the figures. In Figure 2, the thick line is unreadable, and in Figure 1, the x-axis labels are crowded.

      The axis in Fig 1 has been modified to avoid crowdedness. In Fig. 2, the thick (average line) has been modified. We hope they are more visible now.

    1. eLife Assessment

      This manuscript describes a series of studies using four different Go/No Go task variants in combination with fast-scan cyclic voltammetry to determine the role of dopamine release in the ventromedial striatum in action selection, controllability of reward pursuit, effort, and reward approach. The authors conclude that dopamine signals in the ventromedial striatum integrate the invigoration of action initiation with continuous estimation of spatial, but not temporal, proximity to rewards. There is solid support for the conclusions, and the findings are valuable, with theoretical implications for the role of dopamine in reward-driven behavior.

    2. Reviewer #1 (Public review):

      Summary:

      Poh and colleagues investigate dopamine signaling in the nucleus accumbens (ventromedial striatum) in rats engaged in several forms of go/no-go tasks, that differed in reward controllability (self-initiated reward seeking or cue-evoked/quasi-pavlovian), and in the specific timing of the action-reward contingencies. They analysis dopamine recordings made with fast scan cyclic voltammetry and find that dopamine signals vary most consistently to cues that signal a required action (go cues) vs cue signaling action withholding (no go cues). Through various analysis they report that dopamine signals align most clearly with action initiation and with the approach to the reward-delivery location. Collectively these data support aspects of a variety of frameworks related to accumbens dopamine signaling in movement, action vigor, approach, etc.

      Strengths:

      These studies use several task variants that consolidate a few different components of dopamine signal functions and allow for a broad comparison of many psychological and behavioral aspects. The behavioral analysis is detailed. These results touch on many previous findings, larger showing consistent results with past studies.

      Weaknesses:

      The paper is dense and could benefit from some revision to increase clarity of the figures, the methods and analysis. The inclusion of many tasks is a strength but also somewhat overshadows specific points in the data, which could be improved with some revision to focus. There is a lack of strong connection between some of the findings, which if revised would help to emphasize the impact of the work.

    3. Reviewer #2 (Public review):

      Here, the authors record dopamine release using fast-scan cyclic voltammetry in the nucleus accumbens/ ventromedial striatum (VMS) while rats perform variants of a go/no-go task. Two versions are self-paced, in that the rat can initiate a trial by nosepoking at the odor port at any time once the ITI had elapsed, whereas the other two require the rat to wait for a cue-light before responding. Two "long" variants also require either more lever-presses on go trials, or a longer nosepoke time for no-go trials, and also incorporate "free" trials in which the rat is rewarded for just heading straight to the food tray. The authors find that dopamine levels increase more during the response requirement for go than no-go trials, indicating a role for invigorating to-be-rewarded actions. Dopamine levels also steadily increased as rats approached the site of reward delivery, and the authors demonstrate quite elegantly that this was not due to orientation to the food tray, or time-to-reward, or action initiation, but instead reflects spatial proximity to the rewarded location. Contrary to previous reports, the authors did not discern any differences in dopamine dynamics depending on whether the trials were cue- or self-paced, and dopamine release did not scale with effort requirements.<br /> The manuscript is well-written and the authors use figures to great effect to explain what could otherwise be a hard-to-parse set of data. The authors make good use of the richness of their behavioral data to justify or negate potential conclusions.

    4. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      Poh and colleagues investigate dopamine signaling in the nucleus accumbens (ventromedial striatum) in rats engaged in several forms of Go/No Go tasks, which differed in reward controllability (self-initiated reward seeking or cue-evoked/quasi-pavlovian), and in the specific timing of the action-reward contingencies. They analyze dopamine recordings made with fast scan cyclic voltammetry, and find that dopamine signals vary most consistently to cues that signal a required action (Go cues) vs cues signaling action withholding (No Go cues). Through various analyses, they report that dopamine signals align most clearly with action initiation and with the approach to the reward-delivery location. Collectively, these data support aspects of a variety of frameworks related to accumbens dopamine signaling in movement, action vigor, approach, etc.

      Strengths:

      These studies use several task variants that consolidate a few different components of dopamine signal functions and allow for a broad comparison of many psychological and behavioral aspects. The behavioral analysis is detailed. These results touch on many previous findings, largely showing consistent results with past studies.

      Weaknesses:

      The paper could heavily benefit from some revision to 1) increase clarity of the figures, the methods, and the analysis. 2) The inclusion of many tasks is a strength, but also somewhat overshadows specific points in the data, which could be improved with some revision/reworking. 3) Some conclusions are not fully justified. As shown, support for the conclusion "dopamine reflects action initiation but not controllability or effort" is lacking without more analyses and additional context. 4) Further, the notion that the dopamine signals reported here reflect spatial information could be justified more strongly.

      We thank the reviewer for their detailed evaluation and constructive feedback. We have made substantial revisions to address each concern raised:

      (1) Clarity of figures, methods and analyses

      We have revised the organization of the panels in Figure 1 for clarity.

      We have revised Figure 2 and its caption: we have labelled all comparisons depicted in the figure, and now included a line, “Subject-wise comparisons of dopamine data were made for all alignments”, to clearly show that all statistical tests shown in Figure 2c-e were performed between subjects.

      In the caption of Figure 3, we have now added, “... and trial-wise statistical Kruskal-Wallis tests were performed for each task-variant.” to clearly show that statistical tests were performed on the trial-wise level for Figure 3c.

      We have now added a table to the Methods section (Table 1), detailing the sample size in each task variant and the number of trials within each No-go classification.

      We have added more information in the Methods section: we now include Videos to show the classified No-go behaviors and other trial types (Go and Free); and provide a schematic of the DLC workflow in the supplementary materials (Supplementary Figure S9).

      (2) Strengthening specific points in the data

      To improve clarity of our main findings, we have revised the layout of the Results section such as including more descriptive headers:

      “Behavioral performance in Go/No-go (“short task”) was unaffected by controllability of reward pursuit”

      “Behavioral performance in Go/No-go/Free (“long task”) was unaffected by controllability of reward pursuit”

      “Motivation to approach the reward magazine was similar between Go and No-go trials”

      “VMS dopamine release encodes reward-related action initiation”

      “VMS dopamine release does not only reflect reward-related action initiation”

      “Maximum VMS dopamine release encodes spatial but not temporal proximity to rewards”

      “Motivational state reflected by No-go behavioral strategy correlates with dopamine signal size during reward approach”

      (3) (4) Conclusions drawn from data

      We have carefully revised our conclusions to accurately reflect our experimental design and analyses, emphasising our core finding that VMS dopamine was consistently increased in Go versus No-go trials, throughout manipulation of the type of trial start (self- and cue-initiated) and effort manipulation (short and long task variants).

      We thank the reviewer for the feedback regarding our VMS dopamine signals reflecting spatial proximity to reward. We performed additional analyses, which we present in Supplementary Figure S10, S11 and S12, to support our interpretation that VMS dopamine encodes spatial proximity to reward.

      We appreciate the reviewer comment relating to the statement, “Dopamine reflects action initiation but not controllability or effort". We have revised the wording of our conclusion to better reflect our intent, which is to compare the action-selective encoding of dopamine (i.e., action initiation vs action suppression). This subheading is now changed in the Discussion to “Reward-related dopamine depends on action initiation irrespective of controllability and effort”. The additional analyses that we have performed are shown in Author response image 1entitled: “Average Go minus No-go dopamine reveals no effect of controllability (self vs. cue-initiated) or effort (short vs. long).

      Additional details on subjects used in each study, analysis details on trialwise vs subjects-wise data, and other context would be helpful for improving the paper.

      The number of subjects for each task variant was reported in the Methods Section 3: Behavioral procedures in the original submission of this manuscript.

      To improve the paper, we now include a table in Methods Section 6: Statistical Analysis (Table 1), detailing the sample size in each task variant (with FSCV recordings) and the number of trials within each No-go classification.

      To give more context, we made Author response image 1 to illustrate the number of subjects in each task variant (and their overlap):

      Author response image 1.

      Number of subjects included in each Go/No-go task variant (total n = 21). Values (n) depict overlap of each subject between task variants. One animal was recorded in self-initiated Go/No-go and Cue-initiated Go/No-go/Free (dotted line with arrowheads).

      Reviewer #2 (Public review):

      Here, the authors record dopamine release using fast-scan cyclic voltammetry in the nucleus accumbens/ ventromedial striatum (VMS) while rats perform variants of a Go/No Go task. Two versions are self-paced, in that the rat can initiate a trial by nosepoking at the odor port at any time once the ITI has elapsed, whereas the other two require the rat to wait for a cue-light before responding. Two "long" variants also require either more lever-presses on Go trials, or a longer nosepoke time for No Go trials, and also incorporate "free" trials in which the rat is rewarded for just heading straight to the food tray. The authors find that dopamine levels increase more during the response requirement for Go than No Go trials, indicating a role for invigorating to-be-rewarded actions. Dopamine levels also steadily increased as rats approached the site of reward delivery, and the authors demonstrate quite elegantly that this was not due to orientation to the food tray, or time-to-reward, or action initiation, but instead reflects spatial proximity to the rewarded location. Contrary to previous reports, the authors did not discern any differences in dopamine dynamics depending on whether the trials were cue- or self-paced, and dopamine release did not scale with effort requirements.

      The manuscript is well-written, and the authors use figures to great effect to explain what could otherwise be a hard-to-parse set of data. The authors make good use of the richness of their behavioral data to justify or negate potential conclusions. I have the following comments.

      Re: The lack of relationship between effort to acquire reward in the current study and the magnitude of dopamine release, 1) can the authors unpack this a bit more? 2) Why the difference between the Walton and Bouret studies? Were the shifts in effort requirements comparable across the behavioral tasks? 3) What else could be different between the methodologies?

      We thank the reviewer for the feedback and have responded to each of the three questions below (see points 1-3).

      Firstly, we tried to improve the clarity of our research aims. Our primary comparison throughout the manuscript is between Go versus No-go within each task variant. We ask whether the Go/No-go difference in dopamine signaling persists across different response demands. Thus, testing effort was not central to this main question, but rather a feature of the task that did not affect the Go-No-go dopamine difference.

      (1) Consistent with this aim, we show that VMS dopamine differs between Go and No-go actions persistently across all task variants despite differences in response requirements (action was always accompanied by greater dopamine release compared to action suppression). Our behavioral-training data suggest that the ability to perform short and long tasks differed: rats were first trained to criterion on either a short (∼2 s) or long (∼3 s) Go/No-go variant, with the longer variant requiring substantially more training sessions (short: 18.5 ± 7.6 sessions vs long: 41.1 ± 7.9 sessions; see Author response image 2), indicating behavioral demands were higher for the long-task.

      Author response image 2.

      (2) Regarding the apparent discrepancy with Walton and Bouret (2019), we acknowledge that our original description was imprecise (We wrote: “Previous studies have shown that dopamine signals are influenced by the effort required to obtain rewards”). Our intent was not to suggest a direct contradiction, but rather to emphasize that our findings are consistent with the paper’s broader conclusion that effort encoding by dopamine is limited and highly context-dependent. We have now adjusted the manuscript to better reflect our intent by changing the sentence to, “Previous studies have shown that dopamine signals may be influenced by the effort required to obtain reward but only for particular task conditions (Cousins et al., 1996; Gan et al., 2010; see for reviews, Salamone and Correa, 2024; Walton and Bouret, 2019).”

      For added clarity, these were the main results highlighted in the Walton and Bouret review: Gan et al. (2010) demonstrated that VMS dopamine sensitivity to low-effort costs is prominent early in training (≤ 2 training sessions) and diminishes after extended experience (> 9 sessions). Similarly, Hollon et al. (2014) reported that cue-evoked VMS dopamine primarily tracks reward magnitude with minimal modulation by effort. In line with this literature, our rats were highly trained (≥ 9 sessions until the first recording), and exhibited no difference of average Go minus No-go dopamine between short and long task variants within controllability type (see Author response image 3), supporting the idea that extended training exhibits minimal effort-related modulation of VMS dopamine.

      Author response image 3.

      Average Go minus No-go dopamine reveals no effect of controllability (self vs. cue-initiated) or effort (short vs. long). A 2 × 2 Bayesian ANOVA (Cauchy prior: fixed effects r = 0.5; random effects r = 1) consistently favoured the null model over all alternatives. The main effect of controllability and effort showed moderate evidence of absence (controllability: BF<sub>10</sub> = 0.324; effort: BF<sub>10</sub> = 0.309). The model including both main effects performed more poorly (BF<sub>10</sub> = 0.101), and the full model including a controllability × effort interaction was the least supported of all models examined (BF<sub>10</sub> = 0.043). These results provide moderate evidence in favour of H<sub>0</sub>, suggesting that neither controllability, effort, nor their interaction meaningfully predicted average Go minus No-go dopamine responses.

      (3) With respect to task comparability and methodological differences, our behavioral paradigm differs in important ways from those highlighted by Walton and Bouret, where effort was often manipulated by training animals to associate cues with different numbers of lever presses within the same session, and typically involved only action initiation. In contrast, our task required both action initiation and action suppression, and changes in response contingencies occurred across separate recording sessions rather than within-session cue-based manipulations. Although these paradigms are not directly comparable, and only had the same dopamine recording technique in common (FSCV), a key takeaway of our results is that regardless of effort differences, VMS dopamine during action initiation is consistently higher than during action suppression.

      I would argue that the cue- vs self-initiated distinction was pretty minor, given that there was a fixed ITI of 5s. How does this task modification compare to those used previously to show that dopamine release corresponds to behavioral controllability? It would help the reader if the authors could spend more time discussing these disparate findings and looking for points of methodological divergence/commonality.

      We agree that clarifying how our manipulation of controllability compares to prior work improves the manuscript, and we have made the necessary adjustments. However, we would first like to correct an incomplete characterization of the task design.

      While the short-task variant used a fixed 5 s inter-trial interval (ITI), the long-task variant employed a variable ITI ranging from 15–25 s. In the long-task variant, the timing of trial onset was less predictable, and we believe this manipulation reduced animals’ ability to precisely estimate when reward pursuit could begin. Under these conditions, whether trials were Self-initiated or Cue-initiated had a substantial impact on animals’ control over the initiation of reward pursuit. That said, we agree that the Self- versus Cue-initiated distinction overall represents a moderate manipulation of controllability compared to those used in studies that focus on controllability.

      A key source of divergence across studies lies in the definition of controllability. We defined controllability as the animals’ ability to choose the time point of beginning the reward pursuit, rather than whether an action was required, and have now added the following sentence in the:

      - Introduction section: “... controllability of reward seeking, defined as the ability to determine when to initiate reward pursuit (Self- vs Cue-initiated trials)...”;

      - Results section: “We defined controllability as the rats’ ability to choose the time point of reward pursuit. In Cue-initiated trials, the time point at which trials could be started was dictated by a cue light, whereas in Self-initiated trials rats were able to choose intrinsically (control) when to attempt a trial start.”;

      - Discussion section

      Importantly, the action requirements for Go, No-go, and Free trials were identical across these trial-start conditions. We found that the degree to which controllability was manipulated in our task was insufficient to modulate the action-specific VMS dopamine signal (Go vs No-go difference), which remained robust across conditions.

      In contrast, controllability has been defined by others as the presence versus absence of an operant action requirement for reward. For example, Goedhoop et al. (2023) directly contrasted operant (lever press required) and Pavlovian (no action required) conditions, removing action execution as a prerequisite for reward. In that context, cues signaling operant control elicited sustained VMS dopamine release, which was interpreted as reflecting anticipation or preparation for executing a learned action. Similarly, Hamid et al. (2021) demonstrated that dopamine “wave” directionality across striatal regions depends on controllability defined by operant versus Pavlovian conditioning.

      Taken together, these comparisons (results from the present study and in the literature) suggest that dopamine sensitivity to controllability may depend on how it is manipulated. We have clarified these methodological distinctions in the revised Introduction, Results and Discussion, and emphasized that more extreme manipulations (such as removing action requirements entirely or increasing uncertainty over trial timing) may be necessary to reveal controllability-dependent changes in VMS dopamine signaling. Alternatively, the apparent discrepancies across studies may primarily reflect differences in the underlying definitions of controllability rather than conflicting results.

      Reviewer #3 (Public review):

      Summary:

      The manuscript by Poh et al. investigated whether dopamine release in the ventral medial striatum integrates information about action selection, controllability of reward pursuit, effort, and reward approach. Rats were implanted with FSCV probes and trained in four Go/No Go task variants:

      (1) trials were self-initiated and had two trial types (Go vs. No Go) that were auditorily cued,

      (2) trials were cue-initiated and had two trial types (Go vs. No Go) that were auditorily cued,

      (3) trials were self-initiated and had three trial types (Go vs. No Go vs. free reward) that were auditorily cued, and effort was increased,

      (4) trials were cue-initiated and had three trial types (Go vs. No Go vs. free reward) that were auditorily cued.

      The authors report that dopamine levels rose during Go trials and slowly rose in No Go trials, but this pattern did not differ across task variants that modified effort and whether trials were cued or initiated. They also report that dopamine levels rose as rats approached the reward location and were greater in rats that bit the noseport while holding during the No Go response.

      Strengths:

      (1) Interesting task and variants within the task paradigm that would allow the authors to isolate specific behavioral metrics.

      (2) The goal of determining precisely what VMS dopamine signals do is highly significant and would be of interest to many researchers.

      Weaknesses:

      (1) This Go/No-Go procedure is different from the traditional tasks, and this leads to several problems with interpreting the results:

      (a) Go/No Go tasks typically require subjects to refrain from doing any action. In this task, a response is still required for the No Go trials (e.g., continue holding the nosepoke). The problem with this modified design is that failure to withhold a response on No Go trials could be because i) rats could not continue holding the response, as holding responses are difficult for rodents, or ii) rats could not suppress the prepotent go response. This makes interpreting the behavior and the dopamine signal in No Go trials very difficult.

      We appreciate the reviewer raising this important methodological consideration. We acknowledge that our Go/No-go task differs from traditional paradigms used in humans and primates (e.g. Raud et al. 2020, 10.1016/j.neuroimage.2020.11658; Eagle, Bari & Robbins 2008, 10.1007/s00213-008-1127-6; Roitman & Loriaux 2013, 10.1152/jn.00350.2013).

      However, our design addresses the specific constraints of studying dynamics in freely moving rodents while maintaining the core feature of Go/No-go tasks: requiring suppression of a prepotent response. Our task accomplishes the primary aim of our study, which is to compare VMS dopamine dynamics during action initiation and action suppression, and below we list the reasons why. Therefore, we do not believe that this difference compromises the validity and interpretation of our results.

      It has been suggested for decades that the two main processes governed by mesolimbic dopamine are reward learning and motivated action, and our study aimed to better understand how VMS dopamine integrates reward-related information and motivated action, rather than studying them in isolation. To do so, we trained rats in a modified Go/No-go task.

      More recent work (Syed et al. 2016; Hamid et al. 2016; Mohebi et al. 2019) demonstrates that VMS dopamine signaling incorporates both action and reward-related information, rather than either of the two alone. Importantly, in freely-moving rodents, examining this relationship requires preventing the approach response that occurs when reward delivery is anticipated (Pavlovian bias, go for rewards). Traditional Go/No-go designs that simply require "doing nothing" would not achieve this control in freely-moving rats, as animals immediately approach the reward magazine as soon as reward is inferred (as seen in our Free trials). Thus, we require a No-go condition, as we and others have defined (Syed et al. 2016), whereby animals have to actively suppress the ‘initiation’ response. Action initiation is defined at the beginning of the Discussion section: “... at two distinct points after trial start: 1) when rats began lever pressing (Go), and 2) when rats walked to the reward magazine, either without action requirement (Free) or after successful trial completion (Go and No-go)”.

      To further strengthen our interpretation that we compare action initiation and suppression, and to facilitate cross-species translation of our results (i.e., rodent to human), we also include Free trials, where reward delivery requires no specific action (which are essentially like “doing nothing” trials in traditional tasks). This addition allowed us to directly compare No-go and Free trials, where animals must actively suppress responding while maintaining task engagement, to a condition where no overt action is required for a reward, respectively. The dramatic difference in VMS dopamine between No-go and Free trials demonstrates that VMS dopamine reflects active action suppression during No-go trials, rather than merely the absence of action requirements. This has now been discussed.

      Finally, we only report correct Go, No-go, and Free trials, which differs from that of human go/no-go studies that focus on the failure of appetitive no-go trials (i.e., inhibiting the pre-potent response). In the present study, the dopamine signals that we interpret are restricted to successful trials only: action initiation (moving the lever press), action suppression (i.e., suppressing the prepotent Go response while maintaining their position in the nose-poke port), or no action (no lever press, not staying in the port). While this design differs from human Go/No-go paradigms, we believe our study of correctly performed Go, No-go, and Free trials are necessary for isolating action-dependent components of dopamine signaling in freely moving rats (action initiation vs action suppression vs action free).

      (b) Most Go/No Go tasks bias or overrepresent Go trials so that the Go response is prepotent, and consequently, successful suppression of the Go response is challenging. 1) I didn't see any information in the manuscript about how often each trial type was presented or 2) how the authors ensured that No Go responses (or lack thereof) were reflecting a suppression of the Go response.

      We appreciate the reviewer's attention to this important methodological consideration. The originally submitted version of the manuscript already addressed both concerns raised.

      Trial type presentation frequencies

      The Methods section describes our trial presentation approach: "On recording days, the trial types were counterbalanced. Within a session, Go left, Go right, and No-go trials were presented with 33% probability each, without replacement. For sessions with Free trials, trials were presented with a 25% chance without replacement."

      This design results in overrepresentation of Go trials overall (66% in the short-task; 50% in the long-task), which establishes the prepotent Go response as intended in standard Go/No-go paradigms. To improve clarity, this detail has now been included in the Methods section.

      Ensuring No-go responses reflect suppression of Go response

      Our paradigm incorporates multiple features that ensure successful No-go performance reflects suppression of the prepotent Go response:

      First, the overrepresentation of Go trials (addressed above) establishes response prepotency. Second, during No-go trials, rats must maintain their snout in the nose-poke port for the duration of the action cue, which creates the requirement to suppress the natural tendency to approach rewards (i.e., Pavlovian bias; Jones et al. 2017, 10.1016/j.bbr.2017.05.044; Guitart-Masip et al. 2014, 10.1007/s00213-013-3313-4; Dayan et al. 2006; 10.1016/j.neunet.2006.03.002). In Go trials, such natural bias does not require suppression as the lever can be approached and pressed during the action-cue period. This conflict between the instrumental No-go requirement and the Pavlovian-instrumental bias toward action makes action suppression particularly challenging (consistent with computational accounts of similar paradigms; Lloyd & Dayan 2023, 10.1371/journal.pcbi.1011569; Jones et al. 2017, Guitart-Masip et al. 2014, Dayan et al. 2006).

      Figure 3 provides behavioral evidence of this challenge: animals frequently left the nose-poke port and developed spontaneous motor strategies (such as biting and digging) to stay in the port, suggesting Pavlovian bias interfering with response suppression for rewards. Importantly, all reported No-go data include only correct trials (i.e., those without lever presses), ensuring that the dopamine signal reflects successful response suppression rather than failed Go attempts.

      (2) The authors observe relatively consistent differences in the DA signal between Go and No Go trials after the action-cue onset. However, the response type was not randomized between trial type, so there is a confound between trial type (Go/No Go) and response (lever/nosepoke). The difference in DA signal may have nothing to do with the cue type, but reflects differences in DA signal elicited by levers vs. nosepokes.

      As stated in the Introduction section and discussed in our rebuttal to point 1a, the focus of our investigation is how VMS dopamine signals differ during action initiation versus action suppression for rewards, as this is a central unanswered question in the dopamine field. More recent work demonstrates that dopamine incorporates not only RPE but also action initiation (Syed et al. 2016; Hamid et al. 2016; Mohebi et al. 2019), and our goal is to further our understanding of action-dependent VMS signals during reward pursuit.

      The reviewer suggests that dopamine differences may reflect differences in lever vs. nosepoke rather than cue type (Go vs No-go). We respectfully suggest this concern reflects a misunderstanding by the reviewer of our experimental question. The cue-action relationship is the experimental manipulation itself. It is not possible to study how dopamine encodes instructed action initiation versus suppression without linking specific cues to specific actions. The suggestion to 'randomize' action type across cue types would eliminate the very phenomenon we are investigating: how dopamine signals differ when cues instruct different action requirements.

      Our experimental design specifically compares reward pursuit with action requirements (Go trials: lever press; No-go trials: sustained hold) to reward pursuit without action requirements (Free trials: direct magazine approach). This design allows us to isolate how action initiation and action suppression influence reward-related dopamine signaling, which can reveal how the timing of action initiation influences RPE-dopamine. And which is the point of the study: to show how actions influence RPE dopamine signaling.

      Supporting this interpretation:

      Firstly, trial types were randomly interleaved, and each auditory cue explicitly instructed a specific behavioral response. Our design directly follows established methods demonstrating that VMS dopamine encodes whether actions are initiated or suppressed following action cues (Syed et al. 2016). That study, like ours, intentionally linked cue identity to a specific action requirement to assess how dopamine reflects instructed behavioral control. Thus, the fact that Go and No-go cues map onto different actions is inherent to the question being addressed, not an unintended confound.

      Second, as discussed in our response to point 1b, the asymmetry between Go and No-go trials is theoretically essential. Go trials align with Pavlovian approach tendencies (action initiation to reward), while No-go trials create conflict with this bias by requiring action suppression despite the cue being associated with a reward. This Pavlovian-instrumental conflict makes suppression particularly challenging (Lloyd & Dayan 2023, PLoS Comput Biol 10.1371/journal.pcbi.1011569) and allows us to examine the role dopamine in overriding prepotent responses.

      Third, the inclusion of Free trials (discussed in point 1a) demonstrates that our findings reflect instructed action control rather than simply motor execution. Free trials require neither lever pressing nor nose poke maintenance, yet show dopamine dynamics distinct from both Go and No-go trials, confirming that dopamine signals encode action requirements beyond motor output per se.

      Finally, we demonstrate that VMS dopamine differs in the same trial type (No-go) and can be classified based on different movement patterns (Biting, Digging, Calm). Importantly, the difference in VMS dopamine only appeared after the action was completed, particularly during reward approach (Figure 3). This data argues against the idea that VMS dopamine is particularly tied to the specific operant manipulanda as suggested by the reviewer, but rather, may reflect an internal motivational state for reward.

      Together, the aim of the present study is not to redefine Go/No-go paradigms for rodents, but to utilize this task structure to investigate action-dependent dopamine signalling for rewards, which cannot be answered without the cue-action mapping that we have used.

      (3) Both Go and No Go trials start with the rat having their nose in the noseport. One cue (Go cue) signals the rat to remove their nose from the noseport and make two lever responses in 5 seconds, whereas the other cue (No Go cue) signals the rat to keep their nose in the noseport for an additional 1.7-1.9 s. The authors state that the time between cue onset and reward delivery was kept the same for all trial types, and Figure 1 suggests this is 2 s, so was reward delivered before rats completed the two lever presses? I would imagine reward was only delivered if rats completed the FR requirement, but again, the descriptions in the text and figures are incongruent.

      The reviewer asks whether reward was delivered before rats completed the two lever presses and notes incongruence between text and figures. We respectfully note that these details were stated in the originally submitted version of the manuscript (see below).

      Reward delivery timing

      The reviewer asks whether reward was delivered before rats completed the two lever presses, which refers to the short-task variant. No - reward was always delivered immediately after the second lever press for all Go trials. This is described in Methods Section 3: Behavioral procedures - Self-initiated task variant. For added clarity, we have now added the term “immediately”: “... food pellet dispensed into the reward-magazine immediately.”

      In the Results section, we report that the average latency to complete two lever presses was 1.8s, which closely matches the 1.7-1.9s nose-poke hold maintenance required for No-go trials in the “short” variant. Thus, the time point of reward delivery was matched between Go and No-go trial types.

      Representation of reward delivery timing in figure and text

      The reviewer's confusion appears to stem from the schematic representation in Figure 1 and the task variant structure. There were two overarching task variants with different trial requirements:

      “Short-task” variants: Go trials required two lever presses (completed on average in 1.8s);

      No-go trials required 1.7-1.9s nosepoke maintenance

      “Long-task” variants: Go trials required a ‘rewarded’ press to occur 2.7-3.2s after cue onset (completed within ~3s); No-go trials required 2.7-3s nosepoke maintenance.

      For simplicity in depicting action-cue onset in Figure 2c, we used grey shading with a speaker icon at approximately 0-2s and 0-3s to represent these two variants. This schematic representation was not intended to indicate the precise reward delivery time, which (as stated in the Methods) occurred only upon successful completion of trial requirements in the short-task variant, or 2s after successful completion of trial requirements in the long-task variant. To improve clarity, we have adjusted the legend of Fig. 2 for more clarity, adding “Shaded gray area depicts approximate duration of action-cue onset for “short” and “long” task variants.”, and included more information under Results: “In the short-task variants, action-cues switched off after trial completion and a reward was delivered immediately” and “n the long-task variants, action-cues switched off after trial completion, or in the case of Free trials after 3s, and reward was delivered 2s later (Figure 1a).”

      (4) The manuscript is difficult to understand because key details are not in the main text or are not mentioned at all. I've outlined several points below:

      (a) The author's description in the manuscript makes it appear as a discrimination task versus a Go/No Go task. I suggest including more details in the main text that clarify what is required at each step in the task. Additionally, providing clarity regarding what task events the voltammetry traces are aligned to would be very useful.

      We respectfully note that the requested details were already present in the originally submitted version of the manuscript (see below). However, we acknowledge that the task design is complex and may benefit from additional clarity in the main text to aid reader comprehension.

      Behavioral task

      The reviewer suggests our task appears more like a discrimination task than a Go/No-go task. We acknowledge that our paradigm differs from traditional Go/No-go tasks used in humans and primates, as discussed in our responses to points 1a and 2. However, we classify this as a Go/No-go task because it shares the defining feature: requiring action initiation (Go) and suppression (No-go). This classification is consistent with established rodent literature examining action initiation versus suppression (Syed et al. 2016).

      Moreover, as discussed in our response to point 1a, we included Free trials specifically to demonstrate that No-go trials require active suppression rather than discrimination alone. The distinct dopamine dynamics across Go, No-go, and Free trials confirm that our task captures action initiation, action suppression, and action-free states, which is an important contrast that we needed to address our research question about action-dependent dopamine signaling.

      The key requirements for each trial type and task variant are described at the beginning of the Results section. A full description of each step required in the task was provided in the Methods Section 3: Behavioral procedures, to avoid repetition in the Results. Specifically:

      Self-initiated Go/No-go task variant (“short”)

      Self-initiated Go/No-go/Free task variant (“long”)

      Cue-initiated Go/No-go and Go/No-go/Free task variant

      To improve clarity, we have added a reference in the Results section to the Methods: Behavioral procedures for additional procedural details.

      Voltammetry trace alignment

      The events to which voltammetry traces are aligned were stated in the legend:

      “... when traces were aligned to action-cue onset…”

      “... aligned to the time when animals departed the nose-poke port…”

      “... we realigned traces to the moment animals arrived at the reward magazine… “

      Figure 2c legend: "c) Traces aligned to action-cue onset"

      Figure 2d-e legend: “d) Traces aligned to nose-poke exit and e) magazine arrival….”

      However, to improve clarity, we have now added “aligned to action-cue onset.. “ to make the trace alignment more immediately apparent when results are first presented, and added: “Dopamine data were aligned to events of interest: action-cue onset, nose-poke exit, and magazine arrival.”

      (b) How many subjects were included in each task variant? The text makes it seem like all rats complete each task variant, but the behavioral data suggest otherwise. Moreover, it appears that some rats did more than one version. Was the order counterbalanced? If not, might this influence the DA signal?

      The number of subjects for each task variant was reported in the Methods Section 3: Behavioral procedures, where each task variant description includes the corresponding sample size (“Self-initiated Go/No-go task (‘short’; n =9)”, “Self-initiated Go/No-go/Free task (“long”; n = 5)”, “A total of n = 11 and n = 15 were included in the Cue-initiated Go/No-go and Cue-initiated Go/No-go/Free tasks, respectively.”

      For added clarity, we have also added a table for the separation of No-go trials and the number of subjects that it has come from in the Methods.

      Task variant completion

      Not all animals completed all task variants. As stated in Methods Section 4: Real-time dopamine recordings and analysis, animals had to achieve >60% success rate for each trial type on at least two consecutive training sessions to proceed to recording. Other reasons include electrode degradation before all recordings could be completed (See Author response image 1).

      Training order

      We trained four cohorts of animals. One cohort was trained first in the short-task variant

      (Cue-initiated Go/No-go), and the remaining three cohorts were trained first in Cue-initiated Go/No-go (3s) long-task variant (i.e., without Free trials, behavioral and FSCV data not presented in the manuscript). Training order was not fully counterbalanced due to constraints described above.

      Following additional analyses, our data suggest that training order did not influence our core comparison of Go minus No-go dopamine. To directly address whether training order influenced dopamine signals, we separated animals based on whether they were first trained in the Cue-initiated Go/No-go (2s) or the Cue-initiated Go/No-go/Free (3s). We calculated the average dopamine of each rat during the action-cue period and then calculated the difference between them (Author response image 4). We observed absence of evidence of a difference between the groups.

      Author response image 4.

      Average Go minus No-go dopamine reveals no effect of the initial training variant (2s-first vs 3s-first) or test task version (Go/No-go vs. Go/No-go/Free). A 2 × 2 Bayesian ANOVA (Cauchy prior: fixed effects r = 0.5; random effects r = 1) consistently favoured the null model over all alternatives. The main effect of the initial training variant and test task version both showed moderate evidence of absence (initial training variant: BF<sub>10</sub> = 0.378; test task version: BF<sub>10</sub> = 0.367). The model including both main effects performed more poorly (BF<sub>10</sub> = 0.134), and the full model including an initial training variant × test task version interaction was the least supported of all models examined (BF<sub>10</sub> = 0.069). These results provide moderate evidence in favour of H<sub>0</sub>, suggesting that neither the variant animals were first trained on, the task version administered at test, nor their interaction meaningfully predicted average Go minus No-go dopamine responses.

      (5) There is a major challenge in their design and interpretation of the dopamine signal. Both trial types (Go and No Go) start with the rat having their nose in the noseport. An auditory cue is presented for 2-3 s signaling to the rat to either leave the noseport and make a lever response (Go trial) or to stay in the noseport (No Go trial). The timing of these actions and/or decisions is entirely independent, so it is not clear to me how the authors would ever align these traces to the exact decision point for each trial type. They attempt to do this with the nose-port exit analysis, but exiting the noseport for a Go trial (a rat needs to make 2 lever presses and then get a reward) versus a No Go trial (a rat needs to go retrieve the reward) is very different and not comparable.

      We respectfully disagree with the reviewer’s assertion that our data alignment approach is problematic. Aligning neural activity to specific behavioral epochs that occur at different times and across conditions is a widely used method to investigate the relationship between neural activity and behavior. Just to mention some examples: data collected with fiber photometry (e.g., Tan et al. 2026, doi: 10.1038/s41386-026-02368-4; Hart et al. 2024, doi: 10.1016/j.celrep.2024.113828) and voltammetry (Hamid et al. 2016, doi: 10.1038/nn.4173; Syed et al. 2016, doi:10.1038/nn.4187).

      Alignment method

      We intentionally designed the task so that overall action timing is matched between trial types (as described in our response to point 3), while specific behavioral epochs occur at different times. This allows us to compare dopamine dynamics during comparable behavioral events across Go and No-go trials (e.g., nose-poke exit).

      We align data to three critical behavioral epochs, stated in the Methods Section 4: Real-time dopamine recordings and analysis - FSCV measurement and analysis: action-cue onset, nose-poke exit, and magazine arrival. Each alignment addresses a specific aspect of our research question:

      Action-cue onset alignment captures VMS dopamine dynamics when animals have explicit knowledge of trial type and the required action. This allows us to characterize how dopamine evolves following correct action selection, which is central to our research question about how dopamine differs during successful Go versus No-go action execution, as well as no overt action (Free) in the long-task variant.

      Nose-poke exit alignment captures dopamine dynamics at the moment animals initiate movement. The reviewer suggests that exiting for Go versus No-go trials is "very different and not comparable" because subsequent actions differ (two lever presses vs. direct reward retrieval). However, this is precisely our experimental manipulation: we compare dopamine signals when animals exit the nose-poke port to perform different actions. This comparison is both valid and necessary to address our research question (does dopamine encode action initiation?).

      Magazine arrival alignment captures dopamine dynamics at reward approach. This allows us to differentiate between spatial proximity to reward from other concepts including temporal proximity (how soon is reward) and action requirements.

      We acknowledge that we cannot identify the precise moment of decision formation. In fact, the precise moment of decision formation is irrelevant for our question. However, our aim is to characterize VMS dopamine dynamics during successful action execution for rewards, following cues associated with specific actions.

      (6) The voltammetry analysis did not appear to test the hypotheses the authors outlined in the intro. All comparisons were done within task variants (DA dynamics in Go vs. No Go trials, aligned to different task events), but there were no comparisons across task variants to determine if the DA signal differed in cued vs self-initiated trials.

      Our aim was to investigate whether VMS dopamine signals consistently differed between action initiation and action suppression during reward pursuit. To test this within-variant contrast (Go > No-go), we manipulated how reward pursuit is initiated (self- vs cue-initiated) and the “effort” requirements (short vs long task variants), and our results show that they did not affect the differential between Go and No-go.

      The consistent Go > No-go dopamine that we observed across all task variants, together with the consistent increase during magazine approach, supports our conclusion that VMS dopamine integrates motivated action and reward.

      The reviewer suggests that we should have compared dopamine signals across self- vs cue-initiated task variants. We acknowledge this is an interesting, but entirely different, question and have addressed it in the Discussion section. We note that differences in controllability altered the time course of increased VMS dopamine, presumably by triggering earlier positive RPEs in Cue-initiated tasks as compared to Self-initiated tasks (illumination of the nose-poke light being the earliest predictor of reward). However, since our primary research question relates to the difference in VMS dopamine between action initiation and suppression, our results show that this difference was unaffected in two variants (short and long), strengthening our conclusions about the relationship between action and reward-related dopamine signaling.

      (7) Classification of No Go behaviors was interesting, but was not well integrated with the rest of the paper and was underdeveloped. It also raised more questions for me than answers. For example:

      (a) Was the behavior classification consistent across rats for all No Go trials? If not, did the DA signal change within subjects between biting vs digging vs calm?

      (b) If "biting rats" were not always biting rats on every No Go trial, then is it fair to collapse animals into a single measure (Figure 3C).

      (c) Some of the classification groups only had 2 or fewer rats in them, making any statistical comparison and inference difficult.

      Behavioral classification for each rat was consistent across trials (i.e., 100%, see Author response image 5). Upon reviewing the consistency of classifications within individual animals, we found that “Biting” animals exhibited biting behavior across the majority of their No-go trials. Only one animal (in the Self-initiated Go/No-go "short" variant) showed mixed classifications across trials, occasionally exhibiting digging or calm behavior. For all other animals, the predominant behavioral classification was highly consistent within subjects across sessions. The occasional trials where “biting” animals did not bite were too infrequent to permit meaningful within-animal comparisons. Therefore, we believe collapsing animals by their predominant behavioral phenotype in Figure 3C is appropriate and accurately represents stable individual differences in No-go response strategies.

      As stated in the Methods Section 6: Statistical Analysis - Clustering No-go behaviors and regrouping animals (last-line), we specifically avoided between-subjects statistical comparisons for groups with n≤2, as this would be inappropriate (see Author response image 5), and reported qualitative observations only. These exploratory findings at the individual-trial level suggest behavioral heterogeneity during No-go trials, that others may use for future investigation, but do not form primary conclusions.

      The behavioral classification is integrated with our central findings on VMS dopamine encoding spatial proximity. Our results demonstrate that individual variation in action suppression strategy, in particular Biting behaviors, consistently manipulates the timing of max dopamine release during subsequent reward approach, but not during the action itself (Figure 3b-c). This links our observations of action-dependent dopamine (Figure 2c) with spatial reward approach (Figure 2e). We believe that our findings shed new light into the understanding of how action modulates reward-related dopamine dynamics at the individual level. This has been discussed in Discussion section: Dopamine dynamics are linked to motivated action.

      Author response image 5.

      Rats predominantly stick to a particular strategy to perform No-go trials. Each bar represents an individual animal, and colours represent the % of each classification type.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) Figures: It would be helpful to have more panel labels on the figures - 1C, for example, labels 4 different dataset panels, similar to a few other cases. This would really help to improve readability, as there are a lot of task types, trial types, behavioral measures, and labels to sift through. The figures overall are very busy, and it is a bit of a challenge to process the tasks and trial comparisons, as well as interpret what the quantification insets mean.

      We thank the reviewer for their suggestion. We have added more panel labels to Figure 1 to improve the ability to follow with the main text. We hope that the reviewer finds it acceptable.

      (2) Figures: A bit more specifically, in Figure 2, the boxplot insets are pretty hard to see, and it's not clear what scale they are on or what data they reflect. Similarly, it's unclear what the horizontal bars reflect in terms of which conditions are being compared. Why are box plots used for some comparisons, and why are some comparisons based on time series bootstrapping, but others are not clear? I would consider broadly reworking this figure and its description for clarity. Figure 3, by comparison, is easier to understand - the quantifications are clearer and labeled.

      We have now labeled the comparisons being depicted by the horizontal bars in Fig 2c. We have also clarified the boxplot analysis in Fig 2d and Fig 2e by adding 'Max dopamine’ labels to the figure, and have made these analysis methods more explicit in the figure legend.

      We used time-series bootstrap analysis to identify when dopamine signals diverged between trial types, which depicts dopamine differences during distinct action requirements. The latency-to-max quantification provides a summary measure to test specific encoding hypotheses (temporal vs. spatial proximity to reward).

      (3) Broadly, more clarity on the FSCV analysis is warranted.

      (a) Targeting: It looks like the dataset contains a mix of medial shell and mostly core accumbens placements. The paper treats VMS as a uniform dopamine region, but it is more standard to separate core and shell (and also other parts of the shell) into subregions. Many of the reported encoding profiles here are known to differ across the accumbens. So, some consideration of this seems appropriate - a minimal signal-behavior comparison for the shell vs core subgroups, for example.

      We thank the reviewer for raising this point. Upon careful re-examination of our histological analysis, we identified an error that occurred when we made the overlay of electrode placements across rostrocaudal planes (to project placements onto a single plane for the sake of simplicity; Fig 2a): We incorrectly assigned some recordings to the nucleus accumbens shell. We corrected this error, which shows that the vast majority of recordings were in the core, with only 2 animals in the shell (black stars). We have adjusted Fig 2a to reflect this, and added an anatomically more complete illustration of electrode placements across rostrocaudal planes as Supplementary Figure S8.

      While we acknowledge reported differences between core and shell dopamine in some contexts, the small number of shell placements precludes meaningful statistical comparison. However, to determine whether average dopamine concentrations differed between core and shell during the action-cue period, we have plotted the values in Author response image 6. The fact that shell data mostly falls centrally into the overall core-data distribution suggests no consistent difference in dopamine release.

      Furthermore, in our experience (and that of colleagues (personal communication)) with appetitive operant tasks, core and shell FSCV dopamine signals do not substantially differ for action-selective encoding and reward approach. Given the sample distribution and our focus on general VMS function in Go/No-go behavior, we believe pooling these regions is appropriate.

      Author response image 6.

      Average dopamine release during the action cue in nucleus accumbens core (circles) and shell (stars) animals showed no distinct separation between regions. Each symbol represents an animal.

      (b) Design: In my understanding of the design, the main distinction between the short and long task variants is a 2-second versus a 3-second required nose poke hold. 3 seconds here is "long" and more "difficult". I'm not sure I agree that a 1-sec distinction really reflects a difference in task difficulty or effort. Can the authors point to a past paper that demonstrates this variation is sufficient to engage a behavioral difference and/or a neural encoding difference? Broadly, some justification of the validity of this manipulation is needed, I think.

      While we lack direct citations for this specific manipulation, we believe that the 2s and 3s hold requirements represent meaningful differences in difficulty based on our extensive rat-behavior experience and behavioral evidence in Author response image 2.

      Importantly, the difficulty of No-go trials does not stem merely from the required time to hold their snouts in the nose-poke port, but from suppressing the motivational/Pavlovian bias to approach reward-associated cues. In No-go trials, subjects must suppress the prepotent tendency to immediately approach reward-related stimuli and instead maintain active suppression of this approach behavior. Even the 2s hold is challenging as animals tend to perform better on Go trials compared to No-go trials (Fig 1b and 1c), demonstrating the inherent difficulty of response suppression even at the shorter duration. The additional 1-second substantially increases this demand, as it represents a 50% increase in hold duration.

      Our training data clearly demonstrate the difficulty in reaching task criterion when increasing the required action (for Go and No-go trials) from 2s to 3s. Across four cohorts of animals trained in Go/No-go task variants, one cohort that was trained first in the 2s task variant, and the remaining three cohorts were trained in the 3s task variant of Go vs No-go. Animals required 41.1 ± 7.9 (n = 27) sessions to learn the 3s hold (approximately 8 weeks), versus

      18.5 ± 7.6 sessions for the 2s hold (approximately 4 weeks; mean ± SEM). This indicates that despite only a 1-second difference in required action performance, animals needed more than double the number of training days to reach criterion, clearly indicating differential effort demands.

      (c) Figure 2 results: the authors state that because there is a greater DA signal to Go vs No Go cues in all the task variants, this means that controllability of reward pursuit and increased task effort do not affect VMS dopamine. But the magnitude of the signals looks different across the task variants - it looks clearly stronger overall in the self-initiated tasks, for example. Given that dopamine signals are not compared across task variants (I think the tasks are all between-subjects?), I don't think the above conclusion is justified.

      We respectfully clarify that our conclusion does not claim controllability and effort have no effect on dopamine magnitude, but rather that these manipulations do not affect the action-selective difference in dopamine (Go > No-go). Our central finding is that the relative difference between Go and No-go remains consistent across all task variants (within-subjects comparison).

      We did not perform across-variant comparisons of absolute dopamine magnitudes because that was not our primary research question. Our focus was to understand whether dopamine differs between action initiation and suppression, and whether this difference can be modulated by controllability or effort.

      We acknowledge the reviewer’s observation that absolute magnitudes appear larger in self-initiated vs cue-initiated task variants. We believe that this likely reflects differences in RPE timing rather than controllability per se: in cue-initiated tasks, the nose-poke light provides an early trial-start signal, distributing RPE temporally across the trial. In self-initiated tasks, trial-initiation and action requirements are temporally integrated. Though understanding how controllability affects absolute dopamine magnitude is an interesting question for future work (e.g., using sophisticated regression-based encoding models), it was beyond the scope of our current investigation, which focuses on action-selective encoding.

      (4) Broadly, I don't think these data, as shown, support the conclusion "dopamine reflects action initiation but not controllability or effort" without more analysis and additional context.

      We have revised the wording of our conclusions throughout the manuscript to better reflect our intent, which is to compare the action-selective encoding of dopamine (i.e., action initiation vs action suppression). It is now “Reward-related dopamine depends on action initiation irrespective of controllability and effort”.

      (a) Figure 3 - more description of the classified behaviors would be helpful for interpreting this part of the data. When are the behaviors occurring - during the hold cue? Or is the classification related to what they do immediately after holding? Or something in between> I guess I'm not sure what digging and biting are in the context of a nose poke hold. As described, it's not clear what the signal differences relate to - movement differences? Generally, it's not clear what to make of the behaviors. They seem to emerge spontaneously, but it's not clear whether the specific actions mean anything, so it's a bit difficult to know what to glean from the dopamine is greater during "biting". It's a very different movement pattern, so perhaps this result relates to that, rather than task engagement or motivational drive per se?

      We thank the reviewer for the comment. We have added relevant information in the figure caption and in the Results section to clarify that classified behaviors occurred during the action-cue period (for No-go trials, the hold cue; Figure 3a caption). In addition, we have included Videos to better depict the classified No-go behaviors during the action-cue period.

      We agree with the comment that these classified behaviors, such as biting, seem to emerge spontaneously. Our interpretation of these behaviors is that they may represent the motivational state of each subject. Most importantly, whereas the behavioral differences occurred during the action-cue period (while animals had to suppress actions and stay within the nose-poke port), the difference in VMS dopamine was only observable after this behavior was completed. Thus, the movement pattern per se is likely not relevant to the dopamine release occurring after its completion. This has been discussed in the Discussion section: Dopamine dynamics are linked to motivation action.

      Based on our videos, it appears as though Digging could be perceived as more vigorous (i.e., more general movement in the nose-poke port). However, we did not observe more dopamine during the action-cue period of Digging trials as compared to Biting trials. Furthermore, more vigor during the action-cue period (e.g. Digging trials) did not result in more dopamine during the reward approach period. Together, the data suggest that another process may underlie the large increase in VMS dopamine in Biting trials during reward approach, such as varying attribution of incentive salience.

      (b) In some cases, but not all, dopamine measurement comparisons are done on a total trial basis, and in others, it seems to be subject averages. It's not clear why different approaches are used for different parts of the data. But also, for the trialwise analysis, what statistical steps were taken to incorporate the subject as a random factor in the analysis? If that is not done, then a trial-wise analysis artificially increases the power for the stat (n=trial#).

      We used different analytical approaches depending on sample size and data structure. To compute differences in Go vs No-go dopamine within each animal, as intended by our experimental design, we performed subject-level comparisons (Figure 2).

      For the behavioral classification analysis (No-go, Figure 3), we performed trial-level analyses to increase the statistical power and better characterize this unexpected and interesting phenomenon. We explicitly chose not to perform subject-level group comparisons because

      (1) some groups had only n=2-3 animals, making subject-level statistics underpowered, and (2) behavioral classifications were highly stable within individual animals (see Author response image 5). We acknowledge that formal between-group comparisons (across subjects) are underpowered due to small n, but the stability of within-subject No-go behavioral strategy and qualitatively distinct VMS dopamine profile suggest that these differences may be biologically meaningful and worthy of future investigation in larger samples. We have made these limitations more explicit in the Results.

      This relates to Figure 3, where all trial data are shown next to individual subjects - the subject-wise group comparisons are between 2-5 or so rats, which is quite low. In Figure 2, a subject n of 27 is listed, so it's not clear why this analysis is on such a small set of rats. Generally, it's not clear how many rats/subjects are in each data bit. The methods say only 5 rats are in the long self-initiated task, but 15 in the cue-initiated task. Clarity in all this is needed, including in the figures/captions.

      We have now added detail about the statistical test performed in Figure 3’s caption:

      “After action-cue offset: No-go (trials)’: Individual No-go trials classified by No-go behavior, and trial-wise statistical Kruskal-Wallis tests were performed for each task-variant.” We have also added in a table in Methods (Table 1), showing the number of trials in each No-go classification, and animals regrouped based on their predominant No-go strategy (see Methods section: Statistical analysis - Clustering No-go behaviors and regrouping animals).

      For the small sample sizes based on the regrouping of animals based on their predominant No-go strategy, we have now added in the caption, “... Rats classified based on their predominant No-go strategy, with no statistical tests performed.”

      (5) I'm also a little confused about the paper's narrative that the dopamine data reflect spatial (but not temporal) proximity to reward - it seems that this conclusion is based on the dopamine signal peaking at magazine entry, but that is different, I think, than a spatial signal per se (space is not manipulated in this study). I think more analysis of the signals during the magazine approach behaviors would be helpful, and possibly comparing rewarded vs unrewarded approaches. The emphasis, including in the title, that a major take-home of the data is that dopamine encodes reward proximity, is not really borne out by the current analyses. Reward expectation is not manipulated independently of the approach action, so it's hard to pin this on "space" vs "reward is soon". This is admittedly a general complexity in characterizing dopamine ramps.

      As the reviewer notes, we acknowledge that 'spatial proximity' and 'reward is soon' are challenging to fully dissociate in appetitive approach paradigms. However, we believe that our data and new additional analyses, which is now included in Results: Maximum VMS dopamine release encodes spatial but not temporal proximity to rewards and Supplementary Figures S10-12, provide compelling evidence that VMS dopamine primarily reflects spatial proximity to the expected reward location, rather than temporal proximity to reward delivery.

      Evidence against full temporal encoding:

      (1) Max VMS dopamine occurred up to 3s before reward delivery in Free trials (Fig 2e, open circles vs triangles). Furthermore, when we calculated the max values of individual trials of realigned traces, maximum dopamine does not consistently coincide with reward delivery across trial types (new Supplementary Figure S11).

      (2) If dopamine encoded temporal proximity from the earliest reward-predictive cue, we would expect consistent accumulation from cue onset (action cue for self-initiated, nose-poke light for cue-initiated). However, realigned trials also did not show a consistent accumulation around these events (new Supplementary Figure S12).

      Evidence for spatial encoding:

      (1) Max dopamine consistently occurred when animals arrived at the magazine across all trial types and task variants, regardless of when reward was actually delivered (Fig. 2e).

      (2) Across individual trials, max dopamine values concentrated around the magazine-panel, with a striking accumulation when animals were in close proximity to it (new Supplementary Figure S10) showing a distance-dependent distribution.

      (3) Assessment of unrewarded magazine approaches during the intertrial interval (ITI) revealed no increase in dopamine release (Fig 2f), indicating that VMS dopamine requires task-relevant reward expectation.

      (6) Examples of the DLC workflow in a supplement would be appropriate. Also, video examples of the 3 kinds of behaviors from the clustering analysis could be useful for understanding what they are/what they mean.

      We have added supplementary figures showing the DeepLabCut workflow (Supplementary Figure S9) and Videos 1-3 demonstrating the three behavioral classifications (biting, digging, calm) during No-go trials.

      (7) Referencing/scholarship: I would suggest broadening the citation pool for the paper to include more older work that has established the notion that dopamine signaling and the accumbens act as a motivation-action interface, as this has been a longstanding notion since at least the 1980s. There is also a sizable literature on dopamine signaling of effort, some of which would be appropriate to cite here.

      We thank the reviewer for the recommendation and have now expanded our citations to include foundational literature to work from the 1980s-90s. These can be found in the Introduction, Results, and Discussion.

      Reviewer #2 (Recommendations for the authors):

      (1) Lines 353-357- This came as a surprise, as the relevant results are only featured in supplementary information. These should be moved to the main manuscript. As both the biting behavior and faster lever-press completion lead to larger peak dopamine, does this represent response vigor?

      We appreciate the reviewer’s interest in these data. However, we believe that the data that we report in “Supplementary Fig 7: Quartile analysis of Go trials shows coordinated changes in last lever-press timing and VMS dopamine, does not warrant movement to the main manuscript.”

      The purpose of the lever-press timing analysis was to demonstrate that maximum dopamine release coincides with the moment animals arrived at the magazine, rather than with the action period itself. By sorting Go trials based on last lever-press latency, we show temporal coordination between action completion and dopamine timing but critically, dopamine peaks after action completion, not during it.

      This temporal dissociation argues against a 'response vigour' interpretation. If dopamine encoded motor vigour, we would expect the signal to coincide with or precede the vigorous action. Instead, both the lever-press data (Supplementary Figure S7) and the classified No-go trial data (Figure 3, biting behavior) show that changes in max dopamine occur after the actions themselves, during the subsequent approach to reward.

      Together, these findings demonstrate that VMS dopamine reflects spatial approach to the reward location following action completion, not the vigour of the required actions per se. The lever-press analysis serves as supporting evidence for this temporal relationship but does not introduce a novel finding that warrants main figure emphasis. We mentioned this temporal coordination in Results to ensure readers are aware of the converging evidence while maintaining focus on our central findings regarding action initiation versus suppression.

      (2) Line 13- the experimental work cited refers to midbrain dopamine neurons, rather than dopamine release within the VMS. Please correct.

      We thank the reviewer for pointing out this error. We have now corrected the citation to refer to studies measuring striatal dopamine rather than midbrain dopamine neurons.

      (3) Line 166 - Shouldn't this say "consistently delayed for No Go trials"? Looks like peak dopamine occurs later for these trial types.

      This statement refers to traces aligned to nose-poke exit (not action-cue onset). The peak Go dopamine occurred later than other trial types following exit (Green arrows in Figure 2d).

      (4) Lines 391-392 - however however

      Thank you for the comment, we have adjusted the text.

    1. eLife Assessment

      In their important manuscript, Gangadharan, Kober and Rice focus on how Stu2/XMAP215-family microtubule polymerases use their TOG domains to catalytically promote microtubule growth, testing whether their mechanism follows an enzyme-like kinetic model similar to that of actin polymerases. The authors integrate measurements including microtubule polymerization rates and TOG-tubulin binding kinetics to convincingly show that Stu2 follows an enzyme-like model where tight tubulin binding enables efficient polymerization, revealing a shared mechanism with actin polymerases despite their evolutionary divergence. This work will be of general interest to the cell biology and biophysics communities.

    2. Reviewer #1 (Public review):

      This study by Gangadharan and colleagues provides significant progress towards a quantitative biochemical mechanism for Stu2 polymerase activity. A key conceptual advance is the novel application of an enzyme-like model, initially developed for the actin polymerase Ena/VASP, to Stu2.

      Strength:

      New refined affinity measurements for a Stu2 TOG domain using Bio-layer interferometry show more than an order of magnitude higher affinity of TOG domains to tubulin compared to previously published reports.

      The findings reinforce the "concentrating reactants" or, more specifically, for TOG-domain proteins, the "tubulin-shuttling antenna" model, compared to the "polarized unfurling" model, a more speculative structural hypothesis.

      The manuscript builds upon a series of previous manuscripts that showcase the profound intellectual engagement with microtubule polymerization mechanisms by TOG-domain proteins from the Rice lab, a thought leader in microtubule polymerization for over a decade.

      Minor weakness:

      The affinity discrepancy is not fully resolved by side-by-side measurements, which seem to be not feasible as not all buffer conditions are compatible with all assays.

    3. Reviewer #2 (Public review):

      Summary:

      The manuscript from the Rice lab by Gangadharan et al., submitted to eLife, investigates the polymerization mechanism of the yeast microtubule polymerase Stu2. The lab has published a number of articles demonstrating the structural basis by which the two TOG domains of Stu2 each bind free tubulin heterodimers and has developed a tethered polymerization model by which the TOG domains drive polymerization by shuttling those tubulin subunits onto the microtubule plus end. A second model was proposed by Nithianantham et al. (eLife, 2018) based on a closed - to - open transitional state in which Stu2 unfurls and loads two longitudinal associated tubulin heterodimers onto the microtubule plus end. While the second model is not directly tested, the current work aims to further characterize/model the tethered polymerization model using a kinetic framework developed by developed by Breitsprecher et al. for Ena/VASP actin polymerization activity, using a model that is enzymatic (EMBO J., 2011). The general architecture and function of Ena/VASP on actin polymerization versus Stu2 on microtubule polymerization is a reasonable relation and hits upon, as the authors note, potential convergent mechanistic evolution across distinct cytoskeletal networks. The model effectively treats tubulin as the substrate, and the polymerized microtubule plus end as the product. If Stu2 is "enzymatic" in this framework, the model predicts it would behave with Michaelis-Menten kinetics, that there would a Vmax, and polymerase activity would either be "affinity limited" by TOG:tubulin affinity (KD) and/or "kinetically limited" by TOG:tubulin association (Kon) and transfer of tubulin to the microtubule plus end (Kt). The authors find that the Brietsprecher model works well for Stu2 activity, and that Stu2 best aligns with a "kinetically limited" model. The work is interesting and adds to the growing elucidation of the Stu2 microtubule polymerase model. While yeast microtubule polymerases are somewhat distinct in their architecture, there is significant overlap that findings from the manuscript can be utilized to inform the mechanisms of larger, more complex microtubule polymerases such as human ch-TOG.

      Strengths:

      The manuscript invokes the enzymatic model of Breitsprecher et al. used for Ena/VASP and conducts an elegant series of (mostly established) experiments to determine whether Stu2 microtubule polymerase activity aligns with the model - which they conclude does align, supported by the data/results obtained.

      Weaknesses:

      The authors used biolayer interferometry to measure TOG:tubulin affinity. The affinities obtained were significantly higher affinities than the lab obtained in an earlier publication using analytical ultracentrifugation. While differences in buffer and salt conditions may underlie these differences, additional runs using comparable buffer systems, or use of a third independent assay to measure affinities would have added rigor.

      The discussion could be expanded to better compare and contrast the results with both existing polymerase models introduced in the introduction, as well as expanded to look at reversible enzymatic activity (microtubule depolymerization at low to zero tubulin concentrations) and microtubule plus versus minus end activity.

      Comments on revised version.

      The revised submission has addressed these comments adequately.

    4. Reviewer #3 (Public review):

      Summary:

      This study by Gangadharan and colleagues seeks to establish a quantitative biochemical model for the microtubule polymerase activity of Stu2. Stu2 is the budding yeast member of the XMAP215 protein family, which is broadly conserved across eukaryotes. XMAP215 proteins play a wide variety of important roles in cells, and these are attributes to effects on microtubule dynamics. Many studies over the last ~20 years have shown that XMA215 proteins selectively associate with microtubule ends where they increase rates of microtubule assembly and disassembly. More recently, structural biology and biochemical studies by the authors and other groups have shown that the multiple TOG domains on XMAP215 proteins are tubulin-binding domains that selectively bind to curved tubulin, which is present in solution and at microtubule ends, but not to straight tubulin which is present in the walls of the microtubule lattice. This has led to the general model that XMAP215 proteins promote polymerization by delivering soluble tubulin to the growing plus end, and two distinct models have been proposed to explain the mechanism. The 'concentrating reactants' model proposed previously by the authors suggests that TOG domains grab hold of tubulin in solution and concentrate at the microtubule end. The 'polarized unfurling' model proposed by the Al Bassam lab suggests that XMAP215 delivers multiple tubulins to the end, using a stepwise mechanism involving different roles for each TOG domain. The current study seeks to improve our understanding of the mechanism by developing a quantitative model to explain the binding and release of tubulins, the number of Stu2 molecules at the end, and the overall rate of tubulin addition. The authors accomplish this goal using new experimental data. The final model fills in new details of the mechanism. The authors draw a comparison between Stu2 and the actin polymerase which bears similarity to the Ena/VASP and suggest a convergent strategy for cytoskeletal polymerases.

      Strengths:

      This is a focused and clearly written study that incorporates prior knowledge of XMAP215 and draws inspiration from the actin field. The data are clear and convincing, and the study accomplishes its goal of generating a new, quantitative model for Stu2. The model will be important for microtubule researchers to predict and test key points for altering XMAP215 activity across different organisms and potentially for different tubulin substrates. The comparison to Ena/VASP may also inspire similar comparisons across other microtubule and actin regulators, which could lead to new insights across cytoskeletal fields.

      Weaknesses:

      The study is without major weaknesses.

    5. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      In their important manuscript, Gangadharan, Kober and Rice focus on how Stu2/XMAP215-family microtubule polymerases use their TOG domains to catalytically promote microtubule growth, testing whether their mechanism follows an enzyme-like kinetic model similar to that of actin polymerases. The authors integrate measurements including microtubule polymerization rates and TOG-tubulin binding kinetics to convincingly show that Stu2 follows an enzyme-like model where tight tubulin binding enables efficient polymerization, revealing a shared mechanism with actin polymerases despite their evolutionary divergence. This work will be of general interest to the cell biology and biophysics communities.

      Thank you for the favorable assessment of our manuscript.

      Public Reviews:

      Reviewer #1 (Public review):

      This study by Gangadharan and colleagues provides significant progress towards a quantitative biochemical mechanism for Stu2 polymerase activity. A key conceptual advance is the novel application of an enzyme-like model, initially developed for the actin polymerase Ena/VASP, to Stu2.

      New refined affinity measurements for a Stu2 TOG domain using Bio-layer interferometry show more than an order of magnitude higher affinity of TOG domains to tubulin compared to previously published reports.

      The findings reinforce the "concentrating reactants" or, more specifically, for TOGdomain proteins, the "tubulin-shuttling antenna" model, compared to the "polarized unfurling" model, a more speculative structural hypothesis.

      The manuscript builds upon a series of previous manuscripts that showcase the profound intellectual engagement with microtubule polymerization mechanisms by TOG-domain proteins from the Rice lab, a thought leader in microtubule polymerization for over a decade.

      Minor remarks:

      (1) A major new experimental finding of this paper is the affinity of TOG domains, which is more than an order of magnitude lower (10 nM) than previous measurements from the same lab (~200 nM). The authors attribute this change to ionic strength differences between buffer conditions, citing the lab's previous work (Ayaz et al., 2014). This argument left me contemplating what the buffer conditions are in both experiments, and I wonder if other readers would feel the same. After going down the rabbit hole, I believe the difference in ionic strength is ~2.3 fold, and at least on the back of my envelope, this works out beautifully with the measured differences in affinities. A short version of this argument may strengthen the manuscript.

      This is a good comment. We should have been clearer about the different buffer conditions. The revised manuscript now explicitly states how the two buffers in question differ in pH and ionic strength. (Page 8, ‘Tubulin binds rapidly …’ section). We tried to perform comparative measurements of TOG:tubulin affinity in the two buffer systems using biolayer interferometry, then analytical ultracentrifugation and isothermal titration calorimetry, but in each technique one or the other buffer caused aggregation, nonspecific binding, or some other artifact that prevented such an analysis. This is stated in the revised manuscript (Page 9, final paragraph before the ‘Unifying measurements …’ section. Along the lines of the reviewer’s ionic strength calculation, and consistent with the increase in affinity we observed with lower ionic strength, we now also state that prior measurements from the Al-Bassam lab (Nithianantham, 2018) showed that TOG:tubulin affinity decreases ~20-fold with higher ionic strength (100 mM KCl vs 200 mM KCl).

      (2) I am wondering if there may be an alternative explanation to tubulin binding by TOG being the kinetically rate-limiting step for polymerase function:

      TOG + Tubulin ⇌ TOG:Tubulin (fast binding rate, high-affinity binding)

      TOG:Tubulin + MT_end → TOG:MT (tubulin is incorporated into MT, fast transfer rate)

      The binding rate is 3/s, and the transfer rate is 5/s.

      I was wondering if the following step should be considered, which involves a conformational change of tubulin (e.g., straightening) TOG:MT → TOG + MT (ratelimiting straightening and unbinding of TOG from the lattice).

      This is an interesting thought that highlights a gap in the understanding of microtubule dynamics.

      Presumably, the affinity of TOGs for straight tubulin is practically zero for the purpose of this discussion, as there is no lattice binding, which means unbinding is likely very rapid; however, straightening may be the rate-limiting factor here.

      In theory, straightening should also be rapid; however, we lack measurements of how fast or slow this step occurs within the context of a TOG domain, which presumably skews the process towards curved tubulin.

      We agree (based on prior observations) that the affinity of these TOGs for straight tubulin is negligeable in this context. There is much less data about the timescale of tubulin straightening, with or without a TOG domain bound, or even about how tubulin interactions with the microtubule end affect the balance of preferred conformations and/or the rate of conformational change. It’s an extremely interesting topic. Because the straightening process the reviewer envisions is zero-order, the transfer rate in our model could in principle reflect slow straightening (in this view the ‘delivery’ step would need to be very fast, i.e. not rate-limiting). Because there is so little data about this, and because there are not yet methods to study or perturb the timescale of straightening on the microtubule, we prefer not to engage too deeply. We added a sentence to acknowledge this alternative possibility in the revised manuscript (bottom of Page 4 and top of Page 5).

      A hypothetical Stu2, when bound to the microtubule end and with the TOG domain not disengaged from tubulin, would not permit the processivity of that molecule or the binding of a new molecule.

      To emphasize the importance of unbinding, when it is not efficient, as reported for the T238 mutant that results in Stu2 lattice binding (Geyer et al., 2018), the polymerase becomes inefficient.

      The mechanism of polymerase processivity has not been conclusively determined (the Geyer et al. 2018 eLife paper took a step in that direction, though). The model used in this paper is only concerned with how many polymerases are at the microtubule end at steady-state (as opposed to how long a particular polymerase acts before dissociating), so while we appreciate and are interested in these questions, we think it would be better to leave them for future work.

      Reviewer #2 (Public review):

      Summary:

      The manuscript from the Rice lab by Gangadharan et al. investigates the polymerization mechanism of the yeast microtubule polymerase Stu2. The lab has published a number of articles demonstrating the structural basis by which the two TOG domains of Stu2 each bind free tubulin heterodimers, and has developed a tethered polymerization model by which the TOG domains drive polymerization by shuttling those tubulin subunits onto the microtubule plus end. A second model was proposed by Nithianantham et al. (eLife, 2018) based on a closed-to-open transitional state in which Stu2 unfurls and loads two longitudinally associated tubulin heterodimers onto the microtubule plus end. While the second model is not directly tested, the current work aims to further characterize/model the tethered polymerization model using a kinetic framework developed by Breitsprecher et al. for Ena/VASP actin polymerization activity, using a model that is enzymatic (EMBO J., 2011). The general architecture and function of Ena/VASP on actin polymerization versus Stu2 on microtubule polymerization is a reasonable relation and hits upon, as the authors note, potential convergent mechanistic evolution across distinct cytoskeletal networks. The model effectively treats tubulin as the substrate, and the polymerized microtubule plus end as the product. If Stu2 is "enzymatic" in this framework, the model predicts it would behave with Michaelis-Menten kinetics, that there would a Vmax, and polymerase activity would either be "affinity limited" by TOG:tubulin affinity (KD) and/or "kinetically limited" by TOG:tubulin association (Kon) and transfer of tubulin to the microtubule plus end (Kt). The authors find that the Brietsprecher model works well for Stu2 activity, and that Stu2 best aligns with a "kinetically limited" model. The work is interesting and adds to the growing elucidation of the Stu2 microtubule polymerase model. While yeast microtubule polymerases are somewhat distinct in their architecture, there is significant overlap that findings from the manuscript can be utilized to inform the mechanisms of larger, more complex microtubule polymerases such as human ch-TOG.

      Thank you for the nice summary and favorable comments.

      Strengths:

      The manuscript invokes the enzymatic model of Breitsprecher et al. used for Ena/VASP and conducts an elegant series of (mostly established) experiments to determine whether Stu2 microtubule polymerase activity aligns with the model, which they conclude does align, supported by the data/results obtained.

      Weaknesses:

      The authors used biolayer interferometry to measure TOG:tubulin affinity. The affinities obtained were significantly higher than the lab obtained in an earlier publication using analytical ultracentrifugation. While differences in buffer and salt conditions may underlie these differences, additional runs using comparable buffer systems, or the use of a third independent assay to measure affinities, would have added rigor.

      This is a good question that was also raised by reviewer #1. We tried hard to perform comparative measurements of TOG2:tubulin affinity in the two buffer systems using biolayer interferometry, then analytical ultracentrifugation and isothermal titration calorimetry, but in each technique one or the other buffer caused aggregation, nonspecific binding, or some other artifact that prevented such an analysis. This is now stated in the revised manuscript (page 9, final paragraph before the ‘Unifying measurements …’ section). We also added text to state that the affinity of TOG:tubulin interactions have been independently shown to depend on ionic strength in a way that seems consistent with what we observed: prior data from the Al-Bassam lab (Nithianantham, 2018) showed that TOG:tubulin affinity decreases ~20-fold with increased ionic strength (100 mM KCl vs 200 mM KCl) (page 9, final paragraph before the ‘Unifying measurements …’ section).

      The discussion could be expanded to better compare and contrast the results with both existing polymerase models introduced in the introduction, as well as expanded to look at reversible enzymatic activity (microtubule depolymerization at low to zero tubulin concentrations) and microtubule plus versus minus end activity.

      Thank you for the push to be more explicit about the two contrasting models. We made small changes to the introduction (top paragraph on page 3) and added a paragraph to the discussion to be clearer about how the existing models are or are not consistent with the present results (page 12, penultimate paragraph of the main text).

      The ‘transfer’ reaction is treated as irreversible (analogous to catalysis by an enzyme), so the biochemical model we use for the polymerase cannot account for polymeraseinduced microtubule depolymerization at low to zero tubulin concentration. We added text to state that the model is limited to the growth reaction (page 4, last paragraph) but otherwise prefer to not engage too deeply in questions about the reverse reaction.

      These polymerases are thought to be plus-end specific because of the domain organization of the protein: TOGs bind tubulin such that the N- to C-terminal polarity of the TOG corresponds to the plus- to minus-end polarity of the tubulin, and the basic region used to make a ‘slippery’ connection to the microtubule is located C-terminal to the TOGs. These two factors mean that it is only at the plus-end that TOGs can engage αβ-tubulins with the basic region contacting surfaces ‘deeper’ in the polymer. We added text about these issues, citing prior work, to the legend of Figure 5 (page 11). We chose to not elaborate much since it is not the primary focus of the paper.

      Reviewer #3 (Public review):

      Summary:

      This study by Gangadharan and colleagues seeks to establish a quantitative biochemical model for the microtubule polymerase activity of Stu2. Stu2 is the budding yeast member of the XMAP215 protein family, which is broadly conserved across eukaryotes. XMAP215 proteins play a wide variety of important roles in cells, and these are attributed to effects on microtubule dynamics. Many studies over the last ~20 years have shown that XMA215 proteins selectively associate with microtubule ends, where they increase rates of microtubule assembly and disassembly. More recently, structural biology and biochemical studies by the authors and other groups have shown that the multiple TOG domains on XMAP215 proteins are tubulin-binding domains that selectively bind to curved tubulin, which is present in solution and at microtubule ends, but not to straight tubulin which is present in the walls of the microtubule lattice. This has led to the general model that XMAP215 proteins promote polymerization by delivering soluble tubulin to the growing plus end, and two distinct models have been proposed to explain the mechanism. The 'concentrating reactants' model proposed previously by the authors suggests that TOG domains grab hold of tubulin in solution and concentrate at the microtubule end. The 'polarized unfurling' model proposed by the Al Bassam lab suggests that XMAP215 delivers multiple tubulins to the end, using a step-wise mechanism involving different roles for each TOG domain. The current study seeks to improve our understanding of the mechanism by developing a quantitative model to explain the binding and release of tubulins, the number of Stu2 molecules at the end, and the overall rate of tubulin addition. The authors accomplish this goal using new experimental data. The final model fills in new details of the mechanism. The authors draw a comparison between Stu2 and the actin polymerase, which bears similarity to Ena/VASP, and suggest a convergent strategy for cytoskeletal polymerases.

      Thank you for the good summary and favorable comments.

      Strengths:

      This is a focused and clearly written study that incorporates prior knowledge of XMAP215 and draws inspiration from the actin field. The data are clear and convincing, and the study accomplishes its goal of generating a new, quantitative model for Stu2. The model will be important for microtubule researchers to predict and test key points for altering XMAP215 activity across different organisms and potentially for different tubulin substrates. The comparison to Ena/VASP may also inspire similar comparisons across other microtubule and actin regulators, which could lead to new insights across the cytoskeletal fields.

      Thank you for these comments.

      Weaknesses:

      The study is without major weaknesses, but there are several minor weaknesses worth noting. One is that the final model provides new details regarding the Stu2 mechanism, but does not provide a major new advance in our understanding of how the polymerase works. For example, the discussion does not clearly argue for whether the new results and model rule out either of the prior models. This appears consistent with the 'concentrating reactants' model, but does it clearly rule out the 'polarized unfurling' model?

      Thank you for pointing out what in retrospect was an obvious ‘loose end’ in our discussion. The other referees raised the same point. We made small changes to the introduction and added a paragraph to the discussion to be clearer about how the existing models are or are not consistent with the present results (top paragraph of page 3 and new penultimate paragraph of the manuscript on page 12).

      A second minor weakness is that the comparison to Ena/VASP is not developed at a deep level based on the final model. I found these ideas exciting and want more critical consideration here, but perhaps it is better suited for a commentary piece to follow.

      We appreciate the enthusiasm and understand where this comment is coming from. Because there has been a fair amount of recent movement in the understanding of TOG domains and what they can do, and because some of the mechanistically interesting parallels entail speculation, we agree with the referee’s suggestion that a future commentary will provide a better venue.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Other minor remarks:

      (1) Figure 2 C is missing the label for what should probably be TOG1*-TOG2.

      Fixed

      (2) Figure 5, lower left, is oddly cropped, showing residuals that are slightly distracting from the beauty of the model.

      Apologies that the figure did not look good in the initial submission. We adjusted it and it looks much better in the revised submission.

      (3) The dynamics assay buffer composition stated in the protein purification Method section is not the same as the PEM buffer used for the dynamics assay. And both are different from BRB80, which, with the chambers, are rinsed. This may very well be accurate, but it raises the question of why not stick to one version, as they are virtually the same.

      Thanks for asking these questions, and sorry for the confusion. First, we should have used different names for the (barely) different buffers. This has now been corrected. Second, the PIPES concentration was not 90 mM, it was 100 mM as in our prior work and this discrepancy failed to get caught in proofreading. Why the other small differences? It’s a good question. The differences reflect an arbitrary decision made at the beginning of the work, there is not a deeper rationale.

      (4) Out of curiosity: Why 90 mM PIPES and not 80?

      Why not 80 mM PIPES? This is just a historical difference. The early measurements of yeast microtubule dynamics (e.g. Gupta … Himes MBoC 2002 and Bode … Himes, EMBO Rep 2003) that partly inspired us to use yeast as a model system used 100 mM PIPES as the working concentration, and we never deviated from that.

      (5) Please state the source of PIPES.

      Sorry for the oversight, we have added the source of PIPES (it is Millipore Sigma P6757).

      Reviewer #2 (Recommendations for the authors):

      (1) Page 4, last paragraph, the authors call out "Fig 1C" which I believe should be "Fig 1D".

      We fixed this, thank for catching this error

      (2) Figure 2C: The authors subtract basal tubulin polymerization (growth rate) from the rates measured in the presence of Stu2 constructs. One assumption in doing this is that Stu2 polymerization activity does not compete for the ability of tubulin (not bound to Stu2) to polymerize on the plus end. I think this is a logical assumption, but it would be beneficial for the authors to state this assumption.

      We said this explicitly in the revised submission (first full paragraph on page 7), borrowing from the reviewer’s phrasing.

      (3) Figure 2C: the label for the last bar is missing - presumably: " Stu2 (TOG1*-TOG2)".

      Fixed.

      (4) Figure 2D: Many of the KM values determined are at the border for points measured, or in one case, beyond the concentration of tubulin sampled. In this regard, the authors should discuss how well the fitted curves correlated with their data. Also, as Vmax and KM are calculated, it would be beneficial if another panel were produced (e.g., Figure 2E) in which the data were presented as a Lineweaver-Burk plot. Doing so, the authors would be able to test their enzymatic model by doping the system with their nonpolymerizable tubulin mutants, which should yield competitive inhibitor behavior but not change Vmax.

      Thanks for pointing this out, we should have been more explicit about this point. We incorporated into the results section an explicit statement about this limitation (first full paragraph on page 7). The suggestion to use blocked mutants and Lineweaver-Burke plots is an interesting one that we hope to pursue in future work using blocked or other mechanism-specific mutants. But we think to do so is complicated enough to be beyond the scope of the present study.

      (5) Figure 3: The authors quantitate the amount of Stu2-GFP fluorescence at microtubule plus ends using line scan analysis of the kymograph. Since the kymographs are processed images, it is more appropriate to integrate intensity from the original frames collected using a circular area. i.e., ID points on the kymograph, and return to the respective position in the corresponding frame to calculate background-subtracted GFP intensity at the plus end.

      This is a fair point. We chose to stick with the kymographbased analysis because the symmetry of the point-spread function and lack of rapid variation in GFP and/or background intensity means that the kymograph analysis is adequate for the intensity-based comparison we were doing.

      (6) Page 7, last line, the authors call out "Fig 2D" which I believe should be "Fig 1D".

      Sorry for the error, we have corrected it.

      (7) Figure 4A: The authors discuss the "sortase epitope," but technically, an epitope is the binding site specifically for an antibody, not to be used in general terms for proteinprotein interaction sites. As such, the authors should describe this as the "sortase recognition sequence" or something similar.

      Thank you for noticing this; we had indeed used ‘sortase recognition sequence’ elsewhere in the paper but we did not catch this instance during proofreading. We have now used that same language in the legend for Fig. 4.

      (8) Page 8, the authors state "KM is approximately equal to Kt/Kon and KM negligeable," but I think they mean "...and KD negligeable".

      Thank you for noticing this typo. We corrected it.

      (9) Page 9, first paragraph last sentence: the readership would be aided by modifying the sentence as follows (adding "Kon" and "Kt"): " ... must be kinetically limited by either the rate of TOG:tubulin binding (Kon) or by the rate of TOG-mediated transfer to tubulin to the microtubule end (kt)."

      Very good suggestion, we implemented it.

      (10) Page 9, second paragraph, the authors call out "Fig 2D", but perhaps they intended to call out "Fig 1D"?

      Sorry for the error, the reviewer is correct and we fixed this.

      (11) Page 9, second paragraph: The authors mention that the transfer rate of tubulin to the plus end via a TOG domain is close to the transfer rate of free tubulin to the growing plus end. Can the authors expand on why they are mentioning this comparison?

      Thanks for the push to be clearer about this. The basic idea is that each ‘delivering’ TOG contributes 50% of the background (uncatalyzed) polymerization rate. So the presence of multiple TOGs (in a single polymerase or from multiple end-resident polymerases) can substantially increase the rate of polymerization. We added brief text to try to make this clearer (first paragraph on page 10).

      (12) Page 9, second paragraph: "TOG-TOG2 polymerases" would be better phrased as "TOG2-TOG2 dimeric polymerases". Noting as well that "2" is missing from the first "TOG".

      This is indeed better phrasing and we have adopted it (also corrected the missing ‘2’) (first paragraph on page 10).

      (13) Page 13, BLI methods: The authors should list the final pH for the PIPES buffer (was it pH 6.9 as in the polymerization assay?).

      Sorry for the oversight, we have added the pH and it was indeed 6.9.

      (14) Page 13, BLI methods: What is "LR1-457"?

      LR1-457 is lab-notebook-speak that did not get purged in editing; it refers to the polymerization-blocked tubulin mutant that also carried a sortase recognition sequence. We replaced ‘LR1-457’ with more evocative phrasing and corrected another typo we found there.

      (15) In Ayaz et al. (eLife, 2014) Stu2 TOG1 and TOG2 affinities for tubulin were measured using AUC, for which the fitted curves appeared to correlate with the data quite well. As the authors note, the values were KD = 70 nM and 160 nM, respectively. This contrasts with the BLI measurement for TOG2-tubulin (~10 nM), which suggests that at least one of the experiments was off the mark - or, as the authors do note, that different buffer and salt condition was used could account for the differences, but that the BLI conditions align with the polymerization conditions (though not exactly) and thus are more appropriate to use. In a supplemental discussion, the authors should run the AUC values through the equation for their model and state what types of differences these values could imply for Stu2 mechanism. If the differences are significant for the Stu2 model derived, the authors should give thought as to whether a third assay should be employed to determine TOG-tubulin affinity. Based on the BLI reagents, it appears the authors would be well-positioned to conduct an assay using SPR. As a potential alternative, the authors could repeat the BLI experiment using the buffer conditions from the Ayaz et al., AUC work (25 mM Tris pH 7.5, 1 mM MgCl2, 1 mM EGTA, 100 mM NaCl, 20 μM GTP) - noting that BSA and Triton X-100 may need to be added as well. If the authors are able to replicate the ~160 nM affinity for TOG2:tubulin, this would be a reasonable way to bootstrap to the conclusion that the BLI is measuring affinity correctly and that the current PIPES-based BLI experiments yielded accurate data.

      We tried hard to perform comparative measurements of TOG2:tubulin affinity in the two buffer systems. Unexpected challenges and personnel turnover made this slower than anticipated. The reviewer’s suggestions are completely reasonable, but ultimately it was not possible for us to get side-by-side results for TOG:tubulin affinity using the same measurement technique, whether it was biolayer interferometry, analytical ultracentrifugation, or isothermal titration calorimetry. For each technique one or the other buffer caused aggregation, non-specific binding, or some other artifact that prevented analysis. The fact that we were unable to compare the buffer conditions in this way is stated in the revised manuscript (page 9, last paragraph before the ‘Unifying measurements …’ section). We also added text to state that the affinity of TOG:tubulin interactions have been shown to depend on ionic strength in a way that is consistent with the changes we observed: prior data from the Al-Bassam lab (Nithianantham, 2018) showed that TOG:tubulin affinity decreases ~20-fold with increased ionic strength (100 mM KCl vs 200 mM KCl) (page 9, last paragraph before the ‘Unifying measurements …’ section). We also added some text to address the comment about affinity and whether/when the shuttle model would hold (first paragraph on page 11).

      (16) A sentence or two in the discussion, relating how their data aligns (or not) with the Nithiantham model would be beneficial, especially as discussing the two models in the introduction was a central point.

      We completely agree and have now added a paragraph to the discussion to explicitly address the two models and how are or are not supported by the new model and observations (page 12, penultimate paragraph of the main text).

      (17) Discussion: The model in Figure 5 depicts Stu2 engaged with the microtubule, perhaps using its basic linker region (?). The authors could note this in the figure caption for 5A. The authors do not discuss the basis for plus-end polymerization activity versus polymerization activity at both the plus and minus ends. Do the authors propose that this is due to differential Kt values for the two ends and/or differential localization via the basic region to the two ends?

      Thanks for bring this up. The plus-end selectivity of these polymerases is thought to result from the polarity of TOG:tubulin engagement and the positioning of TOG domains relative to the basic region that provides ‘slippery’ binding to the microtubule lattice. We have partially addressed these issues in the legend to Figure 5 (page 11).

      (18) Brouhard (Cell, 2008) demonstrated that XMAP215 can catalyze the depolymerization of GMPCPP microtubules when no free tubulin, or very low levels of free tubulin, are present. This is interesting in that it indicates that the enzymatic activity is reversible. Can the authors comment on how their model would behave in the low-tozero free tubulin concentration regime? Would a different model have to be invoked?

      This is an interesting comment. Because the enzyme-like model treats the transfer step as irreversible, the model cannot account for the kind of ‘depolymerase’ activity Brouhard and others have noted. A more general model that could also encompass the depolymerase activity at low-to-no free tubulin would need to explicitly model the step(s) involved in microtubule association and dissociation. These steps remain a major open question in the field and while it would be quite interesting, trying to address this in a model is beyond the scope of what we can confidently do given the data we have. To be more explicit about this assumption/limitation, we now point this out in the results section where the model is introduced (page 4, last paragraph).

      Reviewer #3 (Recommendations for the authors):

      (1) Figure 1D, legend. "...and a transfer rate constant kf that describes how fast...". Should kf be replaced with kt?

      We made this correction, thanks for pointing the problem out

      (2) Figure 2C. The x-axis label under the blue bar is missing. Also, I find the arrows to the left of the bars confusing and unnecessary.

      We fixed the legend problem. We sympathize with the dislike of the arrows but respectfully prefer to keep them in the hopes of avoiding confusion about the fact we are fitting ‘growth rate attributable to Stu2’, not simply growth rate. The figure legend has been expanded to hopefully smooth this over.

      (3) Figure 3. The kymographs are convincing, but it may be helpful for future studies to state here what the polymerization rates are for 0.6 µM and 1.4 µM yeast tubulin. These values are probably different than what one might expect for mammalian tubulin at those concentrations, and the authors could simply determine them from the slopes in the kymographs.

      Good suggestion, we added the growth rates to the legend (as the reviewer expected, they differ from expectations based on mammalian tubulin).

      (4) Results, page 9, line 11: "...yielded a value of 9.6 nM...". Should this be 8.9 nM, which is that value stated in Figure 4C?

      Actually, these different values are correct. We just wanted to point out that whether we used response amplitudes or measured on- and off-rates, we get very similar values for K<sub>D</sub>. We changed wording to hopefully make this clearer: “Calculating the dissociation constant K<sub>D</sub> from the measured rate constants (K<sub>D</sub> = k<sub>off</sub>/k<sub>on</sub>) instead of from the amplitudes yielded a value of 9.6 nM, in good agreement with the amplitude-based determination of 8.9 nM.” (page 9).

    1. eLife Assessment

      This valuable study explores changes in the Drosophila microbiome in response to environmental temperature over more than ten years. The evidence that temperature leads to diversification of bacterial clades is solid, despite the need for greater clarity in defining and tracking strain competition. The work will interest researchers working with microbiomes, microbial ecology, and evolutionary biology.

    2. Reviewer #1 (Public review):

      Summary:

      The factors that create and maintain diversity in host-associated microbiomes remain poorly understood. A better understanding of these factors will help in the efforts to leverage the adaptive potential of the microbiome to help solve pressing problems in health and agriculture.

      Experimental evolution provides a promising path forward as we can track the causes and consequences in the emergence of novel variants, but experimental evolution remains underutilized in host-microbiome interactions. Here, Gracia-Alvira utilizes a long-term experimental evolution study in Drosophila simulans under hot and cold temperature regimes to identify strain-level variation in an important fly bacterium, Lactiplantibacillus plantarum. They identify three strains of L. plantarum, which are most prevalent in their respective three temperature regimes, suggesting that these are locally adapted bacteria. Then, using a combination of genomics, in vitro, and in vivo, Gracia-Alvira et al attempt to understand the factors that led to the differentiation of the hot and cold L. plantarum and their impacts on the fly host.

      Strengths:

      This is an excellent use of experimental evolution to track the emergence of novelty in the microbiome. The genomic analyses are all solid and appropriate for the data sets. It is especially striking that the comparisons with the other, independent experimental evolution studies in different labs (and across continents between Portugal and South Africa) show a consistent response to temperature. Many have disregarded the microbiome as it is something that is too sensitive to seemingly innocuous variables (particularly in the fly microbiome), such that we cannot find generalities. However, this finding highlights the potential for experimental evolution to uncover these dynamics. The question of how strains emerge and are maintained is timely and is one of the key open questions in host-microbiome evolution currently.

      Comments on revised version:

      I thank the authors for their thoughtful responses to my concerns, and I appreciate the additional experiments to help resolve the questions about subspecies competition. The manuscript remains strongest in the genomic assessment of changes in the L. plantarum genomes, and it is striking and noteworthy that the isolates across multiple countries but same temperature conditions group together phylogenetically.

      I appreciate the additional clarity also incorporated in this revision, but there are still a few key concerns that are unresolved about the microbial ecology described here. I will also note that I apologize if I missed something in the text as no line numbers were provided to point me to where the changes were incorporated in the revised manuscript.

      (1) Competition has many different meanings and many different measurements (see Hart 2018 https://doi.org/10.1111/1365-2745.12954) -and incorporating the effects of competition in shaping an ecological community is, has been, and will continue to drive much research in community ecology. Measuring strain level competition is one of the major questions in host-associated microbiomes, and it is difficult-though there have been significant advances in doing so (see isogenic barcodes, e.g., Daniel 2024 doi: https://doi.org/10.1038/s41564-024-01634-9b, Ordon 2024 https://doi.org/10.1038/s41564-024-01619-8, as well as my previous suggestion to track the outcomes of competition). The inability to directly track and measure competition of the isolates remains a limitation of this manuscript. The authors' explanation of measuring competition is unusual, simplistic, and at times inconsistent.

      They need to be crystal clear about their definitions, logic for making these inferences, and weaknesses in their approach. I think what the authors mean is that competition between the unevolved and C or H in their respective regimes leads to the decrease of the U clade over experimental evolution. But it is not clear how the authors are thinking about competition between C and H clades in the different temperatures.

      The authors state that competition is inferred because changes in relative abundance across the time series-and this is unusual because there are alternative explanations that require no ecological interactions among sub-strains, as I described in my comments on the prior version. This is then combined with in vitro work that shows that the H and C clades can both grow in their mismatched temperature regimes-and thus I think it is to be inferred that because they can grow alone in vitro (and C isolates show lower growth than H isolates in hot temperature), then changes in the relative abundance over fly generations can be attributed to competitive interactions among C and H clades. But then the logic is inconsistent because then the authors just say that in vitro growth curves don't support the differences in relative abundance observed in the flies (lines 224-225). Then the authors argue is it about a combination of diet/sugar metabolism and temperature (line 373), which doesn't make any sense because temperature previously didn't matter (lines 224-225).

      All of this is to say is that the authors need to make clear their logic to the readers-and explain these inconsistencies appropriately. To me, it suggests that there are clear methodological weaknesses that inhibit the ability to track competitive microbial dynamics. Because you can't really assess the microbial dynamics in vivo, it remains further unresolved why clade C isolates have such strong negative fitness effects on the fly but reach such high relative abundances in the C evolving flies. I find that this series of logical inconsistencies (and see my point #2) distracts from the important finding that the C and H clades evolved to utilize sugars differently from the U clade, which is an interesting finding!

      (2) There are also inconsistencies in the patterns observed between the text and the figures. Some of this arises because the authors are not clear what comparisons they are making. For example, line 450 says that clade C outcompeted the other clades, which I presume means only in the cold temperature. Line 456 says that C and H isolates grow faster in the sugar-rich lab diet, but that is not really true because U and C have similar growth rates in Fig. 5, and U and H have similar growth rates in Fig. S4. The text about microbial load is a bit misleading (lines 271-273), as it is confusing that clade C is significantly higher load in both hot and cold temperatures (Fig. S6), which is counterintuitive given Fig. 4, 5, S4. But it is also overly speculative to say that these results suggest that fitness effects depend on microbial load of clade C without connecting the load to the fly fitness measures (and also given the inconsistency with the time series data from evolving lines). Please take care to more carefully phrase these statements to ensure the inference is supported by the experiment design (e.g., clarifying comparison) and statistics (e.g., ensuring agreement with what the figure shows).

      (3) I understand the concern about focusing the reader on the L. plantarum strains. However, it should be clear to the readers that you did not examine the other parts of the microbiome, and that L. plantarum is often very rare in lab and wild fly populations. The data presented on Table S4 (cited line 552, I think citation at line 176 is incorrect) is confusing. If these were colonies picked and then identified, this should be explicit. If it is based off on colonies, then please clarify if this was sampled randomly or occurred when trying to enrich/focus on L. plantarum isolates. If the data was computational (e.g., Kraken to classify), then only taxa richness is not necessarily relevant, but please also include to the relative abundance of each taxa.

      To me, this is relevant information to contextualize these results, particularly because you test this in both D. mel and D. simulans (apologies for the confusion over Mazzucco & Schlotterer 2021), and we have insight into how combinations of Lactobacillus and other taxa impact fitness (Gould PNAS 2018). If the results from D. melanogaster are not applicable to D. simulans, then the authors need to explain this. I understand if incorporating analysis of the broader microbiome is beyond the scope of this manuscript, but at least acknowledging the general rarity in Lactobacillus frequency in Drosophila microbiome and variation in fitness effects will more accurately contextualization these results.

      One small point is that line 452 the citations are OK, but there are fly-specific examples to support this statement, like Gould PNAS 2018, Henry Proceedings B 2025.

    3. Reviewer #2 (Public review):

      Summary:

      In this manuscript, Gracia-Alvira et al. investigated how environmental temperature affects competition among members of the microbiome, with a focus on intraspecific diversity, using the Drosophila model.

      Notably, the authors identified three clades of Lactiplantibacillus plantarum from a natural population of Drosophila simulans collected in Florida. They tracked the dynamics of these three bacterial clades under two temperature conditions over the course of more than ten years. Using comparative genomics and phylogeny, they showed that these three bacterial clades likely adapted to their host independently in a temperature-specific manner. Further, by combining in vitro culture and in vivo mono-association assays, they demonstrated the functional divergence of these three bacterial clades phenotypically, including their growth dynamics and effects on host fitness. Lastly, they performed pathway analysis and speculated on key genomic variance supporting such functional divergence.

      Strengths:

      The laboratory evolutionary experiment in response to cold or hot environmental temperature is impressive, given its more than ten years of experimental time period. This collection of achieved microbiome samples paired with the fly host data can be a valuable resource for the field.

      Comments on revised version:

      The revised version has addressed my major points raised in the original review.

    4. Reviewer #3 (Public review):

      Summary:

      The study presents an analysis of 297 pangenomes derived from 20 populations of Drosophila simulans, at 19 time points for fast-reproducing individuals in a hot environment, or at 10 time points for slow-reproducing individuals in a cold environment, over a period of more than 10 years. The authors select a particular microbial component of the pangenomes and study the dynamics of Lactiplantibacillus plantarum strains in two environments. They discover that the revealed operational taxonomic units could be divided into three phylogenetic clades, which have their own genomic and genetic features, different adaptive capabilities that depend on the environment, and have a distinct impact on the fitness of the host.

      Strengths:

      The authors prove that bacterial microbiome components are sensitive to the environment and could rapidly (years) be fixed in eukaryotic populations. This study establishes a tractable model that potentially enables the study of variability of the physiological influence of distinct strains of an important commensal species, Lactiplantibacillus plantarum, on the Drosophila host. It is clearly shown that this single species consists of several phylogenetically and functionally diverse strains. The authors did not limit their interest to their own model, but rather they have integrated a comparative approach by analysing phylogenetic relationships among 92 described L. plantarum strains.

      Overall, the study is novel and delivers important discoveries of a longitudinal, well-replicated experiment, generating a substantial amount of genomic data. It highlights an important dimension of research that environmental selection operates at the subspecies level.

      Weaknesses:

      Even though the authors show only one particular example by conducting their longitudinal experiment, they honestly acknowledge failures important for interpretation of the biological significance of the results (gnotobiotic mono-association experiments was done with D. melanogaster, but not D. simulans) and therefore they state limitations of their conclusions (weaker effects in the non-axenic flies are due to the presence of other taxa or to higher-order interactions with other members of the microbiome). These interactions could significantly affect bacterial growth, metabolism, and physiological influence on the host.

      The authors exploit the results of their experiment to speculate about a wide range of evolutionary phenomena, like within-species competition, ecological adaptation and evolution of the host, fitness advantage of bacteria to the host, the benefits of parasitism or mutualism, the domestication of the microbiome, etc. At the end, they conclude that their study "highlights that even subspecies diversity plays a key role in adaptation to environmental temperature". However, the potential mechanisms of such adaptation are barely discussed, so that the focus of the study shifts from the temperature-induced changes in microbial population structures toward metabolism-related adaptations of clade representatives that enable them to diversify their carbon and nitrogen sources. The role of the temperature factor remains elusive.

      In addition to that, the paper has a clearly minimalistic experimental approach to address functional properties of the revealed L. plantarum strains, so that their own fitness, or their relationship with the Drosophila host, is characterised superficially. Therefore, the authors' discourse can be speculative rather than factual (especially when the authors use the expression "likely" to share their guesses in the "Results" section). Nevertheless, these minor drawbacks do not underscore the novelty of the discovered phenotypes and the importance of their further investigation.

      Comments on revised version:

      I have read the authors revisions and find them compelling and they address fully the minor points raised in my review.

    5. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The factors that create and maintain diversity in host-associated microbiomes remain poorly understood. A better understanding of these factors will help in the efforts to leverage the adaptive potential of the microbiome to help solve pressing problems in health and agriculture.

      Experimental evolution provides a promising path forward as we can track the causes and consequences in the emergence of novel variants, but experimental evolution remains underutilized in host-microbiome interactions. Here, Gracia-Alvira utilizes a long-term experimental evolution study in Drosophila simulans under hot and cold temperature regimes to identify strain-level variation in an important fly bacterium, Lactiplantibacillus plantarum. They identify three strains of L. plantarum, which are most prevalent in their respective three temperature regimes, suggesting that these are locally adapted bacteria. Then, using a combination of genomics, in vitro, and in vivo, Gracia-Alvira et al attempt to understand the factors that led to the differentiation of the hot and cold L. plantarum and their impacts on the fly host.

      Strengths:

      This is an excellent use of experimental evolution to track the emergence of novelty in the microbiome. The genomic analyses are all solid and appropriate for the data sets. It is especially striking that the comparisons with the other, independent experimental evolution studies in different labs (and across continents between Portugal and South Africa) show a consistent response to temperature. Many have disregarded the microbiome as it is something that is too sensitive to seemingly innocuous variables (particularly in the fly microbiome), such that we cannot find generalities. However, this finding highlights the potential for experimental evolution to uncover these dynamics. The question of how strains emerge and are maintained is timely and is one of the key open questions in host-microbiome evolution currently.

      Weaknesses:

      (1) The framing in the title and throughout the discussion about "subspecies competition" does not match the data that was collected. The subspecies competition requires actually tracking the competitive outcomes between the hot, cold, and unevolved L. plantarum. In the in vivo work, I can see that mixes of the strains were made, but they did not track whether the cold strain outcompeted the hot strain in vivo under cold conditions, for example.

      We thank the reviewer for the honest concern and take this opportunity to defend our claim of "subspecies competition used across the manuscript. As the reviewer states, subspecies competition requires tracking the competitive outcomes between the three clades, and this is what we did by sampling and sequencing across ten years of experimental evolution (Figures 4 and S3). For this reason, we point that the subspecies competition assessment comes from the direct observation of changes in relative abundance across the time series, and not from the follow-up experiments in vivo or in vitro.

      While Figure 4 is suggestive that there is ongoing competition in the hot temperature regime, this is not necessarily shown in the cold, which is dominated by the C clade. It could also be that the bacteria cannot survive in the flies at the different temperatures. The growth curve assays hint that the bacteria can grow, but the plate reader couldn't actually maintain the 18 {degree sign}C temperature (line 455). So all of this evidence is very indirect and insufficient to say that strain competition is driving these patterns.

      We thank the reviewer for the alternative hypothesis that could explain the observed subspecies dynamic. We rule out that dominance of clade C in the cold occurs because the other two clades cannot grow in this regime based on three pieces of evidence:

      (1) In the time series, clades H and U decrease, but never disappear (Figures 4 and S3), even showing some peaks of abundance in specific replicate populations (Figure S3).

      (2) We isolated individuals belonging to clade H in the cold-evolved populations, as shown in figure 2. This is a direct evidence that clade H prevails in the cold-evolved populations, although in low abundance.

      (3) We did grow the three taxa in fly food Petri dishes incubated at both temperature regimes, observing growth in all cases.

      We will include the food growth experiment in the revised manuscript as further supporting evidence for growth in both regimes.

      (2) The in vivo results are interesting in that there appears to be a fitness cost of clade C, but the explanation is underdeveloped. I say under-developed because in Figure 4, the cold L. plantarum remains much higher throughout adaptation to the hot temperature regime than the hot L. plantarum in the cold regime. The hot L. plantarum is low abundance throughout the cold regime. I felt like this observation was not explained, but it seems relevant to understanding the strain dynamics.

      We acknowledge that a strong fitness cost of clade C is observed in axenic D. melanogaster. In the native host, D. simulans, with reduced microbiome, we observed delayed development that could even be an advantage depending on the situation, as pointed out by reviewer 3 in the recommendations.

      Even if we assume that flies colonized with clade C are less fit in the experimental evolution, another caveat is whether the flies can actively select for the L. plantarum clade. Under this assumption, a clade that imposes a fitness cost to the fly (clade C) should be selected against over time because the flies colonized by this clade will have less offspring or develop later than the rest. Alternatively, as the microbiome is shared among all the individuals in the population, the host might not be able to “purge” the pernicious clade, and L. plantarum dynamics might be controlled solely by the relative fitness between clades in the given experimental treatment. We will discuss this hypothesis in the revision as a way to explain the relationship between the abundance of each clade and the effect on the host.

      I will also note that this is not the first time that L. plantarum or other Lactobacillus have been shown to exert fitness costs to Drosophila. Gould, PNAS, 2018, shows that both Lactobacillus plantarum and Lactobacillus brevis in mono-association have lower fitness (measured through Leslie matrix projections using lifespan and fecundity) than axenic flies. Many studies of wild Drosophila fail to find Lactobacillus, or it is low abundance (e.g., Chandler, PLoS Genetics, 2014; Wang, Environmental Microbiology Reports, 2018; Henry & Ayroles, Molecular Ecology, 2022; Gale, AEM, 2025). This might help provide useful context for the in vivo results.

      We thank the reviewer for the references. These observations are compared to our phenotypic results and discussed in the revised version of the manuscript.

      (3) The data in Figure 4 are compelling to focus on the L. plantarum variants. However, I can see from the methods that the competitive mapping included only other strains of Wolbachia.

      We appreciate the thorough reading of the methods by the reviewer. The competitive mapping comprised two steps: first we discarded the reads that mapped to Drosophila, Wolbachia and additional potential contaminants from sequencing facitilies (human, dog...). This step leaves the reads originated from whole the external microbiome of the flies, including L. plantarum. The second competitive mapping step recruits the reads that map any clade of L. plantarum.

      It is not clear how other members of the microbiome changed in response to the temperature regimes. As I note in point #2, given that Lactobacillus is often rare, it is not clear what the rest of the microbiome looks like over the course of adaptation. Indeed, it seems like Mazzucco & Schlotterer, PRSB, 2021 did a broader analysis of the microbiome and found that Acetobacter is by far the most common bacterium (I think this data is also part of the data shown here?). Expanding on why or why not in this context is important and will improve this study, particularly if the focus is on connecting these evolutionary dynamics to ecological competition to explain the emergence of strain diversity.

      We acknowledge that the rest of the Drosophila microbiome is not addressed in this study, as we wanted to focus the storyline around the intraspecific dynamics found in L. plantarum. We consider that a complete characterization of the whole Drosophila microbiome would unnecessarily elongate the paper and thus we treat it as a constant biotic factor.

      We must point out that our dataset is not the one reported by Mazzucco & Schlötterer, which was done in D. melanogaster, rather than D. simulans. Nevertheless, both experiments share the same infrastructure, temperature regimes and fly maintenance.

      We have included a list of taxa that were isolated from the populations, as well as to report L. plantarum prevalence and abundance across the experiment in order to provide context of the microbiome, beyond L. plantarum, to the readership.

      Reviewer #2 (Public review):

      Summary:

      In this manuscript, Gracia-Alvira et al. investigated how environmental temperature affects competition among members of the microbiome, with a focus on intraspecific diversity, using the Drosophila model.

      Notably, the authors identified three clades of Lactiplantibacillus plantarum from a natural population of Drosophila simulans collected in Florida. They tracked the dynamics of these three bacterial clades under two temperature conditions over the course of more than ten years. Using comparative genomics and phylogeny, they showed that these three bacterial clades likely adapted to their host independently in a temperature-specific manner. Further, by combining in vitro culture and in vivo mono-association assays, they demonstrated the functional divergence of these three bacterial clades phenotypically, including their growth dynamics and effects on host fitness. Lastly, they performed pathway analysis and speculated on key genomic variance supporting such functional divergence.

      Strengths:

      The laboratory evolutionary experiment in response to cold or hot environmental temperature is impressive, given its more than ten years of experimental time period. This collection of achieved microbiome samples paired with the fly host data can be a valuable resource for the field.

      Weaknesses:

      The laboratory evolutionary experiment can be limited due to its artificial experimental setup. For example, wild flies rely on a more diverse set of food sources and are constantly exposed to new bacterial inoculations, whereas under laboratory conditions, flies live in a more restricted ecosystem. In addition, environmental temperatures differ among different locations, but they also involve seasonal changes within the same region. This manuscript can be strengthened with further discussions that elaborate on these limitations.

      As the reviewer has correctly noted, our experimental setting is not exempt from limitations. Lab-reared flies are fed with a defined standard diet. Furthermore, although the system is not completely closed to bacterial migration, this is limited as replicate populations are not allowed to mix during the maintenance of the flies. For this reason, we consider our laboratory setting as a compromise between observing wild populations, which undergo all biotic and abiotic stresses but cannot be manipulated, and evolving the bacteria in absence of the host, or in gnobiotic hosts, in which biotic interactions are not fully considered. We will extend on this in the new version of the manuscript.

      Moreover, the extent of host effects involved in these experiments remains ambiguous, because it is unclear whether these Lactiplantibacillus plantarum mostly reside within fly guts or on Drosophila medium. The laboratory evolutionary experiment possibly favored better colonizers on Drosophila medium under either cold or hot temperatures, which subsequently can saturate fly guts. As fully dissociating these variables can be experimentally tedious, the authors may want to comment more on these aspects in the discussion. Or they may want to consider some measurements. For example, measuring the growth rate of these bacteria on Drosophila medium under different temperatures, in addition to the current MRS culture experiments, or measuring the portion of the Lactiplantibacillus on Drosophila medium versus these stably colonizing fly guts.

      The reviewer's point was briefly addressed in the Results chapter: "Phenotypic differences in liquid culture".

      Reviewer #3 (Public review):

      Summary:

      The study presents an analysis of 297 pangenomes derived from 20 populations of Drosophila simulans, at 19 time points for fast-reproducing individuals in a hot environment, or at 10 time points for slow-reproducing individuals in a cold environment, over a period of more than 10 years. The authors select a particular microbial component of the pangenomes and study the dynamics of Lactiplantibacillus plantarum strains in two environments. They discover that the revealed operational taxonomic units could be divided into three phylogenetic clades, which have their own genomic and genetic features, different adaptive capabilities that depend on the environment, and have a distinct impact on the fitness of the host.

      Strengths:

      The authors prove that bacterial microbiome components are sensitive to the environment and could rapidly (years) be fixed in eukaryotic populations. This study establishes a tractable model that potentially enables the study of variability of the physiological influence of distinct strains of an important commensal species, Lactiplantibacillus plantarum, on the Drosophila host. It is clearly shown that this single species consists of several phylogenetically and functionally diverse strains. The authors did not limit their interest to their own model, but rather they have integrated a comparative approach by analysing phylogenetic relationships among 92 described L. plantarum strains.

      Overall, the study is novel and delivers important discoveries of a longitudinal, well replicated experiment, generating a substantial amount of genomic data. It highlights an important dimension of research that environmental selection operates at the subspecies level.

      Weaknesses:

      Even though the authors show only one particular example by conducting their longitudinal experiment, they honestly acknowledge failures important for interpretation of the biological significance of the results (gnotobiotic mono-association experiments was done with D. melanogaster, but not D. simulans) and therefore they state limitations of their conclusions (weaker effects in the non-axenic flies are due to the presence of other taxa or to higher-order interactions with other members of the microbiome). These interactions could significantly affect bacterial growth, metabolism, and physiological influence on the host.

      We agree with the reviewer in that the use gnobiotic animals is a limitation, as by "tuning" the flies' microbiome we are modifying the interactions between members, which can potentially change the phenotypic outcome. Nevertheless, we use it as a complementary approach, rather than the only inference in our study.

      The authors exploit the results of their experiment to speculate about a wide range of evolutionary phenomena, like within-species competition, ecological adaptation and evolution of the host, fitness advantage of bacteria to the host, the benefits of parasitism or mutualism, the domestication of the microbiome, etc. At the end, they conclude that their study "highlights that even subspecies diversity plays a key role in adaptation to environmental temperature". However, the potential mechanisms of such adaptation are barely discussed, so that the focus of the study shifts from the temperature-induced changes in microbial population structures toward metabolism-related adaptations of clade representatives that enable them to diversify their carbon and nitrogen sources. The role of the temperature factor remains elusive.

      We acknowledge that our study does not fully resolve the mechanism by which a different clade ends up dominating each temperature regime. The MRS liquid experiment was an attempt to answer whether differences in optimal growth temperature could explain the temperature-specific abundance of the two clades. Our experiments showed, however, that this was not the case. Beyond this point, it is hard to disentangle the role of the temperature, as it could also act indirectly on the bacteria, for example, through the host or the food.

      A second observation in our time series was that a third clade, U, was unfit in both regimes despite starting the experiment in high abundance. For this reason we also studied what made this clade less fit. Based on our analyses, we propose that the decrease of clade U was driven by the shift to a laboratory diet, shared by all experimental populations.

      In addition to that, the paper has a clearly minimalistic experimental approach to address functional properties of the revealed L. plantarum strains, so that their own fitness, or their relationship with the Drosophila host, is characterised superficially. Therefore, the authors' discourse can be speculative rather than factual (especially when the authors use the expression "likely" to share their guesses in the "Results" section). Nevertheless, these minor drawbacks do not underscore the novelty of the discovered phenotypes and the importance of their further investigation.

      We consider the reviewer's concern and toned down the phrasing when reporting our findings in the revised version of the manuscript.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) One solution to resolve the "competition" issue would be to check that the L. plantarum strains are established at similar or different titers in the in vivo work. Fly phenotypes can be sensitive to microbial load (Keebaugh, iScience, 2018), which might explain some of the counterintuitive in vivo results. In line 227, the authors mention that "bacterial load" is contributing to the magnitude of the effect, but I don't see the data reported anywhere. If this is from the in vitro assays, then the authors need to show that in vitro predicts in vivo L. plantarum abundance.

      Bacterial load inoculated in the in vivo experiment was normalized to OD=0.05 (~5*10<sup>6</sup> CFUs/ml) for the three clades at the beginning of the experiment. Thus, all vials were inoculated with the same titer of L. plantarum. Only the genotype varied between treatments. However, it is possible that, once inoculated, each clade grew to a different titer (as they have different growth rate and carrying capacity).

      Statement in line 227 comes from the differing results in transfers 1 and 2. In transfer 1 we inoculated a fixed load of ~2.5*10<sup>5</sup> CFUs. In transfer 2, however, we inoculated no bacteria to the food, and the flies seeded the vial. Our statement comes from the assumption that bacterial load in transfer 2 has to be lower than in transfer 1 as bacteria seed the vial solely by defecation of the parents.

      Following the reviewer's suggestion, in the revised version of the manuscript we have included a new experiment in which we quantified the bacterial load of each clade in individual flies.

      (2) Tracking the competitive outcomes is tricky, though it could be done with whole genome sequencing. An alternative would be to label the strains with fluorescent proteins (e.g., Obadia Current Biology 2018 has done this in Lactobacillus) and track fluorescence to better understand the results of the "mix" treatment in Figure 6.

      We appreciate the feedback of the reviewer, but consider this rather labour-intensive approach as an interesting option for future follow-up experiments.

      (3) That being said, my main concern with this is the "competition" claim. If the paper were reframed appropriately, this paper could still make an important contribution to the evolution of host-microbiome interactions, but the authors would need to consider what they can and cannot do with this interesting dataset.

      The "competition" claim comes from the changes in relative abundance observed in the time-series data, not from any of the follow-up experiments. Thus, we consider the use of the term "competition" appropriate.

      (4) The text on the figures is very small and hard to read.

      We increased the size of the text in all figures.

      Reviewer #2 (Recommendations for the authors):

      (1) Have you conducted the in vitro culture experiments following the "cold" conditions?

      We have conducted the experiment in "cold" conditions, but with some modifications to the experimental settings, as the plate reader did not have cooling capacity. Instead, we grew a subset of the isolates (four per clade) in glass vials at constant 20 °C, and measured their OD twice a day. We have included the results in the revised version of the manuscript.

      (2) How many technical and biological replicates were measured for the in vitro culture experiments (Figure 5)? Please add this information to the figure legend and method.

      We measured the growth of four isolates from clade U, nine isolates from clade C and sixteen isolates from clade H. Each isolate was grown three times.

      We have included this information, as requested by the reviewer.

      (3) Making the labels in Figures 2, 3, 5, and 6 bigger would be helpful.

      We have increased the font size of all figures.

      Reviewer #3 (Recommendations for the authors):

      (1) Line 268: "Based on our results in experimentally evolved fruit flies, we propose that within-species competition, thus far largely overlooked, could contribute to ecological adaptation and evolution of the host". Overstatement should be avoided, since the evolution of the host was not directly studied here.

      Our results show that reproductive traits of the host differ upon colonization with each clade. Although we don't test the host's evolution, we speculate that flies differing in their offspring number and developmental time might differ in their overall fitness. Finally, we consider the Discussion section as the right place for speculation and development of hypotheses that can be tested in future work.

      (2) Line 258: "These differences do not explain the clade-specific selection, but reflect the different evolutionary histories of the clades". The temperature factor and its possible role in clade selection would be better discussed at least a little bit.

      In this paragraph we described potential metabolic differences between clades using comparative genomics. We did not find enrichment in a function or group of functions that could explain the different dynamics between clades H and C in the temperature regime.

      In the revised version of the manuscript we highlight that we did not find temperature-specific differences from this analysis.

      (3) Line 252: "...This could explain why clade U, which displayed a high growth rate and carrying capacity in liquid culture". The statement could be further developed with a caution. Even if the isolates that belong to the clade U are outcompeted by H or C, it should be noted that the strain U cannot be used as a true reference for fitness, since it could possess its hidden adaptive properties, not being simply "a loser". Such a hypothesis could explain the maintenance of this strain in the wild.

      We agree with the reviewer in that fitness is relative to the selective environment. Clade U is less fit than H and C in our specific experimental conditions, but it could outcompete them in other conditions, such as wild flies or MRS liquid medium. In the revised version of the manuscript we have rephrased this statement to clarify that we specifically refer to clade U's fitness under the new laboratory conditions.

      (4) In a cold environment, association with the clade C induces developmental delay and produces less progeny, which potentially allows the host to survive in case of harsh conditions and potential food limitation. Could the authors speculate and not exclude that this phenotype could be potentially adaptive? It would be curious to check in further studies whether flies associated with C strains are more stress-resistant, for example.

      We thank the reviewer for this alternative hypothesis. In our manuscript we used the Darwinian definition of fitness; reproductive success of an organism in the focal environment. And thus, both higher progeny per female and shorter developmental time would be beneficial in direct competition with other individuals. It is true that delayed developmental time, or less progeny, could be advantageous in specific cases. This could be the case for D. simulans inoculated with clade C. However, we consider that the fecundity levels observed in D. melanogaster upon inoculation with Clade C (average of 0.06 offspring/female*day in the cold) are too low to sustain a population.

      We have included this hypothesis in the Results section.

      (5) It would be highly recommended to add an experiment to complete the story by measuring the quantity of bacteria in the medium and in the flies. This will resolve the hypothesis (Line 785): "Thus, the ability to exploit this ubiquitous source of carbon and nitrogen could be very advantageous in the fly microbiome context, but would not affect the fitness in liquid culture".

      Following the reviewer's recommendation we included two additional experiments. We measured the bacterial load per fly in the native host, D. simulans, inoculated with the three clades. We also compared the clades' growth speed in solid fly food (without host). In the former experiment, we found similar bacterial loads upon inoculation with clades U and clade H. In contrast, in the latter we found delayed growth of clade U relative to H and C in the food. Thus, chitobiose consumption does not seem to provide an advantage in the fly gut to clade H. We attribute the fitness advantage of H and C to their advantage growing on the laboratory fly food, regardless of the host.

      Both experimental results have been included in the revised version of the manuscript, and the comparative genomics paragraph and discussion have been modified in consequence.

      (6) The chapter "Extended clade-specific differences in KEGG metabolic pathways" could be presented in the main text as it contains important results. These results are mentioned in the chapter "Functional divergence on the genomic level", which looks rather humble when it stands alone as it currently does.

      We appreciate the interest of the reviewer in this supplementary chapter. To keep the length of the manuscript digestible for a broad set of readers, we decided to only include in the main text the functional differences that could play a role in adaptation to the new laboratory environment.

      We consider that a full description of the metabolic differences between the three clades has to be published, as it might be relevant for researchers interested in L. plantarum metabolism. However, it does not fully follow the storyline, as the differences reported in the supplementary, such as nitrate respiration or synthesis of molybdenum cofactors, might not be involved in the clade-specific selection observed in the time series.

      (7) Line 773: "Clades C and H encode a shared genetic repertoire related to sugar/riboflavin metabolism that is lacking in clade U". This indeed allows us to hypothesise that the fixation of these clades in fly populations was due to their improved metabolic capabilities. However, the analysis of fitness shows similarity in flies associated with clades H and U, meaning that sugar/riboflavin metabolism in H does not provide an obvious adaptive trait to flies. Moreover, one could say that sugar metabolism in clade C is maladaptive not only for flies, but also for bacteria in liquid cultures. It is recommended to more clearly state the respective limitations of the study.

      Here we have to make a distinction between bacterial fitness and host fitness. The three clades differ in their (bacterial) relative fitness, as evidenced by the time-series dynamics (Figure 4). In the cited statement we hypothesize that a more versatile sugar metabolism repertoire could increase the bacterial fitness of clades H and C (relative to clade U) in the sugar-rich laboratory diet.

      This is independent of the fitness effect that L. plantarum could have in the host. Finally, as it was discussed in the recommendation 3, fitness is specific to the environment. Clade C is the least fit in liquid MRS in hot conditions, but the fittest in cold experimental conditions.

      (8) The authors should better explain why growth in MRS was not performed in a cold temperature regime to further support or refute the hypothesis that capacity and inflection time could partially explain the higher fitness of bacterial strains from clade U.

      We did not perform this experiment in cold conditions due to technical limitations of the plate reader, that does not have cooling capacity. Nevertheless, following the reviewers' suggestion, we have included in the revised version of the manuscript a new MRS growth experiment in cold-like conditions (constant 20 °C).

      (9) When mentioning that L. plantarum can "increase larval fitness of Drosophila melanogaster relative to germ-free flies" (line 196), the authors should specify in which specific conditions this phenotype was observed, and how relevant the mentioned phenotypes are to the current study.

      Following the reviewer's recommendation, we have modified the paragraph in order to clarify the conditions used in other papers and those used in our work. The references cited in this section (PMID: 21907145, 29290388 and 28062579) report that L. plantarum increases the host fitness in protein-poor diets (12 g/l of dried yeast or less), but not in high-protein diet (50 g/l of yeast or higher). Since our experimental diet contains an intermediate amount of protein (24.3 g/l of dried yeast) we were agnostic of whether L. plantarum would benefit the host or not in our conditions. Regarding the phenotypes, we chose two reproductive traits that are affected by changes in the microbiome according to the literature. Developmental time is directly affected by L. plantarum in the aforementioned papers. Offspring number is another fitness component affected by Drosophila microbiome (PMID: 30510004).

      (10) Provide a reference for line 205: "In axenic D. melanogaster none of the L. plantarum clades provided a fitness advantage to the host relative to germ-free controls, contrary to the effects reported in the literature". If the conditions were different from those in the studies referred to, then it would be of no use to compare the fitness advantage (for example, in Reference 24 another type of diet was used).

      Already covered in recommendation 9.

      (11) Please provide more context to this statement (Line 210): "The high content of dried yeast 24.3 g/l in the fly food used in our experiment likely provided already sufficient amounts of essential amino acids, which negated the growth-promoting effects of L. plantarum". It is not clear why amino acids are taken into account, and what the evidence is for the fact that the amount of essential amino acids was sufficient to abolish growth-promoting effects.

      The whole paragraph was modified in order to clarify the relationship between protein input and nutritional fitness benefit of L. plantarum.

      (12) Please provide measurements of bacterial quantity which would support the statement (Line 215): "The fitness reduction was stronger in the first transfer of flies, likely due to a higher bacterial load".

      Upon request of the reviewer, we have estimated the bacterial load per individual fly in D. simulans. Additionally, we have specified the CFUs inoculated in the vials in transfer 1.

      (13) Correct the typo (line 220): "However, the developmental time was significantly extended after inoculation with clade C at cold temperature (Dunn's test, p < 0.05 05 for all significant comparisons)".

      Done.

      (14) Specify more precisely the temperature conditions referred to in line 226: "In summary, we observed that clade C, which is dominant in the cold-evolved populations, decreases host fitness when axenic flies are inoculated". Does it decrease fitness both in hot and cold environments?

      For the axenic flies, we did find a decrease in fitness in both regimes, yes. We specified it in the revised version of the manuscript.

      (15) Please provide evidence for line 227, or otherwise rephrase it: "The magnitude of this effect varies depending on the environmental temperature, the bacterial load, and the presence of other microbial taxa".

      Novel evidence was provided regarding the role of bacterial load on host fitness.

      (16) Correct the following statement, so that it reproduces the results of the original work (reference 19, line 229): "In a low-protein diet, strains that were not isolated from Drosophila enhanced larval growth relative to germ-free individuals, whereas another Drosophila-associated strain did not have any effect".

      This statement was removed from the revised version. This reference was cited in the discussion to state that: " the nutritional symbiosis in L. plantarum is strain-specific".

      (17) Please provide a rationale for using KEGG Orthologs. Why was this database chosen as an appropriate one, even though it is known to be a non-exhaustive metabolomic resource?

      KEGG is a well-known metabolic database that is widely used in comparative genomics (PMID: 40177264) and built in state-of-the-art software for microbial ecology such as Anvi'o (PMID: 33349678). Other similar gene-to-function databases are less focused on metabolic pathways, such as COG or GO, or limited to specific enzymatic activities, like CAZy. Furthermore, the hierarchical organization of KEGG Orthologs in modules and pathways allowed us to map clade-specific orthologs to the broad metabolic context. For these reasons, we considered KEGG to be the best option for this analysis.

      (18) Line 250: "Therefore, we speculate that the ability to exploit this ubiquitous source of carbon and nitrogen in the lab-maintained fruit flies, could be a strong target of selection in the lab environment". This statement concludes the "Results" section but would be more appropriate for the Discussion section, since the authors do not provide any experimental evidence that could support this statement.

      We have modified this chapter, as covered in recommendation 5.

      (19) Line 276: "However, the intraspecific richness of L.plantarum in our flies was three times higher than that estimated in human gut microbiomes". Note that there are other recent studies which show the presence of several OTUs within L.plantarum isolates (for example PMID: 41484402).

      We thank the reviewer for the reference. We comment on it in the revised manuscript.

      (20) Line 287: "Our finding shows that the well-characterized nutritional symbiosis between Drosophila and L. plantarum depends on the bacterial genotype and cannot be generalized to the entire species". Note that such a conclusion has already been previously stated (for example, PMID: 30008290 and 28993620).

      We thank the reviewer for the references. Indeed, these papers show that some L. plantarum strains are beneficial for the host while others are neutral. Furthermore, as commented by Reviewer #1 in the public review, L. plantarum has been shown to reduce the host's fitness relative to axenic flies (Gould, PNAS, 2018).

      Our observations are novel in two ways. (1) The fecundity observed in D. melanogaster, 0.06 offspring/female/day in average, is lethal (in Gould et al. 2018 fecundity never decreased below 1 offspring/female/day). (2) Clade C outcompetes the other clades in the cold, despite being detrimental for the host.

      We have modified the Discussion to account for the previous work.

      (21) Line 343: "In addition, we obtained L. plantarum genomes from two other experimental evolution studies. Two genomes from the South African experiment and seven genomes from the Portugal experiment". Merge two sentences into one.

      Done.

      (22) Line 360: "At sampling, the age of the flies varied between four and eight days for the hot environment and between nine and 16 days for the cold environment". Please comment on the fact that different age of flies (different physiology) is not the reason for bacterial community differences.

      During maintenance, flies are sampled at different ages because the temperature affects their developmental time. We cannot rule out the hypothesis that age difference drives microbiome differences. Temperature could affect clade competition directly (differences in optimal temperature between clades) or indirectly, by affecting either the host (e.g. changes in Drosophila developmental time alters L. plantarum fitness), the surounding microbiome, or the food (e.g. increased metabolic activity in the hot regime changes nutrients profile). We ruled out the direct effect of temperature with growth experiments in liquid MRS medium and solid fly food, but disentangling the indirect effects is not feasible.

      (23) The majority of figures have low-quality labels that are not legible due to the small size of the font. Please improve.

      Done.

      (24) Figure 1 - Correct the legend: There is no "10" label on the picture. Probably by 10, the authors mean "Generation", while by x10 - number of isogenic replicates.

      Done.

      (25) Figure 2 - No numbers at nodes are indicated, whereas it is announced in the legend that they represent bootstrap support values. In addition, it is recommended to show a reference pangenome in the middle panel to clearly refer to the total size of the possible black bar.

      We added high bootstrap support as coloured nodes in figures 2, 3 and S2.

      We do not understand the reference pangenome request. In the middle panel, each black/white bar corresponds to an orthologous gene that can be either present or absent in each of the genomes. These orthologs were sorted based on hierarchical clustering of the their patterns of abundance (columns present in the same set of genomes, together), not by synteny. Thus, a reference pangenome would be simply a black bar.

      (26) Figure 2: It would be advantageous to add a figure that represents the frequency of each strain in each replicate (at the last time point, for instance). It would explain why some “blue” strains appear to be within the “red” cluster. Otherwise, it is confusing to find cold-evolved bacterial strains in hot-evolved fly populations.

      The frequency of each clade in each replicate is shown in figure 4. We think that it would be more confusing to follow the suggestion of the reviewer, as the isolates were sampled at different time points of the experiment. We would not like to call it a confusion that "blue" strains appear in the "red" cluster, but rather the logical consequence of the color code used in figure 2, which corresponds to the temperature regime in which the isolate was sampled (regardless of its clade). Whereas in the following figures colour represents the clade. It is thus possible to find clade H isolates in the cold temperature regime, as this clade is in low frequency but not completely absent in this regime.

      (27) Figure 3 – Add a label for the X-axis.

      We rotated the tree to be able to increase the genome IDs. We have added the label to the Y-axis.

      (28) Figure 4 - Please indicate how the clade relative abundance was assessed.

      Clade relative abundance was inferred by mapping competitively the short reads against the three clades’ reference sequences. It is specified in the legend now.

      (29) Figure 6 - Total number of F1 flies eclosed normalised by day (during which period?). What do T1 and T2 correspond to?

      During the respective number of days that females were allowed to lay eggs: one day in the hot settings and two days in the cold settings in transfer 1. One day and three days, respectively, in transfer two.

      T1 and T2 correspond to the first and second transfers, as described in the Materials and Methods. In first transfer, flies laid eggs in vials pre-inoculated with a set load of L. plantarum. After egg laying, same adults were then transferred to a sterile set of vials and allowed to lay eggs again (second transfer). Bacterial load in these vials was solely seeded by the parents.

      In order to avoid any confusion, in the revised version of the manuscript we have modified figure 6 to show transfer 1 for both Drosophila species, and moved transfer 2 dataset to supplementary figure S5.

      (30) Figure S2 - label the X-axis.

      We guess the reviewer means Y-axis. Done.

      (31) Figure S3 demonstrates the real data and its variability, so it would be better used instead of Figure 4 (which seems to be just a derivative from Figure S3, not a separate dataset and separate type of analysis).

      As the reviewer suggested, we have replaced figure 4 with figure S3.

      (32) Figure S4: Improve plot title: (e.g., C:H:U = 3:3:3).

      Done.

      (33) Figure S6: It is stated that N = 10; however, some datasets do not have 10 points represented. Please specify why. Also, please specify the meaning of "T1/T2".

      For the inoculation experiment in Drosophila simulans, we had nine replicates per treatment, not ten. This has been corrected in the figure and in the Materials and Methods section.

      T1 and T2 correspond to the first and second transfers, already covered in recommendation 29.

      (34) Table S3: provide legend for values (1 - present in all strains, but 0.04 - what does it mean?).

      It means that 4% of the genomes from this clade harbour the specific gene. We have specified it in the legend of the revised table.

    1. eLife Assessment

      This manuscript focuses on developing a structural model of how the multidomain ECM protein SVEP1 enables Angiopoietin (ANG) binding to the orphan receptor TIE1, resulting in downstream receptor phosphorylation and signaling. This is an important study, based on solid data. The results will be of interest to scientists working in vascular biology and RTK signaling.

    2. Reviewer #1 (Public review):

      Summary:

      In this manuscript, Uphoff et al. propose a structural and mechanistic model in which the multidomain ECM protein SVEP1 enables Angiopoietin (ANG) binding to the orphan receptor TIE1, thereby promoting downstream receptor phosphorylation and signaling. Using AlphaFold-based modeling, the authors predict that the CCP20 domain of SVEP1 binds to TIE1, creating a composite surface that facilitates Angiopoietin association and TIE1 activation. The resulting ternary model (SVEP1-TIE1-ANG) offers a structural rationale for how SVEP1 converts TIE1 into a functional, ligand-responsive receptor. Additional models and biological assays suggest roles for other domains of SVEP1, such as CCP5-EGF-L7, although these interactions are predicted with low confidence. The authors interpret these findings as the first structural framework for how SVEP1 enables ANG-TIE1 signaling.

      Strengths:

      (1) The central hypothesis - that SVEP1 enables ANG binding to the orphan receptor TIE1 - is biologically compelling and addresses an important question in vascular biology.

      (2) The AlphaFold-predicted ternary complex (SVEP1-TIE1-ANG) is plausible, high-confidence, and structurally consistent with prior functional data (e.g., poly-Ala scanning from Sato-Nishiuchi et al.).

      (3) The authors' model offers a potential explanation for the previously observed role of SVEP1 in enhancing ANG signaling through TIE1 and may represent the first structural insight into TIE1's transition from orphan to ligand-activated receptor.

      (4) The potential clinical implication - that a combinatorial ligand (ANG+SVEP1) can activate TIE1- could have translational relevance for vascular leak and inflammatory disease.

      Comments on revised version:

      The authors have adequately addressed my concerns.

    3. Reviewer #2 (Public review):

      Uphoff and colleagues present the results of a study focused on characterizing the binding of SVEP1 to TIE1 along with Angiopoietin-2. Starting with computational prediction of SVEP1 binding to TIE1, the authors identify the region of SVEP1 that serves as a high-affinity ligand for TIE1. Advanced studies identify a weak secondary binding site within SVEP1 that appears to be sufficient but not necessary for its interaction with TIE1 based on in vivo rescue experiments. The most novel contribution of the manuscript seems to be the identification of angiopoietin-1 and -2 as co-factors that seem to enhance the binding of SVEP1 with TIE1 and impact downstream AKT signaling. They propose a complex in which SVEP1 binds to TIE1 and ANG2.

      Although the first set of results is essentially confirmatory, the identification of ANG-2 as a "co-factor" enhancing the binding of SVEP1 to TIE1 and associated downstream signaling (i.e., Figures 3 and 4) is novel and is of interest. However, the manuscript and its conclusions would greatly benefit from some clarifying details and additional experiments to ensure rigor and support specific claims.

      Comments on revised version:

      I have no further comments. The authors have addressed my concerns.

    4. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      In this manuscript, Uphoff et al. propose a structural and mechanistic model in which the multidomain ECM protein SVEP1 enables Angiopoietin (ANG) binding to the orphan receptor TIE1, thereby promoting downstream receptor phosphorylation and signaling. Using AlphaFold-based modeling, the authors predict that the CCP20 domain of SVEP1 binds to TIE1, creating a composite surface that facilitates Angiopoietin association and TIE1 activation. The resulting ternary model (SVEP1-TIE1-ANG) offers a structural rationale for how SVEP1 converts TIE1 into a functional, ligand responsive receptor. Additional models and biological assays suggest roles for other domains of SVEP1, such as CCP5-EGF-L7, although these interactions are predicted with low confidence. The authors interpret these findings as the first structural framework for how SVEP1 enables ANG-TIE1 signaling.

      Strengths:

      (1) The central hypothesis - that SVEP1 enables ANG binding to the orphan receptor TIE1 - is biologically compelling and addresses an important question in vascular biology.

      (2) The AlphaFold-predicted ternary complex (SVEP1-TIE1-ANG) is plausible, high-confidence, and structurally consistent with prior functional data (e.g., poly-Ala scanning from Sato-Nishiuchi et al.).

      (3) The authors' model offers a potential explanation for the previously observed role of SVEP1 in enhancing ANG signaling through TIE1, and may represent the first structural insight into TIE1's transition from orphan to ligand-activated receptor.

      (4) The potential clinical implication - that a combinatorial ligand (ANG+SVEP1) can activate TIE1- could have translational relevance for vascular leak and inflammatory disease.

      Weaknesses:

      (1) Lack of structural validation and mechanistic follow-up: Despite the promising AlphaFold model, there are no figures of the predicted interface, no residue-level interactions shown, no ipTM values reported, and no experimental follow-up to test the interface. PAE plots are incorrectly used as confidence justifications, which is not appropriate for complex predictions.

      We have appended the data showing AlphaFold-predicted interfaces, including residues, hydrogen bonds, and surface complementarity. We also added ipTM scores and confidence plots for the predicted complexes.

      (2) Biophysical validation is missing: No surface plasmon resonance (SPR), ITC, or biochemical assays are included to confirm ternary complex formation or quantify binding kinetics. Given the manuscript's structural focus, this is a major gap. For instance, an SPR experiment where ANG is immobilized, and TIE1 binding is measured {plus minus} SVEP1, would directly test the model. And allow direct comparison to ANG-TIE2.

      We have addressed this question and performed ELISA assays to measure binding affinities between SVEP1 and TIE1 in presence or absence of ANG1 or ANG2, thus confirming that the affinity is increased in the presence of ANG1 or ANG2.

      (3) Missed opportunity for mutagenesis-driven validation: The manuscript does not include any interface-targeted mutations, despite clear opportunities. For example, mutating T2595 in SVEP1 (to R) or mutating the TIE1-specific residues (residues PL 202-203 to LF) could strongly test the model and potentially reveal dominant-negative behaviors. E.g. A T2595 mutant should block ANG binding but not TIE1 binding.

      We have depicted figures of the interfaces including P202-L203 and included the TIE1 P202L L203F mutant, as well as the previously described SVEP1 (E2568A - G2569A) mutant in our experimental data. The T2595 mutant was not included in the current study, for the following reason: Modeling suggested that replacing T2595 with an Arg will cause steric and charge clashing with 469GKL471 of ANG1 and 467NKFN470 of ANG2, thus reducing its binding to ANG2 although T2595 does not interact with ANG1/2. A SVEP1 protein comprising CCP15 to the C-terminus with the T2594R mutation shows reduced binding to ANG2, but also reduced binding to TIE1. As the mutation hinders interaction with both TIE1 and ANG2, the data is not included in the manuscript.

      (4) Overinterpretation of weak models: The additional AlphaFold model involving the CCP5-EGFL7 domains binding TIE1 has extremely low confidence (ipTM < 0.15) when reexamined by this reader and should not be emphasized. There is no biophysical evidence or binding data (SPR) to support this interaction, and its inclusion detracts from the much stronger CCP20 model.

      We agree with this point made by both reviewers and have removed the data on CCP5-EGFL7 from the manuscript.

      (5) Language around modeling is overstated and potentially misleading: Terms like "unequivocal," "high-affinity," or "affirms strong binding" in reference to AlphaFold predictions are inappropriate. These are hypotheses -not confirmations - and must be tested at the biochemical level. This should be clarified throughout the manuscript to ensure non-experts do not misinterpret modeling confidence as binding affinity.

      We agree with the reviewer, and have adjusted the wording.

      (6) Negative stain EM data is not informative due to low resolution and lack of defined interfaces; unless replaced by higher-resolution Cryo-EM, this should be omitted. Better would be co-gel filtration, AUC, or SEC-MALLs with ANG-SVEP1-TIE1.

      We have now appended the data by adding gold-labelled TIE/ANG proteins, thus enhancing clarity.

      (7) Disjointed narrative: The manuscript presents a compelling mechanism involving CCP20-driven ANG binding to TIE1, but then becomes fragmented by introducing the low-confidence CCP5-EGFL7 model and speculative higher-order polymerization models that are not experimentally supported.

      We agree with this point and have have centered the manuscript around CCP20. We removed data concerning CCP5-EGFL7 as suggested by both reviewers.

      Reviewer #2 (Public review):

      Uphoff and colleagues present the results of a study focused on characterizing the binding of SVEP1 to TIE1 along with Angiopoietin-2. Starting with computational prediction of SVEP1 binding to TIE1, the authors identify the region of SVEP1 that serves as a high-affinity ligand for TIE1. Advanced studies identify a weak secondary binding site within SVEP1 that appears to be sufficient but not necessary for its interaction with TIE1 based on in vivo rescue experiments. The most novel contribution of the manuscript seems to be the identification of angiopoietin-1 and -2 as co-factors that seem to enhance the binding of SVEP1 with TIE1 and impact downstream AKT signaling. They propose a complex in which SVEP1 binds to TIE1 and ANG2.

      Although the first set of results is essentially confirmatory, the identification of ANG-2 as a "cofactor" enhancing the binding of SVEP1 to TIE1 and associated downstream signaling (i.e., Figures 3 and 4) is novel and is of interest. However, the manuscript and its conclusions would greatly benefit from some clarifying details and additional experiments to ensure rigor and support specific claims.

      We have addressed the reviewers concerns and significantly appended the manuscript. Most importantly, we provide structural validation of AlphaFold models reporting interfaces, residue-level interactions and ipTM values. We have included new biophysical validation of binding kinetics of SVEP1 and TIE1 in the presence or absence of ANG1 or ANG2. Furthermore, we have removed the data on CCP5-EGFL7 from the manuscript in order to retain focus on the CCP20 domain.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      The AlphaFold modeling for the CCP20-based interactions is strong (as determined by this reader rerunning and getting ipTM values and visually inspecting the interactions “because this is not in the manuscript”). As presented, the manuscript stops at the hypothesis-generation stage. Validation is needed to fulfill the paper's title and claims. The structure-function link is not demonstrated, despite an obvious and achievable experimental path (mutagenesis, SPR, kinetics).

      Additional Context and Suggestions

      (1) Show and label AlphaFold-predicted interfaces, including residues, hydrogen bonds, and surface complementarity.

      We have appended the data in the new supplementary figures 1.1, 1.2, 2.1 and 2.2

      (2) Provide ipTM scores and confidence plots for each predicted complex.

      We have added the values and plots in the new supplementary figures.

      (3) Perform SPR assays with ANG-coated surfaces and measure binding of TIE1 {plus minus} SVEP1. Compare to TIE2 binding for context.

      We performed the proposed experiment using an ELISA assay to measure binding affinities between SVEP-1 and TIE1 in presence or absence of ANG2 and included these data in the manuscript in figure 2.

      (4) Test interface mutants: e.g., T2595R in SVEP1 (should impair ANG recruitment but not TIE1 binding), or PL→LF muta on in TIE1 (should disrupt SVEP1 binding).

      We have depicted figures of the interfaces including P202-L203 and included the TIEP202L L203F mutant, as well as the previously described SVEP1 (E2568A - G2569A) mutant in our experimental data. The T2595 mutant was not included in the current study. Our modeling suggested that replacing T2595 with an Arg will cause steric and charge clashing with 469GKL471 of ANG1 and 467NKFN470 of ANG2 thus reduce its binding to ANG2 although T2595 does not interact with ANG1/2. A SVEP1 protein comprising CCP15 to the C-terminus with the T2594R mutation shows reduced binding to ANG2, but also reduced binding to TIE1. As the mutation hinders interaction with both, TIE1 and ANG2, the data is not included in the manuscript.

      (5) Clarify in the Introduction that SVEP1 is a large, multidomain ECM protein to help readers contextualize the domain names early on. Do not use terms like CCP before defining them.

      We added: “Svep1 encodes a 3571 amino acid long extracellular matrix protein containing different domains such as Willebrand factor type A domain (vWF), ephrin-receptor like domains, complex control protein (CCP) domains, and Hyalin repeats at the N-terminus. The C-terminus mainly consists of CCP and EGF domains. Svep1 is expressed in mesenchymal cells, but not in endothelial cells, and functions non-cell-autonomously (Karpanen et al. 2017; Morooka et al. 2017)”.

      (6) Replace or remove negative-stain EM unless higher-resolution cryo-EM data are available.

      We would like to retain the EM data, but have now replaced the negative-stain EM by adding gold-labelled TIE/ANG Proteins to verify that the proteins we show are the ones we expect. The reason we would like to retain the data is that TIE1 has been considered for so many years as an orphan receptor, and thus we consider it appropriate to demonstrate SVEP1/TIE1 binding using multiple methods.

      (7) Reframe claims around modeling to avoid overstatement. For example: "The model suggests a plausible mechanism consistent with prior biochemical data" is more appropriate than "unequivocal".

      We agree with the reviewer, and have adjusted the wording.

      (8) Consider narrowing the focus: the CCP20-TIE1-ANG model is a strong story on its own. The CCP5EGFL7 model and polymerization hypotheses are not essential and may dilute the impact.

      Since this point was made by more than one reviewer, we have removed the data on CCP5EGFL7 from the manuscript.

      (9) Properly define TIE1: Tyrosine kinase with Ig and EGF domains.

      We have corrected the full protein name for TIE1 and added: “TIE1 and Tie2 exhibit a high degree of homology with a globular head domain consisting of three immunoglobulin-like (Ig) domains and three epidermal growth factor-like (EGF) modules and a short stalk formed by three fibronectin type III repeats, while the N-terminal two Ig domains of Tie2 harbor the angiopoietin binding site (Macdonald et al. 2006).” We also added: “D1 and D2 refer to the two N-terminal Ig domains, D3 refers to the three EGF domains and D4 to the third Ig domain of TIE1 or TIE2.”

      (10) Use RTK, not tyrosine kinase receptors TKR.

      Tyrosine Kinase receptor was replaced by receptor tyrosine kinases (RTKs)

      (11) PDBs (.cif and .json files) of the models must be supplied for the readers so they don't need to rerun the AlphaFold jobs.

      We are providing all PDBs with the revised manuscript.

      Reviewer #2 (Recommendations for the authors):

      Major comments:

      (1) In several locations, the authors state that alphafold detects an "unequivocal" and/or "high affinity" interaction. Experts on computational structure prediction can weigh in, but I am not sure it is accurate to say that alphafold predicts affinity. Quantitative estimates of the prediction confidence or other parameters of the alphafold output are not provided.

      Thank you for the comment, we agree with the reviewer and have adjusted the wording. We have also appended the data showing interfaces, ipTM scores and confidence plots for the predicted complexes.

      (2) In Figure 1C, what concentrations of TIE1 and TIE2 are being used in the SPR experiments shown?

      We have added the concentrations of the proteins to the methods section

      (3) In Figure 1C, what is the affinity constant (KD) of the interaction between SVEP1 and TIE1 and SVEP1 and TIE2?

      We have added the values of affinity constants to the manuscript text.

      (4) In Figure 1C, the authors immobilize a "70kD" fragment of human SVEP1 to determine the interaction between SVEP1 and TIE1. How does the affinity they measure between the SVEP1 fragment and TIE1 compare to the affinity between immobilized full-length SVEP1 with TIE1?

      The largest molecule we used in any assay is not full-length SVEP1, but consisted of a C-terminal SVEP1 protein spanning from the first EGF domain to the C-terminus (approximately 295 kDa) as previous studies have shown that SVEP1 is proteolytically cleaved N-terminal to the first EGF domain, generating a protein of this size. In our ELISA assays, this molecule has a higher affinity to TIE1 than the 70 kD fragment. We have included data for the 70 kDa as well as for the 295 kDa fragment in the manuscript. ELISA assays have used larger SVEP1 fragments (as indicated in the figure legends), the SPR assay was performed with the 70 kDa fragment.

      (5) How does the alphafold structure prediction for the "70kD" fragment of hSVEP1 compare to the prediction of the same 70kD fragment from the full-length protein prediction?

      All modellings using different SVEP1 fragments including the 70 kD and full-length version identify the putative binding site at position CCP20. While AlphaFold3 can predict the correct domain folds in both the 70kDa fragment and full-length SVEP1, the orientation of these domains is highly variable due to flexible linkers between each domain module. This flexibility effects the output confidence metrics, thereby hampering our interpretation of the models. Therefore, we conducted the structural prediction of the complexes with smaller fragments and not the full-length SVEP1.

      (6) What are the amino acids for the 70kD fragment?

      The relevant amino acids are 2261-2890. The accession number and amino acids of each protein have been listed in the key resources table.

      (7) In Figure 1D, what is being depicted by the red stars? This is not explained in the text or figure legend.

      We have replaced figure 1D.

      (8) By itself, Figure 1D is not terribly informative and in my opinion does not support the statement that the authors "were able to directly visulalize the attachment of SVEP1 and TIE1." As a minimum, the authors should repeat the same set of images with SVEP1 and TIE2, but other approaches, such as labeling, could be performed.

      We replaced figure 1D with new data and gold-coated protein enhancing clarity. We think it beneficial to demonstrate SVEP1/TIE1 binding using multiple methods as TIE1 has been considered as an orphan receptor for so many years. We have not performed these experiments with TIE2 proteins as we were not able to show binding of TIE2 to SVEP1 with other assays.

      (9) What is being stained in Figure 1D? Full-length SVEP1/TIE1? Or fragments of these proteins?

      We replaced figure 1D by a new assay with labeled proteins using the 150 kDa version of SVEP1 and the ectodomain of TIE1 as well as ANG1 or ANG2 (new figure 1D and new supplementary figure 2.3). TIE1/ANG proteins were gold-labelled. The protein fragments used in this assay are described in detail in the methods section.

      (10) The authors discover CCP6-EGFL7 as a low-affinity binding region of SVEP1 for TIE1. Is this region in physical proximity to CCP20 (the high-affinity binding region for TIE1) based on alphafold prediction? How would the authors think this is binding TIE1?

      We have removed this data set (see comment to reviewer 1’s request).

      (11) What is the affinity constant (KD) for CCP6-EGFL7 with TIE1?

      We have removed this data set (see comment to reviewer’s 1 request).

      (12) The authors claim that ANG1/ANG2 increase affinity between SVEP1 and TIE1 based on immunoblotting. Immunoblots are semi-quantitative at best. If the claim is higher affinity, I think the authors should measure this by SPR and determine the KD between immobilized SVEP1 with TIE1 in the absence and presence of ANG1 and/or ANG2.

      We conducted ELISA assays (figure 2) showing that the affinity is increased in the presence of ANG1 or ANG2 and agree with this reviewer that this strengthens the data.

      (13) In Figure 3, can the authors explain why ANG1/2 does not pull down with SVEP1/TIE1?

      We noticed that upon transfection of TIE1 into HEK cells, ANG1/2 is almost undetectable anymore in the total lysate. Thus, we believe that after the pull down we are below the detection limit.

      (14) In Figure 3, what is "TL"? I assume total lysate, but this is not specified.

      Thank you, we now specify TL as total lysate.

      (15) In Figure 3 "TL" panel (again, I assume this is total lysate), why are the ANG1/2 immunoblots so weak when co-transfected with TIE1?

      We consider it likely that in the presence of TIE1, ANG1/2 proteins are internalized and digested. Another reason for low signals could be that upon transfection of two plasmids, the amount of ANG1/2 protein is reduced as the cell has limited capacity for transcription and translation.

      (16) In Figure 3B, why is the SVEP1 fragment now 150kD when 70kD fragment was previously used? What domains are contained in this 150kD fragment?

      We now better define the domains of the 150kD SVEP1 fragment. The 150 kD fragment was the one produced first and available in high amounts in our laboratory and thus used for functional assays. The 70 kDa fragment together with ANG2 also induces phosphorylation of AKT, but it was not used in as many conditions/replicates as the amounts we had available were lower.

      (17) In Figure 4A, signaling with SVEP1 by itself and ANG2 by itself should be shown to support the claims being made.

      We added the lines for SVEP1 and ANG2, and also the quantification. SVEP1 itself already affects the phosphorylation of TIE1, most likely because hdLECs produce ANG2 by themselves. We show this with the ANG2 blocking antibody for pAKT.

      (18) For pAKT, what are the concentations of proteins being used and the times of incubation?

      This information is provided in the Materials and Methods section. We added the concentration of the anti-ANG2 antibody, which was missing.

      (19) It seems that p-AKT and AKT are being blotted on different gels. If this is correct, loading controls need to be shown for p-AKT blot.

      We added HSC70 as a loading control for both blots.

      (20) It is interesting that anti-ANG2 antibody inhibits SVEP1-induced p-AKT signaling. As the authors may know, SVEP1 has been identified as a receptor for PEAR1, which also leads to downstream p-AKT signaling, which seems to be independent of ANG2. Do the LECs being used here express PEAR1? If these cells express PEAR1, how do the authors think ANG2 silencing will eliminate SVEP1-associated p-AKT signaling?

      hdLECs express PEAR1. However, we show that p-AKT signaling is attenuated after siRNA KO of TIE1. Thus, the downstream signaling is dependent on TIE1 (Figure3).

      (21) Again, experts on computational structure prediction can weigh in, but I am not sure how to interpret the prediction of the 2:2:2 stoichiometry for the theoretical SVEP1/TIE1/ANG1-2 complex. Are there quantitative estimates of the confidence that can be provided? Did the authors attempt to model this complex with different stoichiometries? It is difficult to know how relevant this model is without any experimental results supporting this result.

      Since 1:1:1 is the smallest possible triple complex, it is our starting point. We can model a 2:2:2 version, but anything larger than this AlphaFold will not run. Furthermore, we now provide quantitative estimates of confidence with the pLDDT, PAE, pTM, ipTM scores for all models including the 2:2:2 complexes.

      Minor comments:

      (1) The authors could consider including a reference to alphafold on line 102.

      We have added a reference for AlphaFold2 and 3

      (2) The authors should refer to surface plasmon resonance (SPR) assays by this term as opposed to using the brand name Biacore.

      We agree with the reviewer and have changed the term Biacore to SPR.

    1. eLife Assessment

      This useful study examines whether microsaccade direction primarily indexes shifts rather than the maintenance of covert spatial attention, offering a potentially informative account of inconsistencies in the prior literature. However, the evidence remains incomplete because the key effects are based on relatively sparse microsaccadic events, while concerns about event detection, fixation control, subject-level robustness, and the interpretation of gaze-density analyses remain unresolved. The correlational design and absence of a neutral condition or independent measure of attentional shifting further limit the central claim. The work will be of interest to researchers studying attention, eye movements, and visuomotor mechanisms.

    2. Reviewer #1 (Public review):

      Summary:

      This manuscript describes a study examining the relationship between microsaccades and covert attention. This question has been widely investigated, with numerous studies showing that during sustained fixation, when subjects covertly attend to a peripheral stimulus, microsaccades tend to be biased toward the attended location. Here, the authors ask whether this microsaccade bias reflects a shift of covert attention or the maintenance of covert attention. They conclude that the bias is primarily driven by attention shifts, a finding that also helps reconcile the seemingly conflicting results of prior research, where the bias was questioned in paradigms that largely involved attention maintenance rather than shifting.

      Strengths:

      A large sample size was used.

      Weaknesses:

      The main weakness is that the authors' response does not adequately resolve concerns about the robustness of the microsaccade analyses. The newly reported event counts reveal that the number of microsaccades per participant is very low, especially in Experiment 2, and highly variable across subjects. Because the key analyses rely on proportions of microsaccades toward versus away from the attended location, estimates based on so few events are likely unstable and may not provide reliable subject-level measures.

      A second major concern is that several additional analyses introduced in the revision appear to suffer from the same limitation. The permutation analyses and angle-partition analyses may give the impression of statistical rigor, but if the underlying averages are based on very few microsaccadic events, the resulting probabilities are difficult to interpret. Further subdividing already sparse data into narrower angular bins likely makes the estimates even less reliable.

      A third concern is that the authors have not fully addressed issues related to microsaccade detection and fixation control. The presence of very small-amplitude events with relatively high velocities raises the possibility that some detected microsaccades may be artifacts. The authors also did not implement the requested exclusion of microsaccades smaller than 5 arcmin or the suggested reanalysis using stricter fixation criteria. These omissions leave open the possibility that the reported effects are influenced by detection errors.

      A fourth weakness is that some of the requested analyses or clarifications were addressed only superficially. The comparison with Brandolani et al. remains minimal, despite being highly relevant to interpreting whether the observed microsaccade-direction effect is transient or sustained. Similarly, the gaze-density plots do not show the raw gaze-position distributions that were requested and may therefore be misleading, because difference maps cannot determine whether subjects were actually fixating centrally.

      Overall, the revision raises additional concerns rather than resolving the original ones. The main conclusions remain insufficiently supported unless the authors can demonstrate that the effects are robust at the individual-subject level, based on adequate numbers of microsaccadic events, reliable detection criteria, and appropriate controls for fixation behavior.

    3. Reviewer #2 (Public review):

      Summary:

      This study aims to test the hypothesis that microsaccades are linked to the shifting of spatial attention, rather than the maintenance of attention at the cued location. In two experiments, participants were required to judge an orientation change at either a validly cued location (80% of the time) or an invalidly cued location (20% of the time). This change was presented at varying intervals (ranging from 500 to 3,200 ms) after cue onset. Accuracy and reaction times both showed attentional benefits at the valid versus invalid location across the different cue-target intervals. In contrast, microsaccade biases were time-dependent. The authors report a directional bias primarily observed around 400 ms after the cue, with later intervals (particularly in Experiment 2) exhibiting no biases in microsaccade direction towards the cued location. Noteworthy, it would have been interesting to observe whether directional biases in microsaccades are also evident when compared to a neutral condition. The authors argue that this finding supports their initial hypothesis that microsaccade biases reflect shifts in attention, but that maintaining attention at the cued location after an attention shift is not correlated with microsaccade direction.

      Strengths:

      The results are straightforward given the chosen experimental design. The manuscript is clearly written, and the presentation of the study and its visualisations are of a high standard.

      Weaknesses:

      The link between attention and microsaccades has been the subject of extensive research over the past two decades. The authors present a potential solution to the conflicting past findings, arguing that attention should be considered a dynamic process that can be broken down into an attention shift and a sustained attention phase. To differentiate between the two components, the authors varied the interval between the onset of the attention cue and the test stimulus. It would have been nice to use a theory-driven criterion (or an independent measure), in addition to their data-driven approach, to distinguish between these components of a dynamic attention concept. Moreover, it is important to note that the current experiments take a purely correlational approach.

    4. Author response:

      The following is the authors’ response to the current reviews.

      We thank the reviewers for their time and for their valuable inputs throughout the review process. We wish to clarify, one final time, the primary scope and empirical grounding of our work for prospective readers.

      Our study was designed to evaluate whether microsaccades track (in a correlative manner) covert visual-spatial attentional shifting, maintenance, or both. We did so within a single dedicated paradigm, across a large sample (N = 48 human participants). Despite remaining criticisms concerning per-participant event counts and microsaccade classification criteria, the key observation remains a striking dissociation (of the link between microsaccades and covert attention) during the initial shifting and the subsequent maintenance of visual-spatial attention. Moreover, we note how the robust effect observed during shifting (but not maintaining) attention, mitigates residual concerns regarding microsaccade sparsity or signal-to-noise ratio.

      We thus remain confident in the empirical foundation of our work and we invite readers to examine the full paper, supplementary materials, and open-access data to evaluate these findings independently.


      The following is the authors’ response to the original reviews.

      We sincerely thank the reviewers and the editors for their careful evaluation of our article and for their valuable input. Building on these suggestions, we were able to further corroborate our main conclusions, make our article more comprehensive, and thereby substantially strengthen the manuscript.

      We have one additional point of our own: we noticed that in our original submission, we had smoothed the data more than intended. Having caught this, we have now reduced the smoothing employed by 2.5 times compared to the original amount of smoothing (the exact smoothing values have also been added to the methods section). Importantly, however, while this has affected how the results look, this has not affected any of our original conclusions.

      General summary

      We would like to first respond to the major points brought forward by both the editorial summary and the public reviews. As we understand, the two main points that were raised regard: (1) the novelty and theoretical importance of our work and (2) the (in)completeness of our results. We start by providing our response to both of these main points below.

      Novelty and theoretical relevance of the work

      Regarding the novelty of our work, we believe the reviews and, by extension, the editorial summary underappreciated the main theoretical value of the question we addressed. Our work set out to investigate whether microsaccades track covert attentional shifting, attentional maintenance, or both. We fully recognise that there are ample prior studies that investigated and reported a link between microsaccades and covert attention, but also underscore how other studies report seemingly contradicting evidence by reporting that there is no such link. One such example is a recent paper by Willett & Mayo in PNAS (2023). Prompted by the recent hypothesis that this seemingly conflicting evidence may be due to prior work investigating attention ‘in different stages’ (van Ede, PNAS, 2023), we set out to address precisely this using a dedicated task that we designed for this purpose. As acknowledged by the summary and public reviews, this helps to reconcile seemingly opposing views in the literature. In our view, such reconciliation has substantial theoretical value.

      While we appreciate that our reported insights may resonate and appear plausible to those working on this topic, we are not aware of any prior studies that directly addressed whether the link between covert attention and microsaccades may fundamentally depend on the ‘stage’ of attentional deployment (‘shift’ vs. ‘maintain’). To fill this key gap and address this timely issue, we developed a dedicated experiment designed to evaluate the relationship between microsaccades and the different stages of attention within a single paradigm. We did so by varying the cue-target intervals to uniquely incentivise early shifting (by having short intervals), while also being able to assess microsaccade biases during subsequent maintenance (in the longer trials). To our knowledge, no previous task has jointly examined these components in this manner. 

      Finally, our inclusion of two widely adopted approaches to fixational control provides yet another source of novelty. Together, we believe that these features position our work as a substantive advance that reconciles seemingly opposing theoretical views.

      Completeness of results

      Regarding the completeness of our results, the editorial summary points to “the absence of independent measures, single-trial analyses, and neutral-condition controls needed to substantiate the central claims”. While the raised points are valuable, they pertain to issues that are tangential to our primary question and stem from unfortunate misunderstandings of key analytical choices, as we now better clarify. We consider our results complete and comprehensive with regards to the main question our studies set out to answer.

      First, regarding the portrayed “need” for independent measures to define the ‘shift window’ of interest, we clarify how our main analysis is completely agnostic to predetermined time windows, as we employ a cluster-based permutation approach to assess our rich time-resolved data across the full time axis. For the complementary analyses that address the ‘shift’ and ‘maintain’ windows more directly, we use a priori defined windows that are based on ample prior literature (from prior literature studying microsaccade biases, as well as from prior literature on the time course of top-down attention as studied through SOA manipulations). Accordingly, even these ‘zoomed in’ analyses rely on time windows that are empirically grounded in prior research. 

      Second, regarding the use of single-trial analyses, we want to emphasise that single-trial predictability is not where our theoretical question resides. We start from the perspective that the relationship between covert visual-spatial attention and microsaccades is inherently probabilistic. Our aim is not to address or question this. Rather, our aim is to determine whether this probabilistic relationship behaves similarly during attentional shifting and maintenance— an issue our analyses directly address. In addition, we also explicitly discuss how the link between microsaccades and attention is fundamentally probabilistic at the single-trial level in our discussion, and prompted by the valuable feedback, we have expanded on this important contextualisation as part of our revision.

      Finally, regarding the portrayed “need” for a neural-attention control condition, we agree that inclusion of a neutral attention condition could be informative for disentangling the ‘benefits’ versus ‘costs’ of attentional cueing. However, such disambiguation is tangential to our central aim. Rather, our behavioural data primarily serve to verify attentional ‘allocation’ also at later cue-target intervals. Observing a difference between valid and invalid cues suffices for this central aim. We also note how inclusion of a neutral condition would have reduced trial numbers and statistical power for our critical conditions of interest. Accordingly, we do not see this as a limitation that challenges our main conclusions. Having clarified this, we embraced this valuable reflection and revised the article to ensure that we do not mention selective ‘benefits’ or ‘costs’ of our cueing manipulation, but refer to ‘the presence of an attentional modulation’ instead.

      Taken together, the explicit design and analysis choices that we made align with the theoretical aims of our study, and the central question we set out to address. The raised points are valuable and we are grateful to have been able to leverage them to improve our article, but we hope to have also clarified how they do not render our findings “incomplete” (as currently portrayed) with regards to the key goal of our article.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This manuscript describes a study examining the relationship between microsaccades and covert attention. This question has been widely investigated, with numerous studies showing that during sustained fixation, when subjects covertly attend to a peripheral stimulus, microsaccades tend to be biased toward the attended location. Here, the authors ask whether this microsaccade bias reflects a shift of covert attention or the maintenance of covert attention. They conclude that the bias is primarily driven by attention shifts, a finding that also helps reconcile the seemingly conflicting results of prior research, where the bias was questioned in paradigms that largely involved attention maintenance rather than shifting.

      Strengths:

      The paradigm and conclusions appear sound and supported by the results. A large sample size was used.

      We thank the reviewer for this clear and supportive summary of our work.

      Weaknesses:

      Weaknesses are mostly related to how the authors enforced fixation in the task, and clarifications are needed regarding some methodological details. A more direct comparison of the effect in the two experimental conditions is missing.

      We thank the reviewer for raising these valuable points. We have now added a direct statistical comparison of the main effect in the two experimental conditions (page 5: “When directly comparing this spatial saccade bias between experiments, we observed that the effect was significantly larger in Experiment 1 than in Experiment 2 from 535 to 745 ms (cluster p=0.013) and 1153 to 1309 ms after cue onset (cluster p=0.029)”). Regarding the fixation points, we will address them in our point-by-point replies below.

      Reviewer #2 (Public review):

      Summary:

      This study aims to test the hypothesis that microsaccades are linked to the shifting of spatial attention, rather than the maintenance of attention at the cued location. In two experiments, participants were required to judge an orientation change at either a validly cued location (80% of the time) or an invalidly cued location (20% of the time). This change was presented at varying intervals (ranging from 500 to 3,200 ms) after cue onset. Accuracy and reaction times both showed attentional benefits at the valid versus invalid location across the different cue-target intervals. In contrast, microsaccade biases were time-dependent. The authors report a directional bias primarily observed around 400 ms after the cue, with later intervals (particularly in Experiment 2) exhibiting no biases in microsaccade direction towards the cued location. The authors argue that this finding supports their initial hypothesis that microsaccade biases reflect shifts in attention, but that maintaining attention at the cued location after an attention shift is not correlated with microsaccade direction.

      Strengths:

      The results are straightforward given the chosen experimental design. The manuscript is clearly written, and the presentation of the study and its visualisations are both of a high standard.

      We thank the reviewer for this clear summary of our work.

      Weaknesses:

      The major weakness of this paper is its incremental contribution to a widely studied phenomenon. The link between attention and microsaccades has been the subject of extensive research over the past two decades. This study merely provides a limited overview of the key insights gained from these papers and discussions. In fact, it attempts to summarise previous work by stating that many experiments found a link, while others did not, and provides only a relatively small number of references. To make a significant contribution, I believe the authors should evaluate the field more thoroughly, rather than merely scratching the surface.

      We thank the reviewer for this valuable reflection. For an elaborate response to the perceived novelty, please see our general summary reply above. In addition, we have added a more thorough evaluation of the field to the introduction (page 2, find relevant paragraph below). We hope that this will provide more context for the manuscript and strengthen its contribution to the field.

      Revised paragraph from introduction:

      “This link between microsaccades and covert visual-spatial attention has been demonstrated repeatedly. Early studies linked the direction of microsaccades to the deployment of covert attention [14, 15] and these findings were later replicated and extended. For example, it has been demonstrated in both humans [14–25] and non-human primates [26, 27]; during both externally directed perceptual attention [14, 15, 17, 20, 21, 24–27] and internally directed attention within visual working memory [16, 18, 19, 22, 23]; and in both perception and action tasks following directional cues [28]. Several studies have further linked the directional microsaccade bias to task performance [18, 21, 25, 28, 29]. For example, following spontaneous microsaccades, perception of visual targets presented in the same direction is better [25], and visual discrimination benefits may start already prior to microsaccade execution [21]. Recent evidence further suggests that microsaccades may even play a causal role in shaping the perception of peripheral stimuli [30].”

      The authors then present a potential solution to the conflicting past findings, arguing that attention should be considered a dynamic process that can be broken down into an attention shift and a sustained attention phase. Although the authors present this as a novel concept, I cannot think of anyone in the field who considers spatial attention to be a static entity. Nevertheless, I was curious to see how the authors would attempt to determine the precise timing of the attention shift and manipulate the different stages individually. However, the authors only varied the interval between the onset of the attention cue and the test stimulus, failing to further pinpoint their dynamic attention concept.

      The current version of the experiment, therefore, takes a correlational approach, similar to initial studies by Engbert and Kliegl (2003) and Hafed and Clark (2002). Meanwhile, we have learned a great deal about the link between microsaccades and attention. Below, I will list just a few of these findings to demonstrate how much we already know. It is important to note that, while the present study cites some of these papers, it does not provide a clear overview of how the current study goes beyond previous research.

      (1) Yuval-Greenberg and colleagues (2014) presented stimuli contingent on online-detected microsaccades. A postcue indicated the target for a visual task, and the target could be congruent or incongruent with the microsaccade direction. The authors showed higher visual accuracy in congruent trials. The authors cited that paper, but it is still important to emphasize how this study already tried to go beyond purely correlational links on a single trial level.

      (2) The Desimone lab (Lower et al., 2018) showed that firing rates in monkey V4 and IT were increased when a microsaccade was generated in the direction of the attended target.

      (3) However, attention can modulate responses in the superior colliculus even in the absence of microsaccades (Yu et al., 2022)

      (4) Similarly, Poletti, Rucci & Carrasco (2017) observed attentional modulations in the absence of microsaccades, or comparable attention effects irrespective of whether a microsaccade occurred or not (Roberts & Carrasco, 2019).

      Thus, in light of these insights, I believe the current study only adds incrementally to our understanding of the link between microsaccades and spatial attention.

      We thank the reviewer for this insightful comment, and for pointing out several important studies on this topic. While we appreciate that our reported insights may resonate and appear plausible to those working on this topic, we are not aware of any prior studies that directly addressed whether the link between covert attention and microsaccades may fundamentally depend on the ‘stage’ of attentional deployment (‘shift’ vs. ‘maintain’).

      To fill this key gap and address this timely issue, we developed a dedicated experiment designed to evaluate the relationship between microsaccades and the different stages of attention within a single paradigm. We did so by varying the cue-target intervals to uniquely incentivise early shifting (by having short intervals), while also being able to assess microsaccade biases during subsequent maintenance (in the longer trials). To our knowledge, no previous task has jointly examined these components in this manner. Moreover, our inclusion of two widely adopted approaches to fixational control provides yet another source of novelty. Together, we believe that these features position our work as a substantive advance that reconciles seemingly opposing theoretical views.

      Regarding the use of single-trial analyses, we want to emphasise that single-trial predictability is not where our theoretical question resides. We start from the perspective that the relationship between covert visual-spatial attention and microsaccades is inherently probabilistic. Our aim is not to address or question this. Rather, our aim is to determine whether this probabilistic relationship behaves similarly during attentional shifting and maintenance— an issue our analyses directly and appropriately address. In addition, we also explicitly discuss how the link between microsaccades and attention is fundamentally probabilistic at the singletrial level in our discussion. Prompted by the reviewer’s valuable feedback, we have expanded on this important contextualisation in our discussion section (page 8: “Therefore, even if microsaccades may more reliably track shifting than maintaining attention, as our current findings show, our findings should not be taken as evidence that microsaccades reliably track attentional shifts at the single-trial level.”). We also incorporated the valuable reference suggestions in our revised manuscript, including in the revised paragraph in our introduction where we provide a more extensive overview of the prior literature, as shown in response to the preceding comment and in the discussion where we discuss the relationship between microsaccades and attention on a single-trial level.

      In general, it is important to have an independent measure of the dynamics of an attention shift. I think a shift of 200-600 ms is quite long, and defining this interval is rather arbitrary. Why consider such a long delay as the shift? Rather than taking a data-driven approach to defining an interval for an attention shift, it would be more convincing to derive an interval of interest based on past research or an independent measure.

      We thank the reviewer for their question. We wish to clarify how our main analysis is completely agnostic to predetermined time windows, as we employ a cluster-based permutation approach to assess our rich time-resolved data across the full time axis. For the complementary analyses that address the ‘shift’ and ‘maintain’ windows more directly, we use a priori defined windows that are based on ample prior literature (from prior literature studying microsaccade biases, as well as from prior literature on the time course of top-down attention as studied through SOA manipulations). Accordingly, even these ‘zoomed in’ analyses rely on time windows that are empirically grounded in prior research.

      The present analyses report microsaccade statistics across all trials, but do not directly link single-trial microsaccades to accuracy. Similarly, reaction times and accuracy were analyzed only with respect to valid vs. invalid trials. Here, it would be important to link the findings between microsaccades and performance on a single-trial level. For instance, can the authors report reaction times and accuracy also separately for trials with vs. without microsaccades, and for trials with congruent vs. incongruent microsaccades?

      We thank the reviewer for their sincere interest in our findings and for the great suggestion of an additional analysis. We have now investigated whether trials with a congruent, incongruent or no microsaccade in the shift window (where congruent or incongruent was determined as based on the first microsaccade within the shift window) have, on average, different reaction times. This analysis did not show significant differences between these three conditions (congruent microsaccade, incongruent microsaccade, no microsaccade).

      In interpreting this observation, we would like to stress that our experiment is not particularly well-suited to this analysis, as the amount of time between cue onset and the target events are highly variable across trials. Because of this clear drawback, we have decided not to include these analyses.

      The study would benefit greatly from including a neutral condition to substantiate claims of attentional benefits and costs. It is highly probable that invalid trials would also demonstrate costs in terms of reaction times and accuracy. It would be interesting to observe whether directional biases in microsaccades are also evident when compared to a neutral condition.

      We thank the reviewer for this valuable reflection. We agree that the inclusion of a neutral attention condition could be informative for disentangling the ‘benefits’ versus ‘costs’ of attentional cueing. However, such disambiguation is tangential to our central aim. Rather, our behavioural data primarily serve to verify attentional ‘allocation’ at later cue-target intervals. Observing a difference between valid and invalid cues suffices for this central aim. We also note how inclusion of a neutral condition would have reduced trial-numbers and statistical power for our critical conditions of interest. Accordingly, we do not see this as a limitation that in any way challenges our main conclusions.

      Prompted by this reflection, we have ensured to not mention selective ‘benefits’ or ‘costs’ of our cueing manipulation throughout the article, but refer to this only as ‘the presence of an attentional modulation’ instead (such changes were made on pages 3 and 9, and we kept this phrasing consistent in our additions on pages 5 and 22).

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) The results resemble recent findings by Brandolani et al. (2025), who also showed that the microsaccade bias-in a similar task and using a comparable analysis approach-was restricted to a narrow time window, primarily around the time of the attention shift. The authors should reference this work, discuss similarities and differences, and, given their larger sample size, consider whether they observe a similar correlation between response times and microsaccade rate.

      We thank the reviewer for pointing out this useful reference to us. We have included this article in our introduction, when sketching the current state of the field (page 2) and also mentioned this related article in the discussion (page 8).

      In addition, prompted by this comment, we have investigated whether we observe a similar correlation between response times and microsaccade rate in valid trials. We have investigated this for both possible timeframes: (1) the ‘shift’ timeframe and (2) the ‘maintain’ timeframe. For both experiments, there was no consistent correlation between response times and overall microsaccade rate, as shown in Author response image 1.

      Author response image 1.

      Relationship between the reaction time and the overall saccade rate. This figure shows the relationship between the average reaction time and the average overall saccade rate for both the ‘shift’ period (from 200 to 600 ms after cue onset) and the ‘maintain’ period (from 600 to 1400 ms after cue onset). Each dot represents one participant. Throughout the entire figure, the following significance levels were used: *: p< 0.05, **: p < 0.01, ***: p < 0.001, ****: p < 0.0001.

      (2) I could not find information on the average number of trials per condition and the average number of microsaccades per subject per condition. Ideally, these numbers should be reported (e.g., in a supplementary table). Since the analysis is based on microsaccade direction, knowing how many microsaccades each subject contributed per condition is critical. Microsaccade rates vary substantially across individuals, and subjects with very few events may add noise to the analysis, as proportions of toward/away microsaccades become unreliable.

      We thank the reviewer for pointing out that this relevant information was missing. We have now included these numbers in a supplementary table as suggested (page 17).

      (3) Relatedly, it was unclear whether the time-course analyses were based on collapsing all microsaccade events across subjects or on subject-level averages. In the Methods, the authors state that "the permutation distribution of the largest cluster size was acquired by randomly permuting the trial-average data at the group level 10,000 times," but this is ambiguous. Please clarify.

      We thank the reviewer for pointing out this ambiguity. We have changed the methods section to reflect more clearly that we first obtain time-courses of the microsaccade rate per participant, and subject these time courses to second level statistics using cluster-based permutation analysis (page 13). We have also reworded the sentence you quoted to remove any ambiguity (page 13): “A permutation distribution of the largest cluster size was acquired by randomly permuting the condition labels of each participant’s trial-averaged time course data (i.e. randomly flipping the sign of the difference in rate of toward vs. away saccades) 10,000 times and identifying the size of the largest clusters observed in these randomised data after each permutation.”

      (4) The authors analyze only downward microsaccades, but the cutoff definition is not specified. Presumably, this includes all directions between 180{degree sign} and 360{degree sign}, which may also include nearly horizontal events. This should be clearly stated. In addition, the rate of upward microsaccades should still be shown, divided into up-left and up-right quadrants to parallel the toward/away analysis. This would provide informative context on whether upward microsaccade rates change systematically over time.

      We thank the reviewer for pointing out that this information was missing, and for suggesting this valuable additional analysis. In the methods section, we now explicitly state the angular cutoffs used for the main analysis (page 12). Additionally, we have added a supplementary figure that shows the time course of upwards microsaccades over time (page 21), please see Supplementary Figure S4.

      (5) If the dataset contains enough microsaccadic events per subject, it would be useful to test more conservative angular cutoffs for defining "toward" versus "away."

      We thank the reviewer for this insightful suggestion. We have included an additional analysis, where only microsaccades were included with a direction within a 45° angle around the exacttoward and exact-away directions. This replicated our main finding. The results from this analysis are now included in the supplementary materials (page 24), please see Supplementary Figure S8.

      (6) Figure 2C: It is unclear what the "Center" and "Border" lines represent. The figure is also potentially confusing because it shows microsaccade amplitudes rather than landing positions. Small amplitudes may still bring gaze close to the target; this distinction should be clarified.

      We thank the reviewer for pointing this out. We have changed the “centre” and “border” labels to include more information (pages 6 and 19). We have also added an in-text clarification of the distinction between saccade amplitude and landing position (page 5: “Note that Figure 2C does not show saccade landing positions. While it is theoretically possible for multiple small unidirectional saccades to lead to a larger change in gaze position, a complementary analysis shows that fixation was maintained during the period of peak microsaccade rate in both experiments (see Supplementary Figure S6).”). In addition, in response to the related comment below, we have now also added heatmaps of gaze showing that gaze overall remained close to fixation in our tasks.

      (7) From the Methods, it appears that in Experiment 1, there was no automatic criterion for discarding trials in which gaze deviated from fixation. In Experiment 2, trials were terminated if gaze left a 2{degree sign} window, but given that the target was only 5{degree sign} from fixation, this seems a relatively loose criterion. It would be important to show the distribution of gaze positions during the task to assess whether fixation control was adequate.

      We thank the reviewer for this great suggestion. We have now added a figure to the supplementary materials (page 23) that shows the probability density of gaze position throughout the ‘shift’ period, for left cued trials and right cued trials separately, please see Supplementary Figure S6. We hope that this will further show that even in Experiment 1, fixational control was successful. We also show the difference between left cued and right cued trials, which again shows a gaze bias towards the cued item.

      (8) Did the authors examine whether there was a response time benefit (e.g., RT in congruent microsaccade trials minus RT in incongruent microsaccade trials, as in Brandolani et al., 2025) or an accuracy benefit when microsaccades were directed toward the target?

      We thank the reviewer for their interest in our findings and for the great suggestion of an additional analysis. As we discussed also in response to the general summary from reviewer #2 above, we have now investigated whether trials with a congruent, incongruent or no saccade in the shift window (where congruent or incongruent was determined as based on the first saccade within the shift window) have, on average, different reaction times. This analysis did not show significant differences between these three conditions (congruent microsaccade, incongruent microsaccade, no microsaccade).

      In interpreting this observation, we would like to again stress how our experiment is not particularly well-suited to this analysis, as the amount of time between cue onset and the target events are highly variable across trials. Because of this clear drawback, we decided not to include these analyses. However, please note that we did now include the outcomes of another analysis that more directly targeted the relation between the spatial modulations in microsaccades and task performance, as we turn to below. 

      (9) Was there a relationship between the size of the attentional effect and the magnitude of the microsaccade bias?

      We thank the reviewer also for this insightful question. We have investigated the relationship between the magnitude of the microsaccade bias during the ‘shift’ period and the behavioural benefit. We have done this separately for a response time benefit and an accuracy benefit. Experiment 1 shows a significant correlation for both reaction times and accuracy with the magnitude of the microsaccade bias, but for Experiment 2 both of these relationships did not survive. Because this relationship did not prove robust across both experiments, but is nonetheless a set of findings our readers will likely be interested in, we have included this figure in the supplementary materials (page 22). Please see Supplementary Figure S5.

      (10) The criteria for minimum microsaccade amplitude and duration are not specified. This should be clarified. I recommend excluding events smaller than ~5 arcmin, as these are likely noise-especially since eye tracking was monocular. Monocular "microsaccades" can be spurious, but this can be determined only with binocular tracking. It is also unclear whether subjects used chin/head rests. A main-sequence plot in the supplementary material would be helpful.

      We thank the reviewer for pointing this out. We have included a main-sequence plot in the supplementary materials (page 23). The main-sequence plot can also be found in Supplementary Figure S7 and suggests that our saccade-detection algorithm worked well with detected saccades following the main sequence. We have also stated more clearly in the methods section that subjects used a chinrest (page 11).

      (11) Please specify the asterisk convention in figure captions (i.e., what * vs. ** vs. *** correspond to in terms of p-values).

      We thank the reviewer for pointing out that these significance levels were not mentioned in every figure caption, so we have added this information to every figure caption where they were missing (page 4, 6, 7 and 19).

      (12) The fact that stricter fixation criteria reduced the size of the effect suggests the possibility that gaze drift toward the target might have conferred an eccentricity advantage in this discrimination task. A direct comparison of the effect in the two experiments would be valuable. The authors should comment on this. It would be informative to plot the average gaze position around the time of peak microsaccade rate in both experiments. Reanalyzing the data post hoc with a stricter trial-selection criterion (e.g., excluding trials where gaze deviated more than 1{degree sign} from fixation) could also be very valuable, as it would systematically test how fixation control influences the observed microsaccade-attention relationship. This would be informative for the community studying this topic.

      We thank the reviewer for these valuable reflections. We have now added a direct statistical comparison of the main effect in the two experimental conditions (page 5: “When directly comparing this spatial saccade bias between experiments, we observed that the effect was significantly larger in Experiment 1 than in Experiment 2 from 535 to 745 ms (cluster p=0.013) and 1153 to 1309 ms after cue onset (cluster p=0.029)”).

      Regarding the average gaze position around the time of peak microsaccade rate: in response to reviewer #1, under point (7), we have included Supplementary Figure S6 that shows the probability distribution of gaze position throughout the ‘shift’ period (the same figure is found on page 23 in the article), which shows that fixational control was successful in both experiments. This period is also the period of peak microsaccade rate in both experiments.

      We wholeheartedly agree that systematically investigating the effect of fixational control is important for the field as a whole, and this is also precisely why we set out to perform the same experiment in two different experimental settings with regards to fixational control, and why we decided to include the results from both experiment variants side-by-side in our article.

      Reviewer #2 (Recommendations for the authors):

      In addition to my general concerns in the public review, I have the following recommendations.

      (1) Did the authors distinguish between the initial and subsequent microsaccades during their analysis? Is it possible to produce multiple microsaccades when shifting attention, or do the authors only consider the first microsaccade to be linked to an attention shift?

      We thank the reviewer for pointing out this ambiguity. We have now stated more clearly in the methods section that we consider all microsaccades for our analyses (page 12: “Crucially, we did not restrict our analyses to initial saccades; rather, all detected saccades were included. This allowed us to examine oculomotor behaviour during later trial phases, where initial saccades are unlikely to occur.”). We also believe this methodological choice is important, as otherwise it would be conceivable that no microsaccade bias can be found during the ‘sustain’ period, simply because no ‘first’ microsaccades occur anymore.

      (2) Two microsaccades had to be separated by at least 100 ms. This is an unusually long delay.

      Could the authors please specify how many microsaccades were discarded using this criterion?

      This inter-saccade-interval is quite large on purpose, as we want to minimise the probability of counting the same microsaccade twice. We have re-analysed the data with a minimum ISI of 50 ms, and this led to an increase of found saccades of a, respectively, 5.1% and 1.9% increase for Experiments 1 and 2. However, two participants in Experiment 1 led to a much higher increase in saccades than all other participants (these participants had z-scores of 3.9 and 2.4 for the number of additionally found saccades with an ISI of 50 ms; all other z-scores for Experiment 1 were between -0.5 and 0.5). When those two participants were removed, in Experiment 1 only 1.6% more saccades were found.

      (3) If I understand correctly, the authors did not use staircase procedures to eliminate differences in task difficulty between participants. Could the authors demonstrate how task difficulty relates to the link between microsaccades and performance? For example, is the time course of the microsaccade direction bias correlated with performance?

      We thank the reviewer for this suggestion (that overlaps with a comment of Reviewer 1). We have investigated the relationship between the magnitude of the microsaccade bias during the ‘shift’ period and the behavioural benefit. We have done this separately for a response time benefit and an accuracy benefit. Experiment 1 shows a significant correlation for both reaction times and accuracy with the magnitude of the microsaccade bias, but for Experiment 2 both of these relationships did not survive. Because this relationship did not prove robust across both experiments, but is nonetheless a set of findings our readers will likely be interested in, we have included this figure in the supplementary materials (page 22). Please see Supplementary Figure S5.

      (4) The authors reported using equiluminant stimuli. Could the authors please specify the exact luminance?

      We thank the reviewer for pointing out this missing information. We have now included this information in the methods section (page 11: “, with a luminance of 88.5 cd/m<sup>2</sup>.”). We have also included the luminance of the background (page 11: “luminance: 29.0 cd/m<sup>2</sup>”).

      (5) Could the authors please provide a full polar plot showing all microsaccade directions, and colour-code those included in the analysis?

      We thank the reviewer for this great suggestion on how to present our results even more clearly and comprehensively. We have included a supplementary figure showing the full polar histograms (with 20 radial bins), for all three timeframes of interest: the whole trial, the ‘shift’ period and the ‘maintain’ period (page 20). As requested, the saccades included in the main analyses are colour-coded. See Supplementary Figure S3

      (6) Can the authors please directly compare the main effects between experiment 1 and experiment 2 (Figure 2B)?

      We thank the reviewer for this great suggestion (that was also made by reviewer 1). We have now added a direct statistical comparison of the main effect in the two experimental conditions (page 5: “When directly comparing this spatial saccade bias between experiments, we observed that the effect was significantly larger in Experiment 1 than in Experiment 2 from 535 to 745 ms (cluster p=0.013) and 1153 to 1309 ms after cue onset (cluster p=0.029)”).

    1. eLife Assessment

      The authors describe a cell-specific mechanism by which glutamate transporters regulate the fidelity with which T-stellate cells in the mouse ventral cochlear nucleus relay information from auditory nerve inputs. The study is supported by solid electrophysiological data. It provides valuable insights into how the rapid binding of glutamate to transporters shapes auditory information processing at specific synapses.

    2. Reviewer #1 (Public review):

      In this article, the authors investigate how glutamate transporter function regulates excitability and synaptic coding in T-stellate cells in the mouse ventral cochlear nucleus. They test this in acute brain slices using whole-cell electrophysiology and artificially raise the relative local concentration of glutamate via pharmacological inhibition of transporter proteins. The main finding is that when sub-saturating doses of DL-TBOA are applied, cells become much more sensitive to synaptic input, diminishing the normally high fidelity of EPSP-spike coupling in these neurons. Notably, high-frequency stimulation in the presence of DL-TBOA reveals a large and slowly decaying AMPA receptor component that underlies persistent/rebound firing in earlier recordings. These effects are not seen in other ventral cochlear neurons, suggesting that rapid glutamate clearance in T-stellate cells, particularly, is important for auditory intensity coding. Overall, these experiments are well-performed, and the findings are robust, though there are some aspects that could be expanded to make the work more impactful. These include a better understanding of the relative contribution of neuronal vs glial transporters and an ability to separate the relative contributions of tonic glutamate concentrations in the cleft vs changes in membrane potential in action potential output. Additionally, there were some minor issues of clarity in both the figure presentation and the main text language that should be addressed.

      Major Points:

      (1) Given the dramatic effect of saturating DL-TBOA on tonic leak/RMP and that the sub-maximal concentration used in most of the experiments still varied between 25-50 uM, Figure 1 would be strengthened substantially by a dose-response curve. Ideally, 5 or 6 concentrations, plotting the effect on tonic current or RMP increase.

      (2) Examining the contribution of glial (EAAT1/2) vs. neuronal (EAAT3) transporters (Fig 8) is intriguing but comes across as incomplete here, especially given the small number of recordings. Using a different non-selective EAAT inhibitor (TFB-TBOA) to chase the EAAT1/2 blocker combo seems like an odd choice, given that you have already characterized the effects of DL-TBOA well. One could also try a lower concentration (~50-100 nM) of TFB-TBOA since it is somewhat selective itself for glial EAAT1/2. Given the data presented, neuronal transporters (presumably EAAT3) appear to dominate the rapid clearance of glutamate at this synapse, but this point isn't emphasized or explored sufficiently.

      (3) Separating the effects of depolarization vs. glutamate clearance was never explored. What effect does depolarizing the cell ~10 mV in control conditions (i.e., without TBOA) have on AP number/fidelity during synaptic stimulation experiments? The authors state that submaximal DL-TBOA generally causes no more than a 5 mV change in RMP, but tonic depolarization could also influence spike fidelity. This experiment could demonstrate that the increase in excitability during/after stimulation is not due to increased engagement of voltage-gated channels.

    3. Reviewer #2 (Public review):

      Summary:

      This manuscript addresses an important and mechanistically interesting question: whether plasma membrane glutamate transporters contribute only to slow clearance of ambient glutamate or whether they can rapidly shape synaptic signaling during high-frequency auditory activity. This manuscript provides important evidence that EAAT-mediated glutamate uptake is not merely a slow background clearance mechanism but is essential for maintaining reliable synaptic transmission and linear stimulus-intensity coding in ventral cochlear nucleus T-stellate cells during sustained auditory nerve activity.

      Strengths:

      The finding that EAATs may be required for rapid, local control of glutamate during high-frequency auditory nerve activity is interesting and could have broad relevance to auditory processing. The electrophysiological evidence is generally strong, particularly the use of patch-clamp recordings, stimulus trains, partial versus complete EAAT blockade, and comparison with bushy cell/endbulb synapses. The comparison between T-stellate cells and bushy cells/endbulb synapses strengthens the manuscript. The authors demonstrate that EAAT blockade disrupts coding in T-stellate cells but has little effect on bushy cell spike transmission, supporting a cell-type- and synapse-specific role of glutamate uptake.

      Weaknesses:

      However, some mechanistic conclusions, especially the specific contribution of neuronal versus glial EAATs and the absence of glutamate crosstalk between auditory nerve inputs, rely mainly on pharmacological and indirect electrophysiological inference and would be strengthened by additional anatomical, genetic, or direct glutamate-sensing evidence.

      (1) Clarification of DL-TBOA concentration.

      The authors used bath application of 200 µM TBOA and 25-50 µM in the other experiments, stating that "sub-maximal concentrations (25-50 µM)". The authors should provide a clearer rationale for why different concentrations were used across experiments rather than a fixed concentration.

      The reversibility of DL-TBOA effects should be demonstrated by washout experiments. In addition, potential off-target effects of DL-TBOA on postsynaptic receptors, intrinsic membrane excitability, or presynaptic release (e.g., PPR measurement) should be carefully considered. It would also be useful to test the effects of the submaximal DL-TBOA concentrations (25-50 µM) on membrane potential and inward currents, shown in Figure 1, to determine whether these concentrations depolarize the membrane potential in current-clamp mode or induce inward currents under voltage-clamp conditions.

      (2) Potential contribution of altered intrinsic excitability.

      In Figures 3B and 3C, DL-TBOA appears to induce additional action potentials even immediately after the first stimulation, whereas Figures 6 and 7 suggest that the first EPSC is not substantially altered. This raises the possibility that the enhanced firing may partly result from a modest depolarization caused by background glutamate accumulation or from other changes in intrinsic membrane properties after drug treatment. To address this, the authors should provide a quantitative analysis of physiological parameters under submaximal DL-TBOA conditions, including spontaneous action potential frequency, resting membrane potential, input resistance, and spike threshold.

      (3) Spillover/ crosstalk between AN-fiber-synpases.

      The authors should provide more explanation of how altering the number of active auditory nerve fibers demonstrates the absence of glutamate spillover/crosstalk between bouton synapses. Strong stimulation likely recruits more AN fibers, but it may also change release probability, axonal synchrony, or stimulation spread. The authors should more clearly justify the interpretation that strong stimulation recruits additional independent AN fibers rather than altering release probability or activating fibers with different intrinsic properties.

      (4) Interpretation of glial versus neuronal EAAT contributions.

      The authors claim that both neuronal and glial transporters contribute to rapid uptake using pharmacological approaches. The pharmacological data demonstrate that glial EAATs play a major role in glutamate clearance at T-stellate cell synapses. The strong increase in EPSC decay time and synaptic charge after UCPH-101/DHK application supports the conclusion that glial transporters contribute substantially to limiting glutamate accumulation during sustained auditory nerve activity. However, the conclusion that neuronal EAATs contribute directly should be stated with some caution. The evidence for neuronal EAAT involvement is indirect and depends on the pharmacological specificity and completeness of glial EAAT blockade. The conclusion would be strengthened by additional evidence, such as EAAT subtype expression/localization in T-stellate cells or auditory nerve terminals, transporter current recordings, immunohistochemistry, or genetic manipulation of neuronal EAATs. In addition, fitting the decay phase with a double-exponential model may help determine whether glial and neuronal EAATs contribute over distinct temporal windows.

    4. Author Response:

      We are grateful for the careful and extensive reviews, and are pleased that the reviewers found the work of broad interest to sensory processing. Please find our proposal for revision based on public reviews:

      Reviewer 1

      1) Request for dose-response curve for DL-TBOA and leak current or RMP. We can provide this, at least for the initial phase of the curve relevant to the concentrations used for synaptic experiments. Prolonged exposure to higher concentrations leads to very large cationic currents (through AMPAR) which appear to be damaging to membrane integrity.

      2) We will increase the N for glial vs neuronal block with the blockers we already used; this seems more practical than doing new experiments with different concentrations of TFB-TBOA. 

      3) We can include data to test the effect of blockers or small depolarizations on excitability.

      Reviewer 2

      1) We differentiated experiments with “25-50 uM” from 200 uM DL-TBOA because the higher concentration clearly led to massive AMPAR activation and depolarization block, as shown in Fig 1. We then chose lower concentrations to minimize background current while allowing glutamate build-up during exocytosis.  We felt we were clear on this point. 

      As to reversibility and “off target effects” like synaptic changes, we will provide this information. See also response to Reviewer 1, comment 1.

      2) See response to Reviewer 1, comment 3.

      3) We are certain that increasing stimulus strength increases the number of stimulated fibers, and this is well accepted. The stimulus electrode is placed in the auditory nerve root, well away from recorded cell and synapses, minimizing current spread to synapses. We can compare PPR for weak and strong stimuli in our current dataset to confirm no effects on release probability. As to variations in the intrinsic properties of myelinated auditory nerve fibers and their sensitivity to stimulation, there is no information about this, and do not understand how it would be relevant, particularly in as much as we report a negative result: no difference in blocker effect with small or large numbers of fibers active. The 3 main types of myelinated auditory nerve fiber, Type 1a,b,c, are known to respond to different sound thresholds, but that is a synaptic issue in the inner ear, and apparently not related to the myelinated axon bundle.

      4) We appreciate the reviewer's caution about a role for neuronal transporters and will revise accordingly.  We cited molecular evidence for expression of subtypes in the pre and postsynaptic neurons and in glial cells. Of course, given how ubiquitous such expression is across the brain, we suspect the kinds of experiments we provided offer more direct evidence for function.

    1. eLife Assessment

      This work provides a valuable contribution by leveraging simulation-based inference to investigate candidate compensatory mechanisms in neuronal network models and offering new insights into how distinct pathological perturbations may require different interventions. The computational evidence is convincing and has been strengthened by additional reproducibility analyses, although some aspects of inference validation and the biological interpretation of posterior dependencies warrant further investigation. The study establishes a helpful computational framework for exploring disease-specific compensatory mechanisms and will be of broad interest to the computational and systems neuroscience communities.

    2. Joint Public Review:

      Summary:

      This manuscript couples a 32-parameter model with simulation-based inference (SBI) to identify parameter changes that can compensate for three canonical hyperexcitability perturbations (interneuron loss, recurrent-excitatory sprouting, and intrinsic depolarisation). The study demonstrates a careful implementation of SBI and offers a practical ranking of "compensatory levers" that could, in principle, guide therapeutic strategies for epilepsy and related network disorders.

      Strengths:

      (1) By analysing three mechanistically distinct hyper-excitable regimes within the same modelling and inference framework, the work reveals how different perturbations require different compensatory interventions.

      (2) The authors adopt posterior estimation to systematically rank the efficiency of different mechanisms in balancing hyperexcitability.

      (3) Code and data are available.

      Comments on revised version:

      I appreciate the authors' extensive efforts in revising the manuscript and responding to the previous review. The revised version is substantially improved in clarity, organization, and presentation. In particular, the addition of schematic figures, the reorganization of the Methods section, the improved explanation of the model, and the inclusion of replication analyses all strengthen the manuscript.

      The manuscript remains entirely computational, and therefore its conclusions should be interpreted as predictions generated by a specific model rather than validated biological mechanisms. I believe the work has the potential to make a useful methodological contribution. However, several concerns remain regarding validation, interpretation of inferred posteriors, organization of the manuscript, and presentation.

      Major comments:

      (1) The manuscript states that simulation-based calibration showed the amortized posterior estimator was unreliable (85-88), but these results are not shown. The manuscript explicitly states that simulation-based calibration demonstrated substantial failures of the amortized posterior estimator, yet the corresponding analyses are not presented. Since these results motivate the transition to sequential NPE and are central to assessing inference reliability, they should be reported quantitatively, either in the main text or supplementary material.

      (2) The authors present two independently trained estimators and show strong agreement between them. This is a useful robustness analysis. However, the rebuttal occasionally presents this as addressing concerns regarding cross-validation and generalization. The new analysis does not constitute cross-validation in the usual sense and does not directly assess generalization to held-out targets or posterior accuracy.<br /> I recommend that the authors explicitly describe Figure 4 as a reproducibility analysis and avoid presenting it as a substitute for validation.

      (3) Posterior correlations are useful for generating hypotheses about compensatory mechanisms, but they should not be interpreted as direct evidence of compensation. The compensatory interpretation should instead be supported by the perturbation analyses (e.g., Figure 6), which provide mechanistic validation.

      The manuscript consistently treats posterior correlations and conditional posterior shifts as direct evidence of compensatory mechanisms. These are consistent with compensatory mechanisms, but they do not by themselves establish that the corresponding biological parameters causally compensate for the perturbation. I recommend clarifying this distinction and emphasizing that the conditional posterior analyses generate hypotheses regarding compensation, which are then partially supported by the perturbation experiments shown later in the manuscript.

      The language throughout the manuscript should therefore be softened.

      (4) The manuscript repeatedly suggests that the inferred conditional distributions may be useful for identifying precise interventions or guiding personalized treatments (examples include lines 24-29, lines 217-223, lines 242-246, lines 277-282, lines 283-286). These claims go beyond what is directly demonstrated.

      The study does not evaluate treatment outcomes, patient-specific inference, intervention efficacy, or clinical decision-making. Rather, it demonstrates differences in inferred parameter distributions within a computational model. While these results are valuable and may generate clinically relevant hypotheses, they do not yet establish predictive utility for treatment selection or precision medicine. I therefore recommend substantially softening these translational claims and emphasizing that the current findings generate hypotheses that could be tested experimentally in future work.

      (5) The revised manuscript still mixes presentation of findings with interpretation.

      For example, lines 217-226 largely continue to describe findings from Figure 6 and would fit better in the Results section. The Discussion would be strengthened by focusing more exclusively on biological implications, limitations, and future directions.

      A similar issue appears later in the discussion comparing posterior correlations and conditional distributions. Much of this section effectively reinterprets Figures 2 and 3 rather than discussing broader implications.

      (6) The discussion around lines 271-282 overstates what can be concluded from the inferred posteriors.<br /> The statement that correlations "discover broadly applicable mechanisms" whereas conditionals "identify specific mechanisms" is stronger than the presented evidence supports. Likewise, the conclusion that conditional distributions are more useful for precision treatments is speculative and not directly demonstrated.

      I recommend reformulating these statements as interpretations or hypotheses rather than conclusions.

      (7) Around line 84, the manuscript introduces q(theta|x) without clearly defining θ, x, or q. Readers unfamiliar with SBI may struggle to follow the notation. All quantities should be defined when first introduced.

      (8) The manuscript equates larger KS distances between conditional posteriors with greater compensatory potential. While KS distance provides a useful measure of posterior redistribution, it is not obvious that it should be interpreted as a measure of biological efficacy.

      (9) The manuscript would benefit from a discussion of parameter identifiability. The inference problem maps 32 model parameters to 7 summary statistics, implying substantial non-identifiability. While complete identifiability analysis is likely beyond the scope of the current work, this limitation should be discussed explicitly.

      All in all, the revised manuscript is significantly improved and addresses several concerns raised in the previous review. However, important issues remain as discussed above.

    3. Author response:

      The following is the authors’ response to the original reviews.

      Joint Public Review:

      Summary:

      This manuscript couples a 32-parameter model with simulation-based inference (SBI) to identify parameter changes that can compensate for three canonical hyperexcitability perturbations (interneuron loss, recurrent-excitatory sprouting, and intrinsic depolarisation). The study demonstrates a careful implementation of SBI and offers a practical ranking of "compensatory levers" that could, in principle, guide therapeutic strategies for epilepsy and related network disorders.

      Strengths:

      (1) By analysing three mechanistically distinct hyper-excitable regimes within the same modelling and inference framework, the work reveals how different perturbations require different compensatory interventions.

      (2) The authors adopt posterior estimation to systematically rank the efficiency of different mechanisms in balancing hyperexcitability.

      (3) Code and data are available.

      We thank the reviewers for their positive comments on our manuscript.

      Weaknesses:

      (1) A highly dense presentation of the simulated models and undefined symbols makes it hard for readers outside the modelling community to follow the biological message. An illustration of the models, accompanied by some explanations and references to the main equations and parameters discussed in this paper, would make the first section much more straightforward.

      Thank you for this feedback. To clarify our methods, we have added Figure 7, which illustrates the dynamics of the point neurons and their synapses. We have also added explanations and definitions of variables right where they appear. These variables were previously defined only in a table on a different page.

      We also moved the methods section to the back of the paper, as is common in many modern manuscripts. We hope that relegating method details to the end makes the manuscript more accessible.

      (2) This methodology appears to be a brute-force approach, requiring millions of simulations to tune 32 parameters in a network of 500-700 cells. It isn't scalable. Moreover, the authors did not use cross-validation, which, with a relatively low increase in computational cost, would provide a quantitative measure as to how well it generalizes; this combination raises doubts about both scalability and reliability.

      Scalability is indeed a key challenge of SBI methods. Amortized neural posterior estimation (NPE) is a brute-force approach in that it samples solely from the prior distribution, which is extremely wide. Many of these samples are therefore not very informative for the biologically plausible dynamics we are interested in, which is a downside of amortized NPE. However, amortized NPE is extremely scalable because once the estimator is trained, it can estimate the parameter distribution of any given output dynamic. We tried to build an amortized NPE for our simulator, but simulation-based calibration (a method to validate posterior estimates using additional simulations) showed that the estimators were unreliable.

      Sequential NPE is not a brute-force approach because it samples from posterior estimates, which are narrower than the prior. Because the amortized NPE failed, we use sequential NPE to create the two estimators for the baseline and the hyperexcitable condition described in the paper. While millions of prior samples are used to generate the initial posterior estimate, which is then sequentially refined, the sequential refinement requires only 80,000 additional simulations. This requires a significant amount of computational resources, which is why we consider the results worth reporting, but we make the simulator, the simulation results, and the trained estimators available, so other researchers can use or train their own estimators without running millions of simulations. We hope our rewrites make the advantages and disadvantages of the approach clearer.

      Regarding reliability and cross-validation, we agree that our initial submission has fallen short. We presented results from only one density estimator per condition, which we considered sufficient given the large number of samples. In the revised version, we present the key results from two additional density estimators trained on partially new training data (Figure 4).

      (3) Several parameters remain so broadly distributed after fitting that the model cannot say with confidence which specific changes matter. Therefore, presenting them as "compensatory levers" is somewhat questionable.

      It is indeed difficult to determine which changes matter because of the simulator’s complexity. Especially the marginal correlation coefficient (Figure 2 C) are small and the pairwise histograms are broad (Figure 2A), because all other parameters are unconstrained. But the conditional correlation coefficients are larger (Figure 2D) and narrower (Figure 2B). We have added the histograms in Figure 2B in the revised version to highlight the difference. We cannot provide a definitive threshold for correlation coefficients to discriminate between important and unimportant mechanisms. Therefore, compensatory mechanisms discovered with SBI should be validated mechanistically, as we do in Figure 5.

      (4) Every conclusion is drawn from simulated data; without testing the predictions on recordings, we have no evidence that the proposed interventions would work in real neural tissue. Because today we cannot diagnose which of the three modelled pathological regimes is actually present in vivo, the paper's recommendations cannot yet be used to guide therapy.

      This is indeed an unfortunate drawback of our current work. We are working to apply this approach to constrain microcircuit simulators with data from epilepsy patients. But that work is currently ongoing and will not fit into the present manuscript.

      Recommendations for the authors:

      Beyond the issues I wrote above, which are methodological, I would like to raise my concern about the way this manuscript is written:

      We highly appreciate this editorial feedback on clarity and style. Such feedback is rare and we have worked to address each point to improve the manuscript.

      (1) Paragraphs - several paragraphs start with: "To quantify/identify/find specific compensatory mechanisms of hyperexcitability with simulation-based inference". It is a good idea to orient the reader with the specific goal of each section, but it is not helpful to repeat the overall message of the paper in every paragraph. Several paragraphs open with "However," or "Additionally,". Please restructure sentences so that connectors appear after a clear topic sentence.

      We have done major rewrites to improve the readability of our manuscript. We start paragraphs with more specific context sentences, rather than the broad research goal, and also made paragraphs much shorter with clearer main messages.

      (2) Section 1: The way you present NPE, it would seem like it's specific to neuroscience (and it's not). The paragraph starting at line 26 is not clear. Please revise it. Line 29 - missing a "." before the next sentence begins. Avoid phrases like "for the longest time".

      We now stress that NPE, like SBI, is used across scientific domains.

      (3) Section 2 was tough to read. Please present each equation in its own numbered display, followed immediately by a plain-language explanation of every symbol and parameter. Provide an illustrative diagram: a small schematic of the AdEx neuron, synaptic connections, and the three perturbations. Even a simple block figure will orient nonexperts. Keep critical methodological decisions (priors, summary statistics, simulation length) in the main text, but move voluminous tables of parameter bounds, learning rates, and hardware specs to the supplement. Remove mentions of which Python functions you used. Readers care about algorithmic choices, not function names. Please reserve specific code references for the GitHub README.

      We have added schematic panels at the beginning of Figures 3 & 4 and added Figure 7, which illustrates the neuron and synapse models of the simulator. We also made major rewrites to the methods section to remove programmatic implementation details and define variables where they appear.

      In general, I think it would be a good idea to have an editor to polish syntax, verb tense consistency, and punctuation. A thorough language edit will improve the readability and impact of this manuscript.

      We have attempted to improve the points raised by the reviewer. In particular, we have carefully rewritten verb tense and punctuation throughout the revised manuscript.

    1. eLife Assessment

      This important study reports a novel phenomenon of maternal growth during pregnancy that is independent of growth hormone (GH), adding a new dimension to maternal biology of reproduction. The evidence is convincing and supported by state-of-the-art methodologies conducted in mice and persuasive observations in humans with hereditary isolated GH deficiency. Revised discussion should focus on possible mechanisms, including the role of IGF2, and on directions for future research.

    2. Reviewer #1 (Public review):

      This work evaluates the impact of reproductive history on growth, body weight and body composition in mammals. In mice, somatic growth is stimulated by the first pregnancy while the second pregnancy increases body weight mainly by increasing adiposity. To probe the role of pituitary growth hormone (GH), the key regulator of somatic growth in these processes, was addressed by comparing the impact of reproduction on growth in normal ("wild type") and genetically GH-deficient females and by detailed characterization of the profile of fluctuations in circulating GH levels in both types of animals. Additional studies addressed the possible role of other endocrine pathways (ghrelin and estrogen) in the pregnancy-related growth. Surprisingly, reproduction-related growth was independent of GH, ghrelin and estrogen. To determine whether these results may apply ("translate") to human physiology, data on various parameters of somatic growth were collected from women with hereditary GH deficiency. The findings indicate that GH-independent stimulation of growth by reproductive events also occurs in women.

      Use of multiple animal models, rigorous characterization of GH levels in normal and GH-deficient females, and inclusion of data derived from a unique and well-characterised population of people with hereditary isolated GH deficiency and no GH replacement therapy are important strengths of these elegant and innovative studies. The results address a broader and clinically significant issue of permanent changes in body size, composition and function that result from pregnancy and lactation. This work also provides important background for further studies aimed at the identification of the mechanism involved and the role of specific reproductive events in the regulation of growth.

    3. Reviewer #2 (Public review):

      This manuscript describes the fascinating phenomenon of growth hormone (GH)-independent growth occurring in the mother during pregnancy. This growth was most pronounced in dwarf mice that are lacking the receptor for growth hormone-releasing hormone (GHRH) and therefore showing isolated GH deficiency. However, the pregnancy-induced growth could also be observed in wild-type mice, suggesting that it is a normal part of the maternal adaptation to pregnancy. The study falls short of identifying the mechanism(s) driving this pregnancy-induced growth response, but it certainly reveals a novel insight into maternal physiology. The authors have completed a range of experiments in mice to prove that, as well as being GH independent, the pregnancy-induced growth also did not require GH signaling in the liver (i.e. not another pregnancy-specific ligand operating through the GHR to promote IGF). They also provided complementary data from a population of humans with untreated isolated GH deficiency that are broadly consistent with the hypothesis. While it is important to consider the significant species differences between rodents and humans, both in terms of growth physiology and also in terms of evolution of placental somato-mammotrophic hormones, this unique population are a valuable resource and adds credence to the study. Overall, I find this a compelling research story, but disappointingly unfinished. There are some areas where additional information could improve the ability to interpret the data, and some additional concepts that could be considered in the discussion. There are also areas where additional experiments might provide important insights. However, I think that such suggestions can be considered as appropriate for future research, rather than delaying consideration of the current manuscript.

      Main comments:

      (1) Data in Figure 1 are remarkable - not so much the growth in pregnancy in the wildtype mice, because while elevated GH is well known in pregnancy, but growth in the dwarf mice is indicative of GH-independent growth. From these data, it seems that there is good evidence that growth in pregnancy is an adaptive function. However, it is possible that growth is achieved in dwarf mice and that in wildtype mice may have been mediated through different mechanisms. The dwarf mice showed an increase in liver and plasma IGF1, suggestive of an additional ligand driving IGF in pregnancy. One could hypothesize that such an effect could be mediated by an additional pregnancy-specific ligand activating the GH receptor. In humans, placental growth hormone could be such a ligand, but as far as we know, there is no placental GH in mice. In contrast, the wildtype animals showed suppression of liver and circulating IGF1, and low levels of pSTAT5 in the liver during pregnancy. These data (in Figure 5) are very surprising. Given the high circulating GH in pregnancy, as well as high placental lactogen (which would be expected to activate STAT5 in the liver through the Prlr), the low levels of pSTAT5 are unexpected and would seem to indicate some sort of acquired insensitivity to GH. Is this entirely driven by down-regulation of STAT5b protein, or could there be activation of other, negative regulators of STAT signalling, such as SOCS? What is causing such a profound suppression of STAT5? Regardless of the mechanism, this suggests that pregnancy-induced growth in wildtype mice is independent of circulating IGF1 (potentially a different mechanism or in addition to that seen in IGHD mice).

      The data shown in Figure 6 are a major strength of the study, showing that the pregnancy-induced changes are not specific to one particular transgenic model, but still occur in a variety of models affecting GH through different approaches. Given the pregnancy-specific nature of the changes, however, it seems an oversight not to have evaluated the role of placental lactogens. Prlr is highly expressed in the liver, but the function of this hormone in the liver is not well established. Could the extremely high levels of PL be mediating this growth response? Given the low expression of STAT5 in the liver and the fact that plasma IGF1 is not markedly elevated, it seems more likely that this growth response may be mediated by locally produced IGF1 in target tissues.

      I think these possibilities could be addressed by an expanded discussion of species variation in placental hormones, to highlight that humans have expansion of the GH locus, but rodents have expansion of the prolactin axis (see Soares, M. J. The prolactin and growth hormone families: pregnancy-specific hormones/cytokines at the maternal-fetal interface. Reprod Biol Endocrinol 2, 51, 2004). Importantly, placental GH and chorionic somatomammotropins (CSM) in humans are all variants of the GH gene, but CSM have preferential activity at Prlr. This seems to be a fundamental species difference in pregnancy biology, but has been interpreted as an example of convergent evolution, with conservation of prolactin and GH-like functions at the maternal-fetal interface, mediated by different mechanisms, likely contributing to the metabolic adaptations of the mother (see Newbern D, Freemark M. Placental hormones and the control of maternal metabolism and fetal growth. Curr Opin Endocrinol Diabetes Obes. 2011; 18: 409-416). While the preceding function has focused on explaining the evolution of placental lactogens (either prolactin or GH variants), the present data suggest that there are also mechanisms to maintain growth in pregnancy, independent of GH (even in the absence of a placental GH).

      (2) The human data are very interesting, and my initial impression was that it seemed unlikely to be the same phenomenon. Was there any real evidence for "growth" in pregnancy? Pubertal maturation of long bone growth might be expected to prevent further growth in adulthood. However, these issues were appropriately discussed, and it seems well justified to evaluate this unique population of women with IGHD who underwent pregnancy. It would be very interesting to know if these women experienced elevated IGF1 during pregnancy, indicative of placental GH contributing to growth. Mechanistically, this might be more like the dwarf mouse situation of IGHD, that the situation in wildtype mice (associated with liver insensitivity to GH and low IGF1).

      (3) It would be useful to include investigations that isolate the effects of pregnancy and the placental hormones. Such studies could include evaluating growth in pseudopregnant mice with IGHD (pregnancy-like changes in hormones but lacking the placental contribution) and in IGHD animals that experience pregnancy but not lactation (pups removed at birth). I accept that this might be too large an additional study to add for the present manuscript.

      (4) It is an important and translationally relevant observation that pregnancy increased the risk of long-term weight gain, and that after the first pregnancy, the pregnancy-induced growth response was more directed to promoting fat deposition. Does this provide any mechanistic insight? Could a metabolic adaptation result in growth?

    4. Reviewer #3 (Public review):

      Summary:

      The study describes an increase in body growth and body composition in both mice and women. In mice, the impact on growth is mainly seen during the first pregnancy, and the changes postpartum on body composition are also different during the first and second pregnancies. The study has used various knock-out models in the growth hormone axis to understand these changes as well as some gene expression analysis related to GH, IGF-1 and estrogen signalling pathways.

      Strengths:

      (1) The inclusion of various knock-out mouse models that allow for exploration of mechanisms related to the above-mentioned changes.

      (2) The investigation of gene expression of GHR, IGF-1R and ER pathways.

      Weaknesses:

      The human findings are dependent on the patient's recollection of bodily changes after their pregnancies.

      Conclusion:

      The authors have partly achieved their aim of describing changes in growth and body composition that remain after pregnancy and the mechanisms behind these changes. This study may have importance for a wide variety of research areas as well as in the clinical setting. The study is also unique in its attempt to bridge findings in mice to a unique human model of congenital GH deficiency.

    1. eLife Assessment

      In this valuable study, Zhang et al. investigated EEG neurofeedback as a method to modulate brain activity prior to painful stimulation and examined its effect on pain perception in a well-powered, double-blind study. Results showed that real, but not sham, feedback enabled learning-dependent enhancement of pre-stimulus α oscillations. However, the evidence for a neurofeedback-specific reduction in pain remains incomplete, as the current paradigm cannot distinguish between genuine neurofeedback effects and placebo effects. Nonetheless, this work is likely to be of interest to researchers in the fields of neurofeedback and pain.

    2. Reviewer #1 (Public review):

      Summary:

      Zhang et al. investigated EEG neurofeedback as a method to modulate brain activity prior to painful stimulation and its effect on pain perception. Neurofeedback was designed to train participants to upregulate alpha power contralateral to the site of painful stimulation. Real or sham neurofeedback was administered to two independent groups. Each group performed two tasks: one in which participants were asked to modulate their brain signals (training task) and another in which they were asked to passively watch the feedback (non-training task). The authors reported an increase in alpha power during real neurofeedback training compared with sham training and non-training conditions. The authors also reported a decrease in pain perception during the training task, both in the real and sham neurofeedback groups. Additionally, in an offline analysis, the authors investigated brain dynamics with microstate analysis during the neurofeedback training. Also, they implemented a mediation analysis to infer which brain responses to neurofeedback training mediated changes in pain perception.

      Strengths:

      (1) The research question is licit and sound. EEG neurofeedback is a promising non-invasive technique with the potential to alleviate at least the sensory component of pain. The rationale for applying neurofeedback at the alpha band in the somatosensory cortex is well justified by the alpha-gating theory in pain modulation.

      (2) The sample size is adequate to capture neurofeedback effects. The effort to conduct a double-blind study with a complex design paradigm and an adequate sample size is valuable and appreciated.

      Weaknesses:

      (1) Reported behavioral effects on pain reduction might be due to the placebo effect rather than neurofeedback, as pain ratings were reduced both in the real and sham neurofeedback groups during training. It is important that authors report this effect appropriately and disclose which information was given to the participants when they enrolled in the study, i.e., whether the paradigm was designed to reduce pain perception.

      (2) The utility of training effects, especially in the sham group, is unclear. I understand that including the non-training condition allows the distinction between neurofeedback effects and arousal effects. However, interpreting training effects should not be the point of this study. What does it tell us that participants who received sham stimulation increased or decreased alpha power in the training session vs the non-training session?

      (3) There might be hidden time effects (habituation/sensitization) on pain responses and/or on brain responses to neurofeedback. A within-session analysis comparing the first half of the training with the second half should be conducted to discard them.

      (4) Connectivity analysis reflects spurious effects. In EEG, deriving phase-based functional connectivity at the sensor level is problematic due to volume conduction effects. EEG functional connectivity should be performed after source reconstruction, and measures discarding instantaneous phase lags should be preferred, which is not the case with magnitude-squared coherence. See (Bastos and Schoffelen, 2015).

      Although neurofeedback is a promising technique for modulating pain perception, the current study adds limited novelty to the field, as its design could not disentangle whether behavioral effects (reductions in pain intensity and unpleasantness) were specific to neurofeedback training or due to non-specific effects (e.g., placebo). Nevertheless, the authors corroborated that brain states before painful stimuli could be modulated with neurofeedback (enhancement of alpha power).

    3. Reviewer #2 (Public review):

      Summary:

      This study uses neurofeedback to modulate alpha-band activity and examines how this influences pain-related processing. The question is timely and methodologically elegant, because it addresses whether noninvasive modulation of ongoing oscillatory activity can causally shape pain perception and/or expectation-related processes.

      Strengths:

      The use of neurofeedback as a tool to modulate alpha activity is a major strength, because it provides a noninvasive and conceptually clean approach to probing the functional role of oscillatory brain activity. The design is also attractive because it links neurophysiological regulation to a psychologically meaningful outcome, namely pain processing. Further, the induced changes were also related to different EEG microstates and ERP components during the processing of the pain stimulus, and therefore the authors demonstrate a clear relation between preparatory prestimulus states and stimulus processing.

      The manuscript appears to address an important and clinically relevant question, and the idea of testing whether alpha regulation can alter pain-related responses is of high interest for systems neuroscience and pain research.

      Weaknesses:

      Methodologically, it is unclear what alpha values were used in the analyses. It is stated that alpha was extracted within 2s windows of the 16s long feedback period. However, the values change across this period. Which value is used for the correlation with the pain ratings and all other analyses? Using the average across the 16s could reflect large values in the first half and low values in the final half, but for the relationship between alpha and pain, the last segments should be more relevant. If the initially elevated alpha activity subsides several seconds before the onset of the pain stimulus, it is difficult to see how it could influence subsequent pain processing.

      Related, after the 16s feedback period, a fixation period is used with a 3-5s length. If alpha band activity is relevant for the consecutive pain processing, the amount of alpha in this period should be relevant. The authors should demonstrate that the induced alpha change during the feedback period remains stable during the fixation period and that the activity in this period is related to pain processing.

      Further, it should be noted that the alpha band modulations related to alpha band training were accompanied by significant effects in other frequencies. Therefore, a clear relationship between alpha and behavioral pain ratings is not the only interpretation. Correlations with other frequencies or combinations of frequency band modulations should be incorporated to allow a more precise interpretation. Furthermore, in the sham feedback group, an increase in alpha band activity was observed (p=0.06), and the small difference in the pain intensity rating may be related to a clear outlier in the Sham group (Figure 4a).

      In both groups, a main effect of training, regardless of sham or real feedback, was reported with a small difference between groups. But the main modulator seems to be related to the instruction to modulate the neural activity, and this large effect should be discussed in more detail regarding, for example, possible attentional processes.

      A further central concern is that the visual feedback signal (the ball movement) may generate expectations that are not specific to alpha activity and that these expectation processes modulate the pain processing (ball down may indicate more pain). It is well known that intensity cues can generate expectations about upcoming perceptions, and the used feedback signal with an increasing or decreasing visual curve clearly signals what intensity should be expected. Therefore, it is important to show that the amount of positive (ball up) and negative visual displays is matched between the sham and real feedback group. Further, the authors should report whether the final ball position can predict the latter pain rating in both groups or differentially. Following this interpretation, alpha band activity is not directly related to pain processing but only serves as a signal that is transformed to a visual stimulus that then generates expectations.

      Finally, the manuscript would benefit from a more explicit analysis of whether individual alpha changes are related to pain ratings within each subject. If higher alpha is truly linked to reduced pain perception, this should be visible at the participant level during learning of the neurofeedback procedure. Relatedly, there is no learning period incorporated, and usually participants are not able to regulate their alpha activity from the first trial on. The authors should include an analysis of the development of alpha band activity over learning and a relation of these individual alpha values and the corresponding pain ratings.

      I cannot find a link to the preregistration in the current manuscript.

      In summary, a "causal" relation of alpha activity with pain perception -that is mentioned several times in the manuscript- is not fully supported by the present results

    1. eLife Assessment

      This important study investigates whether perceived gender is represented in the brain in a category-invariant manner across faces, bodies, and objects, identifying the right middle temporal gyrus (rMTG) as a potential locus. The evidence is incomplete due to major conceptual concerns, weak statistical methods, and unaddressed low-level confounds like stimulus size and motion. This work will be of interest to psychology and social neuroscience researchers in face and person perception literature, provided the authors temper their claims regarding abstract representation.

    2. Reviewer #1 (Public review):

      Summary:

      This manuscript investigates whether the human brain contains a shared category-general representation of gender across faces, bodies, and gender-associated objects. The authors acquired fMRI data while participants viewed male and female stimuli from three categories in a one-back task. They then used searchlight MVPA, cross-category decoding, regression-based RSA, CNN vs. brain representational comparisons, and PPI analyses. Their main finding is that gender information could be decoded from distributed occipitotemporal regions within each category, whereas a cluster in the rMTG showed convergence across cross-category decoding and RSA. The authors concluded that this rMTG representation resembles intermediate layers of fine-tuned CNNs and that face and body gender processing share similar functional connectivity patterns.

      Strengths:

      The question is potentially important, particularly for social cognition, object recognition, and the use of neural network models to interpret high-level visual representations. Previous behavioral studies have shown cross-category adaptation between bodies and faces, and even between gender-associated objects and faces, so the attempt to test for a neural counterpart using fMRI is well motivated. The use of multiple complementary analyses including within-category decoding, cross-category decoding, regression RSA, CNN comparisons, and effective connectivity analyses is also a strength. The convergence of cross-category MVPA and RSA in a right MTG cluster is potentially interesting and deserves attention.

      Weaknesses:

      The largest problem is conceptual. The term gender is used as if it refers to the same construct across faces, bodies, and objects. This is not self-evident. In faces and bodies, the stimuli seem to contain visual cues from which observers infer binary gender categories. In objects, however, the relevant information is almost gender stereotype, cultural association, or learned semantic association. These are not equivalent constructs. The manuscript therefore needs to distinguish much more carefully between perceived gender, biological sex cues, gender-associated visual features, and gender stereotypes. Without this distinction, the title and main conclusion are too broad. The object condition is particularly problematic. Javadi & Wee (2012) showed that gender-associated objects can bias subsequent judgments of ambiguous face gender, and they discussed two possible mechanisms, including shared neural substrates or top-down modulation induced by the gender concept. However, their behavioral adaptation study does not directly demonstrate that objects, faces, and bodies are encoded in the same neural representational format. The present manuscript treats these object stimuli as if they provide evidence about the same kind of gender representation as faces and bodies, but that step requires additional empirical support. Independent ratings of object gender association, cultural familiarity, visual similarity, and semantic category are essential here.

      A second major concern is stimulus control. The face images were taken from Chinese male and female actors, the body images were headless bodies in underwear, and the object images were selected because of prior gender associations. This design introduces many possible confounds: hairstyle, makeup, skin texture, body shape, clothing, color, luminance, object category, object function, curvature, spatial frequency, and cultural familiarity. Cross-category decoding can be significant even when a classifier relies on shared visual statistics rather than an abstract gender code. For example, female-associated stimuli may differ from male-associated stimuli in color, shape, brightness, texture, or semantic category in ways that are consistent across faces, bodies, and objects. The present analyses do not adequately rule out these alternatives. Foster et al. (2019) are especially relevant in this respect. They reported that body sex could be decoded from both body- and face-responsive regions. However, the sex of well-controlled faces, for example faces excluding hairstyle cues, could not be decoded from face- or body-responsive regions. This finding should make the authors more cautious. The fact that the present study used more ecological face stimuli may increase sensitivity to gender-related cues, but it also increases the possibilities that decoding is driven by uncontrolled external features rather than by an abstract gender representation. Accordingly, because no additional visual, semantic, or stereotype-based model RDMs were included in the RSA analysis, this result alone cannot establish an abstract, category-independent gender representation. Any systematic difference between male- and female-associated images will load onto the gender RDM. At least, the authors should include additional model RDMs for low-level visual features. In addition, the current RSA analysis has another limitation. The neural RDMs are based on only six condition-level patterns, producing a 6 × 6 matrix. The theoretical model includes only binary gender and category RDMs. This is too coarse to support the claim of category-independent gender representation. Ideally, all the RSA analysis should be performed at the item level rather than at the condition level.

      The cross-category decoding result in rMTG is promising but not yet conclusive. The authors identify a right MTG cluster by overlapping thresholded maps from three cross-category decoding analyses. This is useful descriptively, but it does not by itself establish a common representational code. The overlap of thresholded maps depends on the chosen threshold. If the authors want to make a formal conjunction claim, they should use a valid conjunction-null approach such as a minimum-statistic conjunction evaluated under the appropriate conjunction null, rather than simply displaying the intersection of thresholded maps. Even if this approach cannot be adopted in this study, the issue should be included as a limitation.

      In the PPI analysis, the reported similarity between face and body connectivity matrices is a little bit small (r = 0.08). The claim of a shared functional network should therefore be softened unless the authors test whether this correlation is significantly larger than the face-object and body-object correlations, correct for multiple comparisons, account for the non-independence of matrix elements, and report participant-level distributions and confidence intervals.

    3. Reviewer #2 (Public review):

      Summary:

      The study tests whether male/female-related information is represented in a form that generalizes across faces, bodies, and gender-associated objects. Using within- and cross-category MVPA, regression RSA, comparisons with fine-tuned CNNs, and connectivity analyses, the authors identify a right middle temporal gyrus region whose patterns generalize across the three stimulus classes. They conclude that this region provides a category-general, mid-level representation of gender and acts as a neural hub.

      Strengths:

      The question is novel and important, while the logic of the study is straightforward. Examining faces, bodies, and objects within the same participants provides a useful extension beyond the predominantly face-based literature. Cross-category decoding is also a stronger test of shared information than simple anatomical overlap between within-category maps. The combination of MVPA, RSA, computational modelling, and connectivity analysis is ambitious, and the replication of the CNN layer profile with both AlexNet and VGG16 is a useful characterization of relevant information.

      Weaknesses:

      (1) The construct labelled "gender" is not equivalent across stimulus classes. For faces and bodies, the male/female label is intended to track a property of the depicted person, albeit one inferred imperfectly from appearance; for objects, masculinity or femininity is not an intrinsic property of the object but a culturally contingent association that may vary across observers and contexts. Treating both as levels of a single binary factor risks conflating person-category information with gender-stereotypic object associations and interpreting their common neural discriminability as evidence for one abstract concept of gender. The term "object gender" could also be confused with grammatical gender in some languages (e.g., French or German).

      (2) The CNN analysis does not isolate the shared male/female component. The authors correlate the complete six-condition neural RDM with the complete CNN RDM. However, rMTG also carries substantial information about whether an image is a face, body, or object. Consequently, the peak correspondence with Conv4 may reflect category structure rather than the representation that supports cross-category male/female decoding. The current analysis does not establish that shared gender-related information specifically depends on mid-level features.

      (3) The connectivity interpretation is overstated. PPI measures task-dependent covariance; it does not establish information transmission, directionality, or an upstream-to-downstream processing sequence. The reported face-body connectivity similarity is also small (r=.08). Also, describing rMTG as a "hub" is not justified without network-centrality measures, lesion evidence, or causal perturbation.

      The authors partly achieve their aims. The results provide credible evidence that patterns in rMTG contain information that generalizes across binary male/female-labelled faces and bodies and masculine/feminine-associated objects. They do not yet establish a genuinely abstract representation of gender, a specifically gender-related correspondence with intermediate CNN layers, or a neural hub that transmits information through a directed network. With more precise framing and targeted reanalysis, the study could make a useful contribution to research on social vision and cross-category representation.

    4. Reviewer #3 (Public review):

      Summary:

      In this work, the authors investigate whether gender information is encoded in the brain in a way that is invariant to the object being perceived. They design an fMRI experiment in which 22 participants perform a one-back repetition detection task in a block design. Images shown are of three types (faces, objects, and bodies) and of two perceived genders, male and female. They perform MVPA, RSA, and functional connectivity analyses to determine whether gender information is invariant to the type of image being perceived. They report an area in the posterior right middle temporal gyrus (rMTG) that is found in their gender decoding analysis across categories. To confirm that this area encodes gender information, they perform a regression-based RSA with category and gender model RDMs, and report that the gender model RDM is significantly correlated with brain representations in that area. Finally, to further investigate the representations in this area, they perform a model-based RSA in which they first fine-tune a deep neural network for gender classification, and then study the correlation between model RDMs and brain RDMs. Consistent with a previous report in face processing (Jiahui et al., 2023), they find that gender information is more consistent with representations in middle-to-late layers of the networks. Additional functional connectivity and PPI analyses are reported to reveal differences in co-fluctuation of brain activity within occipital and parietal nodes when perceiving different types of male/female images. Based on these results, the authors conclude that rMTG represents gender information invariant of the category perceived, although rMTG also afforded decoding of category information.

      Strengths:

      Whether perceived gender is represented in a manner invariant to the category of the stimulus is a legitimate and interesting question, and one of relevance particularly to the face and person perception literature.

      The model-based RSA, in which RDMs from networks fine-tuned for gender classification are compared against brain RDMs, is an interesting approach, and the layer-wise profile the authors obtain converges with a previous report in the face processing literature (Jiahui et al., 2023).

      Weaknesses:

      A substantial number of inferences are drawn on the basis of weak statistical methods and a suboptimal design. My concerns are set out below, ordered by severity.

      (1) The statistical tests are not appropriate for classification and RSA, and are prone to false positives. Classification accuracies and RSA correlations may be positively biased, and the true null distribution may therefore be centered above the nominal chance level, or above zero in the case of RSA. Testing against a theoretical value with a one-sample t-test under these conditions inflates the false positive rate, especially with few test samples per classification, and does not afford valid population inference for information-like measures (Combrisson & Jerbi, 2015; Allefeld et al., 2016). The concern applies to every inferential claim in the manuscript, including the identification of the rMTG cluster on which the paper's central conclusion rests. The established remedy is permutation testing, in which the labels are randomly permuted and the full analysis, including cross-validation, is re-computed so that any bias is captured in the empirical null distribution (Stelzer et al., 2013; Etzel & Braver, 2013). This approach has been applied in comparable face-decoding studies using both classification and RSA (Guntupalli et al., 2017). I raise this methodological concern here because it is the clearest way to convey why the reported statistics cannot be safely interpreted at face value.

      (2) The decoding analyses do not appear to test generalization to left-out stimuli. From my reading of the design, each run contained all six conditions presented three times in random order, with each block containing 12 images (10 unique plus two repetitions serving as catch trials). If all images were presented in every run, the same images would be present in both the training and test sets of the cross-validation. Under these conditions, the interpretation of a general "gender" code is difficult to justify: the classifier may be exploiting low-level image features specific to the particular exemplars rather than gender per se. This bears directly on the paper's central claim, which concerns an abstract, category-invariant representation of gender, a claim that requires decoding to generalize to stimuli the classifier has not encountered.

      (3) There is no evidence that participants perceived the stimuli's gender as the authors assumed. Perceived gender may be subject-specific, yet no norming is reported establishing that participants actually rated or processed the stimuli according to the gender the authors assigned to each image. Some images are likely to be more ambiguous than others. This is a construct validity issue rather than an analysis issue: the class labels used throughout the decoding analyses, and the gender model RDM used in the RSA, both rest on an assumption about the participants' percepts that is never tested against the participants themselves.

      (4) The rMTG ROI reported in Figure 2c appears to overlap almost perfectly with the motion-sensitive area hMT+. The reported effects may therefore be driven, at least in part, by low-level motion signals arising from the rapid on/off changes of the stimuli and the associated optic flow. I am not claiming that the results are fully driven by this, but no control reported in the manuscript rules it out, and this region is the centerpiece of the paper's conclusion.

      (5) Stimulus size is confounded with category in the functional connectivity analyses. The authors report that functional connectivity differed between faces and objects, and between bodies and objects. However, faces and bodies were shown with the same visual extent, while objects were larger. Given that the nodes being investigated are in visual areas, it is unclear how these differences can be attributed to category rather than to the low-level difference in stimulus size. The same confound bears on the behavioral task performed within the scanner: participants can perform the one-back task more easily, simply by detecting size differences, since two images of different sizes are clearly not the same image, rather than by processing the image content. This affects what can be assumed about participants' attention to the stimulus category or gender.

      (6) No motion quality control is reported for the functional connectivity analyses. Functional connectivity is well known to be highly susceptible to head motion, yet the manuscript reports no summary of how much subject motion there was, no indication of whether volumes with excessive motion were removed or censored, and no account of quality control on the measured data more generally.

      (7) The use of famous faces introduces an avoidable confound. The face stimuli were famous faces. Famous and familiar faces are known to recruit substantially more widespread activity than unfamiliar faces, extending well beyond the core visual system (Gobbini & Haxby, 2007; Natu & O'Toole, 2011; Visconti di Oleggio Castello et al., 2017; Kovacs, 2020). For a study focused specifically on gender, this introduces a source of variance that unfamiliar faces would have avoided, and it complicates the comparison of the face conditions against the body and object conditions.

      (8) The rationale and benefit of fine-tuning the deep neural networks are not established. The manuscript does not report the original, non-fine-tuned accuracy of the models that required fine-tuning, so the benefit of the procedure cannot be assessed; given that the final validation accuracy is low, it is unclear that fine-tuning actually helped. AlexNet and VGG are trained for object classification on large datasets, and fine-tuning with 2,000 training images may not be sufficient to genuinely shift the objective. Whether the activation patterns and RDMs changed in any significant manner after fine-tuning is not reported, and the rationale for selecting the specific layers used is not stated.

      (9) Taken together, the analyses as presented do not establish the paper's central claim. My concern is not that the reported effects are necessarily absent, but that the combination of statistical tests that do not account for possible positive bias, a cross-validation scheme that may not guarantee generalization across stimuli, a key region that coincides with a motion-sensitive area, and gender labels that were never validated against participants' own perception leaves too many open questions for the results to be evaluated as they stand.

      (10) I would add one broader consideration. Perceived gender is likely to depend on culture and to vary across individuals. A binary male/female contrast in 22 participants, without evidence that those participants perceived the stimuli as the authors intended, is a narrow operationalization of a construct that is unlikely to be so simple. Even if the analyses were fully sound, caution would be warranted in generalizing from this design to claims about how the brain universally represents gender.

    1. eLife Assessment

      This important study extends a model of cortical normalization (ORGaNICs) to interacting cortical areas and shows that communication through coherence and communication subspaces can arise from a single set of dynamics. The evidence is solid, showing analytically that contrast-dependent gamma dynamics and a low-dimensional inter-areal communication subspace arise from one parameter set, though the comparisons to data remain qualitative and the analytics rest on a linearization that is not checked against numerical simulation. The work will interest neuroscientists and theorists concerned with inter-areal communication, cortical oscillations, and divisive normalization.

    2. Reviewer #1 (Public review):

      In this paper, Pal and colleagues propose a mechanistic unification of two influential accounts of inter-areal communication: communication through coherence and communication subspaces. A major strength of the paper is that it does not treat coherence and communication subspaces as independent phenomena, as typically done, but instead derives both from the same circuit with divisive normalization. In this framework, noise-driven fluctuations around the normalized fixed point determine covariance and cross-power structure (which, in retrospect, makes so much sense to be related). Then, they show how these determine linear prediction performance and the effective dimensionality of the communication subspace. They also show (however not very visually, see recommendation below for a figure) how divisive normalization is crucial to shape inter-areal coherence and the dimensionality of communication.

      I found this conceptual contribution potentially very influential, but somewhat obscured by the technical complexity of the model. The central intuition (I think) is that recurrent normalization can organize cross-area fluctuations, both frequency-specific correlations and cross-covariances. Took me a while to grasp this insight, mostly because I was stuck with the model details. Note that I have some experience with network dynamics, but not with this particular model.

    3. Reviewer #2 (Public review):

      Summary:

      The authors extend the ORGaNICs framework (a recurrent circuit that dynamically implements divisive normalization) to connected cortical areas with explicit top-down feedback. Because the network has a known analytical fixed point that coincides with (or closely approximates) the normalization equation, the authors can linearize about that fixed point and derive closed-form expressions for the power spectral density, inter-areal coherence, and communication subspaces. Using a two-area instantiation (V1 & V2) with a single fixed parameter set and no data fitting, they show the model reproduces: (i) contrast-response functions with steeper slope V2; (ii) gamma-band power and coherence peaks that shift to higher frequency with contrast; and (iii) a low-dimensional inter-areal communication subspace that is lower-dimensional than the within-area subspace. They derive parallel predictions of what happens by changing model parameters: feedback gain enhances inter-areal and suppresses within-area communication, and normalization is necessary for both the oscillatory dynamics and the reduced subspace dimensionality. A three-area extension (V1&V4, V1&V5/MT) is used to argue that differential top-down feedback can dynamically route functional connectivity.

      Strengths:

      (1) Analytical tractability: Deriving power spectra, coherence, and communication-subspace structure in closed form from a known fixed point is genuinely valuable.

      (2) Conceptual unification: Framing coherence and communication subspaces as arising from the same normalization-driven dynamics is an elegant and useful contribution.

      (3) Breadth from few assumptions: A large range of phenomena (contrast gain, gamma dynamics) emerges from normalization-based model assumptions.

      (4) Biological grounding: The mapping of model variables onto identified cell types connects the abstract computation to known cortical microcircuitry.

      (5) The prediction that input-gain versus feedback-gain modulation produce distinct spectral signatures gives experimentalists a clear way to test the framework.

      Weaknesses:

      (1) Comparisons are qualitative, not quantitative: The theory/experiment panels are visual side-by-side comparisons. There is no quantitative goodness-of-fit for any predictions.

      (2) The simulations use τ ≈ 1 ms for all cell types, which the authors acknowledge is unrealistically short; realistic values would shift the gamma peaks to lower frequencies.

      (3) Divisive normalization is a special case and is recovered exactly only for the identity recurrent matrix (self-normalization). Some statements that the circuit implements divisive normalization exactly need softening.

      (4) The element-wise (multiplicative) interaction in the modulator dynamics is not tied to a specific cellular mechanism.

    4. Reviewer #3 (Public review):

      Summary

      The work of Pal and colleagues considers a hierarchical and multi-population version of the "oscillatory recurrent gated neural integrator circuits" (ORGaNICs) model, showing through analytics that the model captures multiple relevant experimental results: first of all, its oscillatory dynamics produce a profile with high resemblance to experimental results, both in terms of decay of power at high frequency and in terms of shifting peak as a function of stimulus contrast. Second, inter-areal communication subspace dimensionality is lower than within-area dimensionality. The authors then proceed to further characterize the model's response properties as a function of input and feedback gain. In particular, they find that frequencies transmitted with higher strength also carry more information, that changing gain modifies the dimensionality of communication subspaces, and that these properties can be used in a three-layer model, where an upstream area can select which downstream area to communicate to, based on the strength of feedback gain.

      Strengths

      This work demonstrates that a single-circuit model with normalization properties can capture both the oscillatory dynamics and the inter-areal communication properties measured in cortical circuits, matching multiple experimental results. The full analytical tractability of the model is highly advantageous, allowing for easier exploration of parameters, replicability, and effective interpretations of results compared to purely numerical approaches.

      The work also makes a useful conceptual link between normalization, coherence-based communication, and subspace-based communication. In particular, it shows how both phenomena can emerge from the same circuit dynamics, where normalization is a key factor.

      Interestingly, the model is also extended to multiple areas, showing how attention (in the form of changes in feedback gain) can synchronize the activity of a downstream area with one of two upstream areas, thus effectively selecting which area to communicate with.

      In general, this is an interesting computational framework and a useful starting point for future modeling work. A particular strength is that it connects normalization, oscillatory dynamics, coherence, and communication subspaces within one analytically tractable model, making it possible to generate mechanistic hypotheses about when inter-areal communication should be stronger, lower-dimensional, or preferentially routed through feedback.

      Weaknesses

      Although I see the analytic approach as a strength, at the same time I regard the lack of any numerical comparison as a big weakness. Circuit simulations would not only confirm the correctness of the analytics, but also offer further insights on the error margins and on the regimes where the analytics are valid. This is because, to my understanding, the analytics are based on a linear approximation around the operating regime, which means deviations might be expected, especially for high gain levels in the input, or in the feedforward and feedback pathways.

      Another problem is that the analytically tractable model seems to rely on effective connectivity weights that break Dale's law. Numerical simulations with explicitly modeled excitatory and inhibitory units might give insights into effects due, e.g., to the additional transmission delays mentioned in the Discussion.

      Another weakness is the use of the term "predictions" to indicate features of the model dynamics that are purely described in the context of the model parameters. Although the model's response properties may certainly lead to predictions, I think the term requires a better contextualization in terms of neurophysiology and experimental neuroscience. The Discussion draws very interesting and valuable bridges between neuron morphology, interneuron types, and model parameters. But it seems it's left to the reader to backtrack and figure out which biological mechanisms or experimental manipulations should correspond to changes in input or feedback gain, and how these should be distinguished from possible changes in feedforward gain.

      Relatedly, the manuscript places substantial emphasis on modulation of feedback gain, but does not comparably explore modulation of the feedforward gain, β2, which regulates the V1-to-V2 drive. This seems important because changes in feedforward gain could also influence communication subspace dimensionality and oscillatory dynamics. Therefore, predictions related to top-down feedback modulations should be taken with a grain of salt.

      Last but not least, the model dynamics are split among multiple elements and nonlinear interactions, reaching a level of complexity far higher than the other ORGaNICs formulations present in the literature. The authors derive these dynamics in the supplementary material, as a dynamical system that converges to a fixed-point solution that includes "exact divisive normalization". I wonder, however, if there could be simpler solutions that also produce normalization, either approximate or in a different form than the one proposed by the authors. Note also that the designation of "excitatory neurons" is misleading: despite the presence of two explicitly inhibitory populations, the "excitatory" units also interact with negative effective weights both recurrently and in the inter-areal interactions, thus breaking Dale's law.

    1. eLife Assessment

      This valuable descriptive study describes the expression of a developmentally relevant transcription factor in the adult Tribolium brain. The evidence supporting the claims is convincing and based on a very detailed and rigorous analysis of light microscopy data, which, however, lacks single-cell resolution. This neuroanatomical study is of interest to the field of insect neural development and neuroscience.

    2. Reviewer #1 (Public review):

      Summary:

      Pang et al. investigated the expression pattern of the transcription factor foxQ2II in an adult beetle brain. They find nine distinct clusters, with many neurons expressing Glut/ChaT and dopamine. Some of the dopamine neurons resemble cell types described in Drosophila. Several neurons seem to project to prominent higher brain regions such as the MB and CX, and might even connect to both.

      Strengths:

      The authors use state-of-the-art labeling techniques for the analysis of individual cell types, such as beetle brainbow, to investigate the until now unknown expression of the transcription factor in the adult beetle brain.

      Rigorous cell reconstruction and image analysis revealed a better understanding of the anatomy of the labeled cells.

      Weaknesses:

      The brainbow labeling seems to include all cells labeled by the enhancer trap line, as well as the ones not expressing foxQ2II. Thus, it is unclear how useful this data is to compare individual cells to other insects.

      The functional relevance of this transcription factor in the adult brain cell is still unknown. It is therefore unclear if the described neurons have any specific function and if they require this transcription factor for normal function.

      Overall, the neural reconstructions are missing single-neuron details; it is difficult to compare the shown cell types to specific cell types in Drosophila based on the presented data, and this finding remains speculative.

    3. Reviewer #2 (Public review):

      Summary:

      The authors provide the first thorough profiling of neurons in Tribolium characterized by the expression of the transcription factor foxQ2, which will be useful for developmental neurobiology. They use state-of-the-art methods convincingly to not only identify the neurons, but also to further characterize them anatomically and neurochemically.

      Strengths:

      Thorough and meticulous application of state-of-the-art anatomical methods in a non-standard laboratory organism.

      Weaknesses:

      No weaknesses were identified by this reviewer.

      Comments:

      I don't really have any major suggestions at all. Loved the work.

      There is only one tiny nitpicking aspect:

      P21: "Biogenic amines are involved in learning and memory and setting arousal threshholds (Davis, 2023), which are functions performed by the mushroom bodies and related to the function of the central complex in goal directed navigation, respectively."

      MBs mainly process olfactory memory. At least in Drosophila, most other kinds of memories are being supported elsewhere.

      https://pubmed.ncbi.nlm.nih.gov/10454381/

      such as, e.g., visual pattern learning in the CX

      https://pubmed.ncbi.nlm.nih.gov/16452971/

      or motor learning in motor neurons

      https://pubmed.ncbi.nlm.nih.gov/38779314/

      or ventral ganglion, antennal lobes, and median bundle for place learning:

      https://pubmed.ncbi.nlm.nih.gov/10706599/

      If the authors focus on MBs, this sentence ought to reflect the fact that the function of the MBs is much narrower than the current sentence appears to suggest.

    4. Author response:

      We are very happy that our work was positively received by the reviewers and editors and we are looking forward sharing our results via eLife. 

      We have added more details on the generation of the Tribolium brainbow-lines and we have submitted the respective plasmids to Addgene and give the respective IDs. Some additional minor changes were done to make the text more clear. 

      Public Reviews: 

      Reviewer #1 (Public review): 

      Summary:

      Pang et al. investigated the expression pattern of the transcription factor foxQ2II in an adult beetle brain. They find nine distinct clusters, with many neurons expressing Glut/ChaT and dopamine. Some of the dopamine neurons resemble cell types described in Drosophila. Several neurons seem to project to prominent higher brain regions such as the MB and CX, and might even connect to both. 

      Strengths: 

      The authors use state-of-the-art labeling techniques for the analysis of individual cell types, such as beetle brainbow, to investigate the until now unknown expression of the transcription factor in the adult beetle brain. 

      We would want to add that this work establishes and introduces the brainbow system for the first time in an arthropod outside Drosophila melanogaster and that we are the first (outside flies) to relate the expression of a neural transcription factor with neural projection and neurotransmitter content.

      Rigorous cell reconstruction and image analysis revealed a better understanding of the anatomy of the labeled cells. 

      Weaknesses: 

      The brainbow labeling seems to include all cells labeled by the enhancer trap line, as well as the ones not expressing foxQ2II. Thus, it is unclear how useful this data is to compare individual cells to other insects. 

      We kindly disagree with the first statement: not all cells of the enhancer trap are labelled but a subset. Therefore, we call it “sparse labelling” in our manuscript while we do not reach “single cell labelling”, which admittedly limits both precision and use.

      The functional relevance of this transcription factor in the adult brain cell is still unknown. It is therefore unclear if the described neurons have any specific function and if they require this transcription factor for normal function. 

      Previously, we published that this gene has an important function in neural development during embryogenesis. Actually, we have done extensive RNAi experiments to test for an e ect during postembryonic development. We found surprisingly small defects when looking at alterations in several imaging lines. However, we found some changes in behavior. Given the extensive data presented in the current paper, we decided to publish these functional data (another 12 figures/suppl. figures) separately. 

      We also note that the identity/function of neurons is determined by a mix of transcription factors. Disentangling the individual role of each of those transcription factors indeed is an exciting question and a major endeavor beyond the scope of this paper.

      Overall, the neural reconstructions are missing single-neuron details; it is difficult to compare the shown cell types to specific cell types in Drosophila based on the presented data, and this finding remains speculative. 

      Indeed, we do not reach single cell resolution, which is below the standards of fly neurobiology. However, compared with all other arthropods we reach a unique level of precision. Specifically, we are the only ones outside fly research that relate the expression of a developmental transcription factor to neural projection and neurotransmitter content. 

      We also think that combining our transgenic line with dopamine-expression was su icient to compare the labelled cells to fly neurons. From what we saw in that analysis, we feel that most homology assessments of single neurons across such large evolutionary distances will remain hypothetical to some degree.

      Reviewer #2 (Public review): 

      Summary:

      The authors provide the first thorough profiling of neurons in Tribolium characterized by the expression of the transcription factor foxQ2, which will be useful for developmental neurobiology. They use state-of-the-art methods convincingly to not only identify the neurons, but also to further characterize them anatomically and neurochemically. 

      Strengths: 

      Thorough and meticulous application of state-of-the-art anatomical methods in a nonstandard laboratory organism. 

      Thank you for this encouraging comment. 

      Weaknesses: 

      No weaknesses were identified by this reviewer. 

      Comments: 

      I don't really have any major suggestions at all. Loved the work. There is only one tiny nitpicking aspect: 

      P21: "Biogenic amines are involved in learning and memory and setting arousal threshholds (Davis, 2023), which are functions performed by the mushroom bodies and related to the function of the central complex in goal directed navigation, respectively." 

      MBs mainly process olfactory memory. At least in Drosophila, most other kinds of memories are being supported elsewhere. https://pubmed.ncbi.nlm.nih.gov/10454381/ 

      such as, e.g., visual pattern learning in the CX https://pubmed.ncbi.nlm.nih.gov/16452971/

      or motor learning in motor neurons  https://pubmed.ncbi.nlm.nih.gov/38779314/

      or ventral ganglion, antennal lobes, and median bundle for place learning: https://pubmed.ncbi.nlm.nih.gov/10706599/ 

      If the authors focus on MBs, this sentence ought to reflect the fact that the function of the MBs is much narrower than the current sentence appears to suggest. 

      Thanks for this clarification – we have rephrased:

      "Biogenic amines are involved in learning and memory and setting arousal threshholds (Davis, 2023). This relates to the mushroom bodies’ function in olfactory memory, and the function of the central complex in visual pattern learning and goal directed navigation, respectively."

    1. eLife Assessment

      This paper describes a valuable tool for the detection and analysis of dentate spikes, network events in the dentate gyrus that are common yet understudied. This tool could help standardize dentate spike detection and analysis across labs and is therefore likely to be of interest to hippocampal neurophysiologists. However, the strength of evidence for its broad usefulness was viewed as incomplete, due to several identified bugs in the program and insufficient explanations of parameter selection and methods.

    2. Reviewer #1 (Public review):

      Summary:

      Esfahany et al. describe a new platform (Toothy) to identify and analyze dentate spikes and sharp wave ripples from silicon probe electrophysiology data. The goal is to facilitate and standardize the extraction of DS1 and DS2 events, which have highly variable properties across recordings from different labs. The manuscript describes the basic workflow of the Toothy pipeline, including loading data, assigning channels along a linear probe, customizing parameters, selecting ideal channels for analysis, and classifying DS1 and DS2 events.

      Strengths:

      The manuscript is clear and easy to follow and does a good job of describing the platform. Overall, this will be a useful analysis pipeline that can help to standardize DS analysis across labs and datasets.

      Weaknesses:

      The current version has several bugs that prevent analysis, and the documentation of analysis parameters needs to be improved.

      (1) In limited testing, the pipeline had several bugs, and I was not able to complete the full analysis of a dataset. Loading data from .mat or .npy files gave errors (it seemed that the metadata was not loaded correctly from the pop-up window). I was able to load a .nwb file, which worked well. The probe configuration tool was a bit difficult to understand, and there was not much documentation to help, although it worked when simply entering the x-y coordinates of the channels. It also crashed several times while trying to make a probe configuration due to it trying to save when a small typo was briefly entered. The initial analysis worked well, and the auto-selected channels matched our recording notes and seemed appropriate. DSs and ripples were extracted. An error came when trying to classify DSs, and the program repeatedly crashed across a variety of parameters. Overall, parts of the pipeline worked well, but others had significant bugs that need to be addressed.

      (2) The authors should provide test data that can be run through the pipeline. Ideally, this could use a variety of data types, probes, and conditions so that it is clear how they differ.

      (3) There are a lot of parameters that can be adjusted, but very little information about how they are chosen and what goes into parameter selection for a dataset. Additional documentation with more information on adjustable parameters, channel selection, and best practices would help improve the utility of the tool. Ideally, this could also integrate citations (either in the manuscript or documentation) to support some of the choices made during parameter selection.

      (4) There is no validation presented against other analysis methods or datasets. While there is no ground truth of when DSs occur, this may limit the ability of this tool to become the standard for DS analysis. A section comparing the analysis used in the pipeline to other published analyses would be helpful.

      (5) In the manuscript, it would be helpful to further describe the rationale for initially detecting DSs and SPW-Rs on all channels, when they are network events that occur across channels.

      (6) A section on what hardware and software are necessary to run the pipeline should be added.

    3. Reviewer #2 (Public review):

      Summary:

      This work provides an open-source, Python-based, graphical user interface for curating the detection and classification of dentate spikes (DSs) from hippocampal local field potential (LFP) recordings. The tool may also be used to detect, but not classify, sharp wave-ripples (SPW-Rs). The tool utilizes previously published Python packages for loading LFP files and creating experiment-specific probe objects. Detection and classification parameters are clearly defined and logged in a parameter file before starting processing. Once LFP data has been mapped to the probe object, event detection occurs across all channels. DSs are detected as qualifying peaks in the filtered DS band LFP, while SPW-Rs are detected as qualifying peaks in the filtered ripple band amplitude envelope. An initial curation step allows visualization of the LFP, instantaneous current source density (CSD), and depth-by-frequency band power plots for determining the approximate channel locations of key anatomical regions (i.e., CA1, the hippocampal fissure, and the hilus of the dentate gyrus). The optimal channel for detection is further refined in the next step by comparing event waveforms and quality metrics across channels. Artifacts and noisy waveforms can also be manually excluded during this step. Finally, DSs detected from the optimal channel are classified by computing the CSD profile around events and then clustering the first two principal components of all CSDs. The authors claim that this customizable tool will standardize DS detection and classification.

      Strengths:

      Toothy's detection and classification algorithms are appropriate and well-validated in the literature. The ability to change many parameters, the CSD calculation method, and clustering algorithm is helpful for precise replication of methodology that has varied previously. Default parameters optimized for mouse recordings provide a standardized starting point for rodent researchers.

      The authors' commitments to transparency and user-friendliness are to be commended (e.g., clear instructions, defined and logged parameters, multiple visualization options, etc.) and are likely to be appreciated by new users. Researchers with little-to-no coding experience should find this tool especially powerful for jumpstarting their own DS analyses.

      While not the focus of the paper, the capability to detect SPW-Rs provides an additional use case for Toothy and streamlines simultaneous analysis of SPW-Rs and DSs.

      Weaknesses:

      I encountered unexpected errors while trying to load LFP data into Toothy for testing, indicating that the "data ingestion" stage of Toothy requires minor code revision.

      Toothy's utility for recordings that do not produce an LFP depth profile is unclear. According to the authors, Toothy allows probe designs with irregular spatial sampling (e.g., tetrodes) to be used. However, recording from a linear probe with electrodes spanning from approximately the hippocampal fissure to the hilus of the dentate gyrus is required for Toothy's full functionality. For example, Toothy uses a DS type classification algorithm that relies on sufficiently sampled CSD depth profiles that tetrode recordings cannot provide. As such, usage is currently restricted to detection only for certain recording setups.

      The documentation on Toothy's output could be improved. Specifically, the work does not state which files different data are saved to or list the properties saved per detected event. Furthermore, the work does not discuss the potential importance of DS properties that are saved besides those related to the timing of the DS and its type.

    4. Reviewer #3 (Public review):

      Summary:

      Esfahany et al present a novel, UI-based tool to detect dentate spikes from hippocampal local field potential recordings, called Toothy. Toothy is easily accessible, compatible with many popular recording formats, and guides users entirely via UI through the dentate spike curation and analysis process. The functional and interactive visualizations enable users to gain a detailed understanding of their data and rigorously analyze dentate spike phenomena. This tool will be broadly useful for anyone who studies hippocampal electrophysiology. Furthermore, by expanding access to dentate spike analysis, it may encourage more scientists to explore this understudied but critical phenomenon.

      Strengths:

      (1) Toothy provides several ways for users to interact directly with parameters, revealing the ramifications of these choices. Most parameters are adjustable and made obvious via a UI panel. Their effects are then visualized across channels and individual events. This will help users think critically when selecting parameters.

      (2) Toothy is fully UI-based and pip-installable, lowering the barrier to entry far below what most electrophysiology analysis tools offer.

      (3) The channel selection tool is broadly useful for identifying DG hilus and CA1 pyramidal locations. Since subregional and laminar localization of electrode sites is critical to correctly interpret hippocampal recordings, this tool could be more generally used to identify site locations across the hippocampus.

      Weaknesses:

      (1) The rationale behind parameter choices is not explained. In order to function "not as a black-box detector", as the authors state, all initial parameter choices should be explained with citations. If possible, these citations would also be available from Toothy directly, alongside citations describing alternative parameter choices. This will help users make informed choices. For instance, a user analyzing data from rats would need to adjust the default ripple frequency band upwards (150-250Hz), and would benefit from guidance to adjust this properly.

      (2) The Results describe the functions of Toothy from the perspective of the user, but there is no Methods section describing what Toothy does between UI displays. This would allow readers to compare the tool directly to analysis pipelines as described in the Methods sections from other papers. Particular attention should be paid to justifying the analysis decisions that cannot be changed by the user, such as detecting events off of a single representative channel instead of across a consensus of multiple channels.

      (3) It's unclear whether or how Toothy evaluates data quality to confirm that its analyses return interpretable results. At a minimum, the tool should confirm adequate sampling rate (e.g. <=1kHz) and inter-site spacing for CSD (e.g. <=50um).

      (4) The paper does not put Toothy into context among the other common open-source electrophysiology analysis toolboxes. Consider Rippl-AI (Navas-Olive & Rubio et al, 2024) or pynapple (Viejo et al, 2023), to give a few examples. The paper would be strengthened by addressing how Toothy extends beyond the capacities of these other tools and how Toothy can be integrated into a workflow that also uses these other tools.

    5. Author response:

      We thank the editors and reviewers for their thoughtful and constructive assessment of Toothy, and for recognizing it as a potentially valuable resource for standardizing dentate spike (DS) analysis across labs. We are especially glad that the reviewers found the manuscript clear and easy to follow, judged the detection and classification algorithms to be appropriate and well-validated, and appreciated the tool's graphical user interface (GUI) based, pip-installable design for lowering the barrier to entry for DS analysis.

      We also understand the concerns raised. Most importantly, we will resolve the data-ingestion and classification errors that reviewers encountered and release an updated version of Toothy that we have verified end-to-end across input formats and datasets. Alongside this, we will provide downloadable demo dataset(s) spanning multiple file formats, probe types, and recording conditions, so that users can confirm a correct installation and see how these cases differ.

      To make the pipeline more transparent, we will add a section describing what Toothy does between user steps, including the rationale for decisions users cannot change, such as detection from a single representative channel. We will also expand the documentation of parameter choices with supporting citations and alternatives, and surface this guidance within Toothy where feasible, consistent with our aim that the tool not function as a black box.

      We will clarify Toothy's scope and current limitations. Recordings with irregular spatial sampling (e.g., tetrodes) are supported for detection but not for CSD-based DS-type classification, which requires a laminar probe spanning approximately the hippocampal fissure to the hilus; we will state this explicitly and evaluate adding an optional waveform-based classification mode (Santiago et al., 2024) to extend type classification to such recordings. We will also add data-quality checks (including sampling rate and inter-electrode spacing) that warn users when a recording may not support reliable results.

      Finally, we will situate Toothy among existing open-source toolboxes, describing how it differs, extends beyond, and interoperates with them, and we will add a comparison of Toothy's outputs to previously published analyses while being explicit about the limits of such comparisons. We will of course also address the remaining technical clarifications and figure edits raised by the reviewers.

      We are confident that addressing these points will make Toothy clearer and more useful to the hippocampal community.

    1. eLife Assessment

      This study reports important and invaluable findings that advance understanding of how attention is distributed between what we look at directly and what lies outside the center of gaze during active visual search. The evidence supporting the main claims is solid, with a large and rich dataset spanning multiple brain areas, although some aspects of the interpretation would benefit from additional controls and clearer separation of attention from eye-movement planning. The work will be of particular interest to researchers studying attention, visual perception, and eye movements behavior.

    2. Reviewer #2 (Public review):

      Summary:

      In natural visual behavior, such as when one is looking for a face in the crowd, the eyes are moved from site to site, seeking possible matching targets. This involves attention both to the current view at center of vision (the foveal location) as well as to upcoming views via attention to targets in the periphery. While it has been established that attention generally enhances neuronal response (compared to simple visual activation) at the attended spatial location, this study provides solid evidence that attention during active visual search leads to neuronal response enhancement only when the eye moves towards targets that exhibit the desired feature and category. This study thus moves the field towards understanding the neural encoding of active vision.

      This study examines the neuronal basis of feature selective attention during active, freely behaving visual search. Traditional electrophysiological studies on visual attention in monkeys commonly used an eye fixation with covert attention paradigm, but have not sufficiently addressed the roles of both foveal and peripheral attention in play during natural looking behavior. Here, the authors present a novel paradigm in which, during eye movement mediated search neuronal receptive fields are recorded in multiple cortical areas (sensory V4, temporal and prefrontal areas). In this manner, as the eye foveates, items in the array fall into foveal or non-foveal recorded sites. Thus, the experimental paradigm is elegant, offering the opportunity to make multiple types of comparisons: target/distractor, towards/away from fovea, areal. Specifically, following a category cue (face, house, hand, flower), freely initiated saccades are made to locate a categorically matching 'target' in an array of distractors. Feature attention is assessed by comparing eye saccades made to targets vs to distractors. Spatial attention is assessed by comparing saccades made 'towards' vs 'away' from targets. Statistics are rigorous and nicely designed. Detailed association of simultaneously obtained eye movement sequences and neural parameters are well done. These are valuable data which will contribute to our understanding of attentional modulation in visual search.

      The significance of these findings is fundamental. Decades of attention research in vision have been based on the paradigm of visual fixation and covert peripheral attention. However, increasingly the field has moved towards understanding how the visual system works during active vision. Here, the authors use an active visual search paradigm and record from key mid-tier (V4) and higher order (IT, PFC) areas. They find enhancement of attention both in the foveal and peripheral locations, and, furthermore, marked by a high degree of feature and categorical specificity. That is, while attention generally enhances neuronal response (compared to simple visual activation) at the attended spatial location, this study provides solid evidence that attention during active visual search leads to neuronal response enhancement only when the eye moves towards targets that exhibit the desired feature and category. This provides valuable data for the concept of a foveal-peripheral spatiotemporal attentional window in natural vision. The controls (comparisons of neuronal response during looks to targets vs distractors and looks towards and away from the target) and statistical rigor make these findings compelling. There will likely be additional future impacts of this study. For example, the eye movement patterns collected in this study may also provide a valuable dataset for future study of understanding search strategies. Goal-directed vs non-goal-directed task comparisons could be designed to test possible circuit models. Although much remains unknown regarding how and where frontal and temporal signals are integrated during active search, these data contribute important guideposts for future models of active visual search.

    3. Reviewer #3 (Public review):

      In this manuscript, the authors investigate the role of attention in foveal processing during a naturalistic task. They record neural activity from extrastriate visual areas V4 and inferotemporal cortex, as well as from the lateral prefrontal cortex, in macaques performing a free-gaze visual search task. In this task, animals searched for a face or house target among multiple complex stimuli, with no constraints on eye movements. Unlike classic studies of visual attention, which often rely on controlled fixation, this work examines neural activity in both foveal and peripheral receptive fields during naturalistic eye movements.

      The main question addressed by the authors is how feature-based attention is distributed and coordinated across foveal and peripheral visual fields during active search, and how this attentional processing influences saccade behavior. The authors show that foveal units in visual areas exhibit feature-based attentional enhancement, with stronger responses when a fixated stimulus is a target compared to when the same stimulus serves as a distractor. Peripheral units in visual and prefrontal areas show both feature-based and spatial attentional modulation, consistent with prior work. Finally, the authors show that attentional modulation depends primarily on stimulus category rather than response magnitude, with neurons showing similar enhancement for all images within the target category regardless of how strongly individual images drive the cell.

      There are several notable strengths of this paper including:

      (1) Disentangling feature-based and spatial attention during naturalistic vision remains a central challenge. This paper tackles both simultaneously, parsing neural populations by object selectivity (face-selective, house-selective, non-selective) and RF position (foveal vs. peripheral).

      (2) The unconstrained search task (Fig. 1A) moves beyond the dominant fixed-gaze, cued-attention designs (Zhou & Desimone, 2011) to study attention as it operates during natural behavior, with sequential fixations and voluntary saccades.

      (3) The scale of the multi-area recordings is a major strength and is well aligned with current trends in primate and human neuroscience toward large-scale, multi-area recordings. Simultaneous recordings from visual and prefrontal areas, comprising over 4,900 foveal units and more than 1,500 peripheral units, enable meaningful cross-area latency comparisons and area-specific analyses of attentional modulation. This study builds on the authors' previous analyses of this dataset by expanding the scope to show that feature-based attention generalizes across neuronal classes and operates on categorical identity rather than response magnitude.

      (4) The combination of simultaneous multi-area recordings and a rich behavioral paradigm provide a dataset that is well suited for population decoding, cross-area interaction analyses, and trial-by-trial prediction of saccade choices, which could substantially deepen mechanistic understanding beyond the largely univariate comparisons presented here.

      While the data broadly support the paper's main conclusions, several issues limit the strength of the mechanistic interpretation and should be taken into consideration:

      (1) Receptive field size is not explicitly quantified and may confound foveal-peripheral comparisons. Units are classified as foveal or peripheral based on responsiveness to the cue versus the search array (Methods, p. 17), but the manuscript lacks essential information about receptive field sizes, eccentricities, and the number of search stimuli falling within each receptive field and related proper controls. This is critical because receptive fields in visual area V4 at foveal eccentricities are relatively small (Gattass et al., 1988; Desimone & Schein, 1987), whereas receptive fields in inferotemporal cortex can span several degrees to tens of degrees and often include the fovea (Op de Beeck & Vogels, 2000; DiCarlo & Maunsell, 2003; Zoccolan et al., 2007). Given the 2{degree sign} × 2{degree sign} stimulus size, multiple search items could potentially fall simultaneously within peripheral receptive fields. This introduces a potential confound, as attentional modulation is known to be strongest when multiple stimuli appear within a single receptive field (Reynolds et al., 1999). Although the authors acknowledge this issue for visual area V4 (p. 17), it is neither quantified nor controlled for. Without explicit receptive field mapping relative to the search array, comparisons between foveal and peripheral units, as well as between visual areas, are difficult to interpret cleanly.

      (2) Attentional modulation is difficult to dissociate from saccade planning and decision-related signals. The free-gaze paradigm enhances ecological validity but introduces a temporal confound: mean distractor fixation durations are approximately 156 ms (p. 9), while attentional effects emerge between 137 and 170 ms after fixation onset (Fig. 2). As a result, the reported attentional modulation coincides with preparation of the subsequent saccade. Neural activity measured in the primary analysis window (150-225 ms; p. 19) therefore likely reflects a mixture of visual, attentional, motor planning, target recognition, and behavioral relevance signals, all of which are known to modulate responses in visual areas at similar latencies (e.g., Chelazzi et al., 1998). Moreover, target fixations (~257 ms) and distractor fixations (~156 ms) occur on fundamentally different behavioral timescales, which may inflate apparent foveal attentional effects. While the authors suggest that these timing differences support the idea that foveal feature-based attention facilitates prolonged fixation on target stimuli, this interpretation is not fully supported by the current analyses. That said, the saccade-aligned analyses of peripheral units (Fig. S3) partially mitigate this concern by demonstrating that feature-based modulation persists through saccade execution.

      (3) The "attention-out" condition for spatial attention lacks directional control. In the spatial attention analyses (Fig. 4D-F), the "attention-out" condition appears to include all fixations followed by saccades directed away from the receptive field, regardless of saccade direction. This differs from classic spatial attention designs, which typically use controlled anti-saccades or saccades to fixed locations opposite the receptive field (e.g., Moore & Armstrong, 2003; Gregoriou et al., 2009). Saccades directed toward locations adjacent to, but outside, the receptive field may still partially engage spatial attention mechanisms near the receptive field via broad attentional fields or motor preparation gradients (Bisley & Goldberg, 2010). In addition, the "attention-out" condition likely contains a heterogeneous mixture of trials in which the stimulus in the receptive field is either a target or a distractor, since feature-based attention effects are derived from this same pool of trials. As a result, spatial and feature attention effects are not fully orthogonal, and variance related to feature attention may already be embedded in the spatial attention baseline.

      [Editors' note: the authors have provided responses to each of these points.]

    4. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This manuscript aims to differentiate between foveal and peripheral attentional mechanisms in visual and frontal brain regions in monkeys engaged in a free-gaze visual search task.

      Strengths:

      The manuscript is clearly written, the question is important, and the behavioral task is interesting.

      Weaknesses:

      I have two major concerns.

      (1) The authors interpret divergence in neural responses to target vs nontarget as attention. But it is not. The subject has to attend to both target and nontarget stimuli to determine the stimulus category and thereby decide on the next action. Thus, divergence between target and nontarget responses could reflect categorical discrimination, but I am not sure this can be interpreted as attentional modulation. While it may be tempting to suggest that finding a stimulus of a specific category is "feature attention", analogous to, e.g., attending to the red stimulus, I don't believe this is correct. For the former, the animals have to attend to a stimulus, and examine the stimulus to determine the stimulus category, unlike a simpler discrimination, which may pop out. Given this, I am unconvinced that the interpretations in this manuscript are valid.

      We thank the reviewer for raising this concern. Selective attention is a process of focusing on goal-relevant stimuli (targets) while ignoring irrelevant distractions. Importantly, attentional selection is not limited to simple visual features (e.g., color, shape, or motion); it can also operate over more complex features. For example, objects themselves can serve as units of attentional selection [1, 2], and feature-based attentional effects have been observed when searching for images that match the cued images or image patches [3, 4]. In this context, attention can be directed either to overall features of an object or to objects as configurations of multiple non-spatial features. Furthermore, attention to the category of stimuli has been extensively investigated in fMRI experiments in humans [5-8], and it has been shown that attention can warp the representations of semantically related categories when participants search for different categories [7].

      Similarly, in our study, monkeys were trained to search for images that matched the category of the cue. The neural responses to targets versus distractors were compared while constrained to the same stimuli across different trials, ensuring that the observed response divergence was not due to the physical category of the targets and distractors. We also included only neural responses occurring prior to fixations associated with target selection, that is, before the monkeys made a behavioral choice, thereby controlling for potential contributions of target detection or decision-related signals to the observed effects.

      We have clarified and addressed this point in the Discussion as follows:

      “Feature-based attention to simple visual features such as color, shape, or motion has been extensively studied [1, 3, 5, 7-9, 11, 12, 64]. Attention can also operate over more complex features. For example, objects themselves can serve as units of attentional selection [65, 66], and feature-based attentional effects have been observed when searching for images that match the cued images or image patches [6, 67]. In this context, attention can be directed either to overall features of an object or to objects as configurations of multiple non-spatial features. Furthermore, attention to the category of stimuli has been extensively investigated in fMRI experiments in humans [68-71], and it has been shown that attention can warp the representations of semantically related categories when participants search for different categories [70]. In this study, the neural responses to targets versus distractors were compared while constrained to the same stimuli across different trials, ensuring that the observed response divergence was not due to the physical category of the targets and distractors.”

      (2) Regarding the RF classification of foveal and peripheral RFs for IT and PFC, prior work suggests that neurons in IT cortex (especially AIT) and PFC have RFs that largely include the foveal visual field. So, it would be important to include figures that show the RFs of neurons classified as foveal versus peripheral for all three areas.

      We thank the reviewer for raising this important point. We agree with the reviewer that neurons in IT cortex and PFC often have RFs that include the foveal visual field. We did record foveal units with both focal and broad foveal RFs; however, in our analysis we only included neurons with focal foveal RFs to exclude the influence of peripheral stimuli. We defined focal foveal-RF units as those that responded solely to the cue in the foveal region and not to items in the search array presented at least 5° away from the central fixation point, ensuring that their RFs did not extend to these peripheral locations. The items were also separated by at least 5° from each other, excluding the possibility that peripheral stimuli fell within their RFs during fixations. By definition, their RFs were restricted to the central point. This is further supported by Fig. S1A-H, which shows no responses to items in the search array at peripheral locations. We have made modifications in the Results and Methods as follows:

      “Notably, the items in the search array were presented at least 5° from the central fixation point and were also separated by at least 5° from each other, excluding the possibility that peripheral stimuli fell within their foveal RFs during fixations.”

      And:

      “In this study, our focus was on units with focal foveal RFs and units with localized peripheral RFs. All further analyses were conducted on these units.”

      We modified Fig. 1 to illustrate the RFs of neurons classified as peripheral, which were also characterized in our previous study using the same dataset [9]. The peripheral population exhibits no responses to the central cue (Fig. S1I–T).

      Reviewer #2 (Public review):

      Summary:

      In natural visual behavior, such as when one is looking for a face in the crowd, the eyes are moved from site to site, seeking possible matching targets. This involves attention both to the current view at the center of vision (the foveal location) as well as to upcoming views via attention to targets in the periphery. While it has been established that attention generally enhances neuronal response (compared to simple visual activation) at the attended spatial location, this study provides solid evidence that attention during active visual search leads to neuronal response enhancement only when the eye moves towards targets that exhibit the desired feature and category. This study thus moves the field towards understanding the neural encoding of active vision.

      This study examines the neuronal basis of feature-selective attention during active, freely behaving visual search. Traditional electrophysiological studies on visual attention in monkeys commonly used an eye fixation with a covert attention paradigm, but have not sufficiently addressed the roles of both foveal and peripheral attention in play during natural looking behavior. Here, the authors present a novel paradigm in which, during eye-movement mediated search, neuronal receptive fields are recorded in multiple cortical areas (sensory V4, temporal, and prefrontal areas). In this manner, as the eye foveates, items in the array fall into foveal or non-foveal recorded sites. Thus, the experimental paradigm is elegant, offering the opportunity to make multiple types of comparisons: target/distractor, towards/away from fovea, and areal. Specifically, following a category cue (face, house, hand, flower), freely initiated saccades are made to locate a categorically matching 'target' in an array of distractors. Feature attention is assessed by comparing eye saccades made to targets vs to distractors. Spatial attention is assessed by comparing saccades made 'towards' vs 'away' from targets. Statistics are rigorous and nicely designed. The detailed association of simultaneously obtained eye movement sequences and neural parameters is well done. These are valuable data that will contribute to our understanding of attentional modulation in visual search.

      Strengths:

      The significance of these findings is fundamental. Decades of attention research in vision have been based on the paradigm of visual fixation and covert peripheral attention. However, increasingly, the field has moved towards understanding how the visual system works during active vision. Here, the authors use an active visual search paradigm and record from multiple areas (V4, IT, PFC). They find enhancement of attention both in the foveal and peripheral locations, and, furthermore, a high degree of feature and categorical specificity. This provides valuable data for the concept of a foveal-peripheral attentional window in natural vision. The controls (comparisons of neuronal response during looks to targets vs distractors, and looks towards and away from the target) and statistical rigor make these findings quite compelling.

      Weaknesses:

      While the study is generally quite strong, there are a few weaknesses to be addressed.

      (1) Little rationale is provided for recording in the selected areas, V4, IT, and PFC. Given the respective roles in sensory, object recognition, and goal-directed behavior, some rationale for this design should be offered, and commonalities/distinctions between these areas should be discussed.

      We thank the reviewer for the suggestion and we modified and added the rationale to the Introduction as follows:

      “V4 and inferotemporal cortex (IT), as the middle and high-level areas of the ventral visual stream, are important for object recognition and categorization [27-34], and their roles have been extensively studied in central vision. At the neuronal level, however, most investigations have largely neglected their functions during active, free-gaze visual search. The prefrontal cortex, including LPFC, has long been implicated as a source of top-down signals that bias the selection of attended features and modulate visual cortical responses [6, 9, 11, 35-40]. Although target-related visual responses have been reported in IT during visual exploration [41], and target-selective responses have been observed in the human medial temporal lobe (MTL) [42] and medial frontal cortex (MFC) [43] during visual search, these studies did not map the receptive fields (RFs) of recorded neurons.”

      We also added a discussion as follows:

      “Some studies have provided evidence for integration between peripheral and foveal feature information across saccades, including features such as stimulus color [58, 59] and object orientation [60, 61], and visual features have been shown to be predictively remapped prior to saccades [62]. Our finding provides a potential neuronal mechanism that may support this integration process [63]. We found that LPFC’s extensive representation of the visual periphery provides a neural substrate for monitoring the broader search array. Crucially, our finding that LPFC activity temporally precedes attentional effects in the visual area—consistent with previous studies [6, 9, 11, 35-40] suggests that it does not merely reflect peripheral sensory input. Instead, LPFC likely acts as a top-down orchestrator, projecting task-relevant templates derived from current foveal goals onto peripheral candidate locations, a possibility that warrants further investigation.”

      (2) Given the reliance of all analyses on saccadic behavior (towards target/distractor, towards/away from target), additional description and summaries of eye movement behavior during single trials and across trials should be provided.

      We thank the reviewer for this helpful suggestion. We have added a description of saccade behavior to the Results as follows:

      “The mean number of saccades monkeys made to find the target after the onset of the search array was 2.25 ± 1.35 (mean ± SD across trials; Table 1) of correct trials, and the mean saccade amplitude was 7.99° ± 3.58° (mean ± SD across saccades; Table 1). Monkeys could fixate on each distractor or the target freely, provided they did not maintain fixation on the target for longer than 800 ms. Across sessions, 42.44% ± 3.6% of saccades were directed to distractors, 57.56% ± 3.6% to targets, and 12.59% ± 3.46% were saccades away from targets (see our previous studies [44-46] for detailed behavioral analyses).”

      We have modified Fig. 1A and its legend to illustrate the saccadic patterns of monkeys during the search task.

      We have also included Table 1, which summarizes eye movement behaviors.

      (3) The dependency of findings on top-down (categorical & feature-specific) task design should be discussed.

      We thank the reviewer for the suggestion and added a discussion as follows:

      “In this task, attention is strongly guided by top-down goals, which bias processing toward behaviorally relevant features and object categories [2, 50, 51]. Top-down attention, including categorical and feature-specific components, has been shown to modulate neural processing across the visual pathway based on task demands and to originate from distributed frontoparietal control networks [11, 35-38, 40]. Our study provides further insight into the mechanisms of goal-directed visual attention, as it is among the first to demonstrate foveal feature attention effects during free-gaze visual search, as well as the distribution of feature and spatial attention across the entire visual field.”

      Reviewer #3 (Public review):

      In this manuscript, the authors investigate the role of attention in foveal processing during a naturalistic task. They record neural activity from extrastriate visual areas V4 and inferotemporal cortex, as well as from the lateral prefrontal cortex, in macaques performing a free-gaze visual search task. In this task, animals searched for a face or house target among multiple complex stimuli, with no constraints on eye movements. Unlike classic studies of visual attention, which often rely on controlled fixation, this work examines neural activity in both foveal and peripheral receptive fields during naturalistic eye movements.

      The main question addressed by the authors is how feature-based attention is distributed and coordinated across foveal and peripheral visual fields during active search, and how this attentional processing influences saccade behavior. The authors show that foveal units in visual areas exhibit feature-based attentional enhancement, with stronger responses when a fixated stimulus is a target compared to when the same stimulus serves as a distractor. Peripheral units in visual and prefrontal areas show both feature-based and spatial attentional modulation, consistent with prior work. Finally, the authors show that attentional modulation depends primarily on stimulus category rather than response magnitude, with neurons showing similar enhancement for all images within the target category regardless of how strongly individual images drive the cell.

      There are several notable strengths of this paper, including:

      (1) Disentangling feature-based and spatial attention during naturalistic vision remains a central challenge. This paper tackles both simultaneously, parsing neural populations by object selectivity (face-selective, house-selective, non-selective) and RF position (foveal vs. peripheral).

      (2) The unconstrained search task (Figure 1A) moves beyond the dominant fixed-gaze, cued-attention designs (Zhou & Desimone, 2011) to study attention as it operates during natural behavior, with sequential fixations and voluntary saccades.

      (3) The scale of the multi-area recordings is a major strength and is well aligned with current trends in primate and human neuroscience toward large-scale, multi-area recordings. Simultaneous recordings from visual and prefrontal areas, comprising over 4,900 foveal units and more than 1,500 peripheral units, enable meaningful cross-area latency comparisons and area-specific analyses of attentional modulation. This study builds on the authors' previous analyses of this dataset by expanding the scope to show that feature-based attention generalizes across neuronal classes and operates on categorical identity rather than response magnitude.

      (4) The combination of simultaneous multi-area recordings and a rich behavioral paradigm provides a dataset that is well-suited for population decoding, cross-area interaction analyses, and trial-by-trial prediction of saccade choices, which could substantially deepen mechanistic understanding beyond the largely univariate comparisons presented here.

      While the data broadly support the paper's main conclusions, several issues limit the strength of the mechanistic interpretation and should be taken into consideration:

      (1) Receptive field size is not explicitly quantified and may confound foveal-peripheral comparisons. Units are classified as foveal or peripheral based on responsiveness to the cue versus the search array (Methods, p. 17), but the manuscript lacks essential information about receptive field sizes, eccentricities, and the number of search stimuli falling within each receptive field and related proper controls. This is critical because receptive fields in visual area V4 at foveal eccentricities are relatively small (Gattass et al., 1988; Desimone & Schein, 1987), whereas receptive fields in inferotemporal cortex can span several degrees to tens of degrees and often include the fovea (Op de Beeck & Vogels, 2000; DiCarlo & Maunsell, 2003; Zoccolan et al., 2007). Given the 2{degree sign} × 2{degree sign} stimulus size, multiple search items could potentially fall simultaneously within peripheral receptive fields. This introduces a potential confound, as attentional modulation is known to be strongest when multiple stimuli appear within a single receptive field (Reynolds et al., 1999). Although the authors acknowledge this issue for visual area V4 (p. 17), it is neither quantified nor controlled for. Without explicit receptive field mapping relative to the search array, comparisons between foveal and peripheral units, as well as between visual areas, are difficult to interpret cleanly.

      We thank the reviewer for the helpful suggestion and apologize for not explicitly providing essential information about the RFs of the units. We added a detailed description of RF properties to the Results as follows:

      “The RFs of these peripheral units were further mapped using a visually guided saccade task and quantified by the number of stimuli that activated each unit (Fig. 1F-K). The eccentricities of the peripheral RFs were 6.22° ± 1.31° (mean ± SD) in V4, 7.04° ± 1.52° in IT, and 6.68° ± 1.56° in LPFC. The sizes of the peripheral RFs were 3.67° ± 1.87° in V4, 6.86° ± 3.11° in IT, and 8.65° ± 3.02° in LPFC. The numbers of items from the search array falling within peripheral RFs were 1.49 ± 0.55 in V4, 2.2 ± 0.72 in IT, and 2.56 ± 0.74 in LPFC (also see our previous study [44]).”

      The reviewer is correct that multiple items from the search array did fall within the RFs of peripheral-RF units. However, for focal foveal units, only the fixated stimulus fell within the RF, due to the design of the search array and the definition of these units used in our analyses (see our reply to Reviewer 1, Public Review, Question 2 for details). We agree with the reviewer that attentional modulation is typically stronger when multiple stimuli fall within RFs. In our design, peripheral RFs, on average, contained more stimuli than foveal RFs. Therefore, this difference in RF size would, if anything, be expected to bias toward stronger attentional modulation in peripheral units. This would make our observation conservative, thereby further supporting rather than undermines our main finding of robust feature-based attentional enhancement in foveal units, challenging the prevailing view that such modulation is predominantly peripheral. However, we agree that, when comparing the latency of attentional effects across brain regions in Fig. 3, we cannot rule out the influence of the number of stimuli arising from differences in RF size.

      (2) Attentional modulation is difficult to dissociate from saccade planning and decision-related signals. The free-gaze paradigm enhances ecological validity but introduces a temporal confound: mean distractor fixation durations are approximately 156 ms (p. 9), while attentional effects emerge between 137 and 170 ms after fixation onset (Figure 2). As a result, the reported attentional modulation coincides with the preparation of the subsequent saccade. Neural activity measured in the primary analysis window (150-225 ms; p. 19), therefore, likely reflects a mixture of visual, attentional, motor planning, target recognition, and behavioral relevance signals, all of which are known to modulate responses in visual areas at similar latencies (e.g., Chelazzi et al., 1998). Moreover, target fixations (~257 ms) and distractor fixations (~156 ms) occur on fundamentally different behavioral timescales, which may inflate apparent foveal attentional effects. While the authors suggest that these timing differences support the idea that foveal feature-based attention facilitates prolonged fixation on target stimuli, this interpretation is not fully supported by the current analyses. That said, the saccade-aligned analyses of peripheral units (Figure S3) partially mitigate this concern by demonstrating that featurebased modulation persists through saccade execution.

      We thank the reviewer for raising this important question. We agree that the temporal overlap of visual, motor planning, target recognition, and behavioral relevance signals with attention can result in mixed activity, which needs to be dissociated. Therefore, when calculating feature-based attention, we did implement a series of controls. We added a discussion as follows:

      “A major challenge in interpreting neural activity related to attentional modulation is the inherent temporal overlap of visual processing, motor planning, and target recognition signals in the free-gaze visual search task [73]. To isolate genuine feature-based attention from potential confounds, we applied several stringent analytical constraints, consistent with prior studies [3, 5, 6]. Specifically, by restricting our analysis to fixations where the subsequent saccade was directed away from the RFs, we dissociated attentional modulation from the preparatory motor activity associated with saccade execution. Furthermore, by comparing responses to the same physical stimulus alternating its role as a target or distractor across trials we eliminated any potential bias introduced by stimulus identity or physical category. We restricted our analysis to fixations preceding target selection that is, before the monkeys made a behavioral choice to minimize contributions from target detection or decision-related signals.”

      We thank the reviewer for pointing out the issue of different timescales for target versus distractor fixations. To address this, we conducted a control analysis by computing foveal feature-based attentional modulation using fixations on targets and distractors with matched fixation durations. We obtained similar results. We have updated Fig. S2 to include this control analysis.

      We also clarified this point in the Results as follows:

      “We also obtained similar results when controlling for fixation durations on targets and distractors (i.e., there was no significant difference between fixation durations on targets and distractors; Wilcoxon signed-rank test, P > 0.05; Fig. S2K–P).”

      Lastly, as the reviewer correctly pointed out, the interpretation that foveal feature-based attention facilitates prolonged fixation on the target was not supported. We have revised the Results as follows:

      “On average, target fixations (256.69 ± 197.44 ms [mean ± SD]) were significantly longer than distractor fixations (156.26 ± 45.94 ms; Wilcoxon rank-sum test, P < 0.0001), and during these prolonged target fixation, foveal feature-based attention modulation was consistently observed.”

      (3) The "attention-out" condition for spatial attention lacks directional control. In the spatial attention analyses (Figures 4D-F), the "attention-out" condition appears to include all fixations followed by saccades directed away from the receptive field, regardless of saccade direction. This differs from classic spatial attention designs, which typically use controlled anti-saccades or saccades to fixed locations opposite the receptive field (e.g., Moore & Armstrong, 2003; Gregoriou et al., 2009). Saccades directed toward locations adjacent to, but outside, the receptive field may still partially engage spatial attention mechanisms near the receptive field via broad attentional fields or motor preparation gradients (Bisley & Goldberg, 2010). In addition, the "attention-out" condition likely contains a heterogeneous mixture of trials in which the stimulus in the receptive field is either a target or a distractor, since feature-based attention effects are derived from this same pool of trials. As a result, spatial and feature attention effects are not fully orthogonal, and variance related to feature attention may already be embedded in the spatial attention baseline.

      We thank the reviewer for this important question. We performed a directional control analysis by computing spatial attentional modulation using paired fixations from the attention-in and attention-out conditions. Only saccades directed in nearly opposite directions—defined as having a saccade direction angle ≥ 170° within the 0–180° range—were included. We obtained similar results (Author response image 1). 

      Author response image 1.

      Peripheral spatial attentional modulation in V4, IT, and LPFC. Population response to stimuli followed by saccades directed into their RFs (attention in) versus directed approximately opposite and outside their RFs (attention out), shown for V4 (A), IT (B), and LPFC (C). Shaded area denotes ±SEM across units.

      We did control for feature-based attention when calculating spatial attentional modulation. We apologize for the lack of clarity and have added a description of this control to the Methods as follows:

      “The saccade-target stimulus in the RF during attention-in fixations was matched to a stimulus in the same location during attention-out fixations; in both conditions, this stimulus always served as a distractor for that trial, except in the “Distractor fixations to T” condition (Fig. 5 and Fig. S4), in which it instead served as the target. This design eliminates differences due to feature-based attention between the attention-in and attention-out conditions.”

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) Figure 3C: Unclear how to compare LPFC vs V4 for foveal units since only data from peripheral LPFC is shown?

      We thank the reviewer for pointing out this mistake. In Fig. 3C, we only compared LPFC peripheral units, V4 peripheral units, and V4 foveal units. We have corrected this in the legend of Fig. 3 as follows:

      “Shown are cumulative distributions of feature-attention effect latencies, computed from individual foveal face-, house-, and non-selective units in V4 and IT, and from peripheral non-selective units in V4,

      IT, and LPFC.”

      (2) On page 8, last para: For units with peripheral RFs ... Is this controlled for whether the saccade is to targets or to distractors?

      We thank the reviewer for the question. We indeed addressed this concern by separating fixations based on whether the subsequent saccade was directed to a target or a distractor, and by analyzing attention modulation within each condition. Therefore, attention effects were evaluated while holding the saccade destination constant, effectively controlling for potential confounds related to saccade target selection.

      (3) Page 9: The authors find that target fixations were longer than distractor fixations and conclude that this supports the idea that foveal feature-based attention increases fixation duration, but this interpretation is pure conjecture, and there is no experimental manipulation presented in this paper that helps to establish this interpretation.

      We thank the reviewer for this important comment. We agree that this observation does not, by itself, support our original interpretation, and we have modified it in the Results. Please refer to the last paragraph of our Reply to Question 2 from Reviewer 3 (Public Review).

      (4) Data analysis: receptive field. The authors state that visual response to a cue and the stimulus array was assessed during the 0-200 ms window after stimulus onset. However, after the array onset, the animal could saccade within the 200 ms window. How do the authors ensure uniform stimulation during the 0-200 ms window?

      We thank the reviewer for this question. The activity of units in V4, IT, and LPFC within the 200 ms window after array onset primarily reflected visual stimulation prior to saccades, because typical saccade latencies were approximately 150–200 ms, and the response onset latencies of these units were around 50 ms.

      (5) On page 19, the authors state that to assess feature attention in peripheral RFs, they divided trials into target and distractor fixations. In the former, there was a target in the neuron's RF. This is confusing. I assume target fixations imply fixating on a target, but the authors may mean fixations where a target is in the RF. Please clarify.

      We thank the reviewer for pointing out this confusion. In the original manuscript, we intended to sort fixations by whether a target stimulus was located within the unit’s peripheral RF. To avoid further confusion, we have revised the description in the Methods as follows:

      “we sorted fixations during the search period, following a procedure similar to that in our previous study [5], into two types: “target” – a target stimulus was located within the unit’s peripheral RF; and “distractor” – the same stimulus appeared in the same peripheral RF location but served as a distractor.”

      (6) Figure S1: Are these example units? How many trials? SEM? The sharp rise and no noise are inconsistent; the former suggests minimal smoothing, while the latter suggests lots of smoothing.

      We thank the reviewer for these questions. We showed average responses across all units in Fig. S1. On average, there were 941.79 ± 182.56 trials (mean ± SD across sessions). Shaded areas indicate ±SEM across units. The sharp rise reflects the synchronous response of neurons to the stimulus, while the smooth appearance and low noise result from averaging across a very large number of units and trials.

      Reviewer #2 (Recommendations for the authors):

      Major comments:

      (1) One weakness of this manuscript is the lack of a rationale for choosing V4, IT, and PFC. Specifically, what are the predictions of the roles of these respective areas in the integration of current and peripheral (future foveal) views? There is a significant literature linking the pre-saccadic peripheral stimulus and the post-saccadic foveal stimulus, suggesting that both spatial and temporal integration occur. However, whether such integration occurs at high or low cortical levels is unknown. By recording from mid-tier (V4) and high-order areas (IT, PFC), the authors have an opportunity to address this question. However, there is no mention of this topic, either in the introduction, results, or discussion. I find this omission surprising. At the very least, it should contribute to experimental design rationale and some discussion.

      We thank the reviewer for the suggestion and we modified and added the rationale to the Introduction and a discussion about this integration. Please refer to our Reply to Question 1 from Reviewer 2 (Public Review).

      (2) As both behavior and neural recordings are collected, a figure on saccadic patterns would enhance the reader's understanding. Questions that come to mind are: What does a single search trial look like? How many saccades are there per trial? How often is the target identified after 1, 2, 3, etc saccades? What is the average size of a saccade? Although this is not a study of search strategy per se, a modicum of description of the search sequences would provide context on the behavior. I suggest an illustration of one or more sample trials; a summary of saccade behavior would also be helpful for understanding the data in relation to behavioral performance.

      We thank the reviewer for this helpful suggestion. We have modified Fig. 1A and its legend to illustrate the saccadic patterns of monkeys during the search, providing an example of a single search trial. Additionally, we have added a description of saccade behavior to the Results and included Table 1, which summarizes eye movement behavior. Please refer to our Reply to Question 2 from Reviewer 2 (Public Review) for further details.

      (3) "Consistently, the probability of making a saccade to a peripheral target was higher following distractor fixations (75.22%) than following target fixations (48.44%, or 63.49% after probability calibration; see Methods), indicating the important role of peripheral feature-based attention in guiding eye movements" It should be noted that this target-oriented visual search is fundamentally a top down task. Once the target is found, the reward is obtained; saccades to distractors are not rewarded, so saccades are more likely. So certainly this task design would increase the post-distractor saccades and decrease the number of post-target saccades. Please clarify the behavioral paradigm: once a reward is obtained, does the task continue, or is a new trial initiated?

      We apologize for the confusion regarding the behavioral paradigm. We would like to clarify that when the target was found and fixated for 800 ms, the reward was delivered and no further saccades occurred. However, if the target was not fixated for 800 ms, the search could continue. It is worth noting that the target fixations in our analyses were restricted to those occurring during ongoing search behavior, excluding target fixations associated with trial termination and reward delivery. Moreover, we compared the probability of making a saccade to the target, rather than the absolute number of saccades, following these fixations. We have modified the Results for clarification, as follows:

      “Two monkeys performed a category-based visual search task, where their objective was to fixate on one of the two search targets that matched the category of the cue (Fig. 1A, B). Specifically, the monkeys were presented with a central fixation point for 400 ms, followed by a cue lasting 500-1300 ms. After a 500 ms delay, a search array appeared with 11 items, including two targets, randomly chosen from 20 possible locations (Fig. 1E). The monkeys had 4000 ms to find one target and maintain fixation on it for 800 ms to earn a juice reward. Fixating on either target completed the trial, and the monkeys did not search for the second target. A new trial began after the reward. It is worth noting that the two target stimuli matched the category of the cue but were different images. The monkeys were required to maintain fixation throughout the cue and delay periods. During search, however, eye movements were unconstrained, and monkeys could revisit each search distractor or target as long as they did not fixate on a target for 800 ms.”

      (4) The fact that there are many more peripheral units in LPFC suggests that this is a region of foveal/periph integration. Combined with the finding that the LPFC leads the attentional effects, this should be a discussion point.

      We thank the reviewer for the suggestion and we added a discussion as follows:

      “Some studies have provided evidence for integration between peripheral and foveal feature information across saccades, including features such as stimulus color [58, 59] and object orientation [60, 61], and visual features have been shown to be predictively remapped prior to saccades [62]. Our finding provides a potential neuronal mechanism that may support this integration process [63]. We found that LPFC’s extensive representation of the visual periphery provides a neural substrate for monitoring the broader search array. Crucially, our finding that LPFC activity temporally precedes attentional effects in the visual area consistent with previous studies [6, 9, 11, 35-40] suggests that it does not merely reflect peripheral sensory input. Instead, LPFC likely acts as a top-down orchestrator, projecting task-relevant templates derived from current foveal goals onto peripheral candidate locations, a possibility that warrants further investigation.”

      Minor comments:

      (1) Figures 2A-D. "These face-selective units also showed slightly enhanced responses to house targets in IT (P < 0.05), but not in V4 (P = 0.89)." It does not appear enhanced.

      We agree with the reviewer that the effect is modest and does not appear strongly enhanced. However, the average response in the 150–225 ms time window to the house target was significantly higher than that to the house distractor in IT face-selective units (Wilcoxon signed-rank test, P = 0.042). We modified the description in the Results as follows: 

      “These face-selective units also showed weakly but significantly enhanced responses to house targets in IT (P < 0.05)”

      (2) Figure 3. For population comparison, a bootstrapped null distribution was used, and a 2-sided permutation test was used to determine the latency difference between the target and distractor; please show these results (described in text) in a figure. Figures 3A-C are described as the latency of individual units. So each of these graphs is the mean of multiple units? So this is also a population analysis? What is the difference between these two comparisons? This is somewhat confusing.

      We apologize for the confusion and thank the reviewer for pointing this out. Each panel in Fig. 3 shows the cumulative distribution of latencies across individual units within each brain region, reflecting the variability of response timing across single neurons. For this analysis, we first calculate the latency of each unit separately. In contrast, population-level latency is measured from the averaged responses of all units within each region (Fig. 2), which captures the overall timing of the population response rather than individual variability. Statistical comparisons at the population level are performed using a two-sided permutation test. We modified Fig. 2 to better illustrate the population-level latency results.

      (3) Did peripheral RFs span more than a single stimulus in the array? If so, how does this impact the interpretation of Figure 5?

      We thank the reviewer for pointing this out. The reviewer is correct that, in peripheral RFs, more than one stimulus from the search array could fall within the receptive field (1.49 ± 0.55 in V4, 2.2 ± 0.72 in IT, and 2.56 ± 0.74 in LPFC). We controlled for this in our analysis of both feature-based and spatial attention effects for peripheral units in Fig. 5. For feature-based attention, we performed the analysis in a stimulus-by-stimulus manner within each category (house and face), such that when a given stimulus served as the target, it was the only target within the RF, and when it served as a distractor, it was the only distractor of its category within the RF. Although additional distractor could still fall within the RF, their identities were random across conditions and thus would be averaged out. A similar approach was applied to spatial attention, where the stimulus-by-stimulus comparison was extended across all four categories, and attention-out stimuli were paired with the corresponding saccade-target stimuli in the attention-in condition, with the effects of other randomly present distractors averaged out. Therefore, the effects shown in Fig. 5 reflect comparisons at the level of individual stimulus, minimizing confounds from other stimuli within the RF.

      (4) Figure 5G: "during "Target fixations to D", there was no significant feature attentional enhancement in response to the peripheral target (Wilcoxon signed-rank test, P > 0.05; Figure 5G-I left panels). It appears that there is some effect of spatial attention during Target Fix to D trials.

      We thank the reviewer for pointing this out and have revised the Results as follows:

      “We further found that spatial attentional enhancements to the saccade target were reduced during target fixations compared to distractor fixations in V4 and IT when activity was aligned to fixation onset (Wilcoxon rank-sum test, P < 0.05; Fig. 5G, H versus Fig. 5A, B), although this effect was not completely abolished.”

      (5) The specific areas of IT and LPFC that were recorded should, as much as possible, be mentioned.

      We thank the reviewer for the helpful suggestions and have added a description of the specific IT and LPFC recording sites to the Methods as follows:

      “Recordings in IT spanned the central IT cortex, encompassing the area between the anterior middle temporal sulcus (AMTS) and the posterior middle temporal sulcus (PMTS), including TE and TEO. Recordings in LPFC were located anterior to the arcuate sulcus (AS) and lateral to the principal sulcus (PS), mainly covering areas 45 and 44.”

      (6) It is often difficult to distinguish the different lines, e.g., red solid vs red dotted, due to their overlap. Would the removal of the error band make this clearer? If so, could put full figure with error bands in the Supplementary Figure.

      We thank the reviewer for this helpful suggestion. To improve visual clarity, we adjusted Fig. 6, Fig. 7, Fig. S2, Fig. S3, Fig. S4, and Fig. S6 by changing the line styles and placing the shaded error bands beneath the traces, allowing the lines to remain clearly visible despite overlap.

      (7) For easy access, the number of saccades to/from targets/distractors should be put into a table.

      We thank the reviewer for the suggestion. We calculated the probability of saccades to and from targets and distractors for each session and report the mean ± SD across sessions in Table 1, as the mean number of saccades per trial was only 2.3. Please refer to our Reply to Question 2 from Reviewer 2 (Public Review) for Table 1.

      Reference

      (1) O'Craven, K.M., P.E. Downing, and N. Kanwisher, fMRI evidence for objects as the units of attentional selection. Nature, 1999. 401(6753): p. 584-7.

      (2) Baldauf, D. and R. Desimone, Neural mechanisms of object-based attention. Science, 2014. 344(6182): p. 424-7.

      (3) Hayden, B.Y. and J.L. Gallant, Combined effects of spatial and feature-based attention on responses of V4 neurons. Vision Res, 2009. 49(10): p. 1182-7.

      (4) Bichot, N.P., et al., A Source for Feature-Based Attention in the Prefrontal Cortex. Neuron, 2015. 88(4): p. 832-844.

      (5) Reddy, L. and N. Kanwisher, Category selectivity in the ventral visual pathway confers robustness to clutter and diverted attention. Curr Biol, 2007. 17(23): p. 2067-72.

      (6) Peelen, M.V., L. Fei-Fei, and S. Kastner, Neural mechanisms of rapid natural scene categorization in human visual cortex. Nature, 2009. 460(7251): p. 94-7.

      (7) Cukur, T., et al., Attention during natural vision warps semantic representation across the human brain. Nat Neurosci, 2013. 16(6): p. 763-70.

      (8) Keller, A.S., et al., Attention enhances category representations across the brain with strengthened residual correlations to ventral temporal cortex. Neuroimage, 2022. 249: p. 118900.

      (9) Zhang, J., et al., Behavioral and Neural Mechanisms of Face-Specific Attention during GoalDirected Visual Search. The Journal of Neuroscience, 2024. 44(46): p. e1299242024.

      (10) Bichot, N.P., A.F. Rossi, and R. Desimone, Parallel and serial neural mechanisms for visual search in macaque area V4. Science, 2005. 308(5721): p. 529-534.

      (11) Bichot, N.P., et al., The role of prefrontal cortex in the control of feature attention in area V4. Nat Commun, 2019. 10(1): p. 5727.

      (10) Cohen, M.R. and J.H. Maunsell, Using neuronal populations to study the mechanisms underlying spatial and feature attention. Neuron, 2011. 70(6): p. 1192-204.

      (11) Maunsell, J.H. and S. Treue, Feature-based attention in visual cortex. Trends Neurosci, 2006. 29(6): p. 317-22.

      (12) McAdams, C.J. and J.H. Maunsell, Attention to both space and feature modulates neuronal responses in macaque area V4. J Neurophysiol, 2000. 83(3): p. 1751-5.

      (13) Motter, B.C., Saccadic momentum and attentive control in V4 neurons during visual search. J Vis, 2018. 18(11): p. 16.

      (14) Sapountzis, P., S. Paneri, and G.G. Gregoriou, Distinct roles of prefrontal and parietal areas in the encoding of attentional priority. Proc Natl Acad Sci U S A, 2018. 115(37): p. E8755-E8764.

      (15) Treue, S. and J.C. Martinez Trujillo, Feature-based attention influences motion processing gain in macaque visual cortex. Nature, 1999. 399(6736): p. 575-9.

      (16) Zhou, H. and R. Desimone, Feature-based attention in the frontal eye field and area V4 during visual search. Neuron, 2011. 70(6): p. 1205-17.

    1. eLife Assessment

      This valuable study characterizes how antibody responses converge on similar functional solutions despite diverse genetic backgrounds, providing a resource that is of importance in understanding immune responses and informing vaccine research. The evidence is solid, with extensive and well-executed analyses supporting the primary findings, although broader conclusions regarding vaccine design would benefit from more cautious interpretation and fuller discussion of the study's limitations. The work will be of interest to researchers studying antibody responses, viral evolution, and vaccine development.

    2. Reviewer #1 (Public review):

      Summary:

      Based on previous work showing that viral evolution follows reproducible patterns in diverse animals, the authors sought to examine whether the antibody response operates under similar constraints. By analyzing over 17,000 B cells isolated from 6 monkeys at 3 different time points, the authors convincingly show that the immune response does follow specific patterns of responses to different classes of epitopes based on the infecting virus. Moreover, each of these clusters has characteristic (cross-) binding and neutralization properties. Importantly, these classes are independent of the underlying immunogenetics, which (as expected) vary significantly between monkeys. This last point is particularly relevant for vaccine design, as it means that immunogens may not need to be as narrowly focused on specific germline genes as previously thought.

      Strengths:

      The large number of B cells cultured for this study is a particular strength, as is the fact that they were isolated in an antigen-unbiased fashion. The experiments are well-designed and comprehensive.

      Weaknesses:

      The genetic element is a relatively minor component overall and more qualitative than quantitative. It would be nice to investigate other properties of the repertoire like CDRH3 length and possible public clones, as well.

    3. Reviewer #2 (Public review):

      Summary:

      Song et al. comprehensively analyzed the SHIV-infected macaque B cell repertoires and commonalities among their antibody responses, despite their diverse genetic background. They suggest these studies would inform HIV-1 vaccine design.

      Strengths:

      This study is well-designed and used proper analysis methods, and the figures are clear and effectively presented.

      Weaknesses:

      However, it tends to overstate its novelty and significance, emphasizing points that are relatively obvious (e.g., different classes of antibodies can recognize a common epitope) and appears to have been overwritten and unnecessarily fancy ("conceptually analogous to ecomorph evolution", "epitopic convergence"). Moreover, some limitations of the rhesus macaque model and the differences between bnAbs and nAbs should be discussed. That said, the underlying data are solid and important in their detail, and the manuscript will be a useful resource for HIV-1 vaccine and pathogen studies.

    1. eLife Assessment

      This is an important study reporting a new phenotype for a gene cluster that has previously been associated with the responses of the Gram-negative opportunistic pathogen Pseudomonas aeruginosa to flow fluid. Expression of the froABCD gene cluster is induced by HOCl in vitro and by activated immune cells, which produce these types of reactive chlorine species and the evidence presented by the authors is in many places convincing. Overall, the authors have been responsive to the previous review, although the exact mechanism of fro-induction by HOCl remains unclear. The high cysteine- and methionine content of the anti-sigma factor FroI hints at a direct oxidative modification of this protein during activation of the operon and a corneal infection model shows that fro is upregulated in P. aeruginosa 20 h after infection, but the evidence that HOCl is the causative agent of fro upregulation under these conditions is at present circumstantial. This study is of interest to infection biologists interested in mechanisms of bacterial pathogenicity.

    2. Reviewer #1 (Public review):

      Summary:

      Foik et al. report that hypochlorous acid, a reactive chlorine species generated during host defense, activates the transcription of the froABCD in P. aeruginosa. This gene cluster had previously been associated with a potential role during flow of fluids and appears to be regulated by the sigma factor FroR and its anti-sigma factor FroI. In the present study, the authors show that froABCD is expressed both in neutrophils and macrophages, which they claim is likely a result of HOCl but not H2O2 production. Fro expression is also induced in a murine model of corneal infection, which is characterized by immune cells invasion. Expression of the fro system can be quenched by several antioxidants, such as methionine, cysteine, and others. FroR-deficient cells that lack froABCD expression during HOCl stress, appear more sensitive to the oxidant.

      Strengths:

      The authors provide a number of data supporting their claim that transcription of the froABCD system is induced by reactive chlorine species. This was shown by RNAseq, qRT-PCR, and through microscopy using a transcriptional reporter fusion. Likewise, elevated expression of froABCD was shown in vitro and in vivo, excluding potential in vitro artifacts. The manuscript, while mostly descriptive, is easy to follow and the data were presented clearly and convincingly. The authors have also been responsive to concerns from the previous review.

      Weaknesses:

      (1) Line 10: "HOCl preferentially oxidizes....". Please consider modifying the language to: "the second-order rate constant of HOCl is significantly higher with Met/Cys compared to other aa."

      (2) I am not sure I completely understand Fig 1B. Is the promoter right upstream of yfp or is yfp located downstream of froA? If the latter is the case, wouldn't this be a translational fusion?

      (3) My previous comment regarding why fro expression is higher during phagocytosis in macrophages compared to neutrophils has been somewhat (albeit not convincingly), addressed by the authors in the response to the reviewer, but this discussion should be part of the manuscript as the macrophage data were shown.

      (4) Line 122: The statement "The degree of fro inhibition by 4-ABAH...." is incorrect unless the authors can provide experimental evidence. Fro expression is not upregulated because MPO is inhibited by 4-ABAH, which results in less hOCL production.

      (5) Can Supp Fig. 1 be quantified in a similar way it was done for HOCl to allow for a better comparison if HOCl or flow is the more potent inducer?

      (6) Overall, the fro expression (YFP/mCherry) seems highly variable for treatment with HOCl (Fig. 2C: ~65; 2D: ~20; why is fro expression 3x lower?

      (7) The authors should provide evidence that N-chlorotaurine can activate fro expression also. They said they weren't able to obtain chlorinated taurine, but this is quite simple to produce: PMCID: PMC1219228

      (8) Fig. 4 supplement 1: Please provide concentrations for the oxidants used in these experiments.

      (9) Lines 251/252: change to: upregulation of instead of in

      (10) Chaperones and other heat-shock genes are more upregulated in ∆froR, indicating elevated HOCl-mediated oxidative damage, which supports their findings.

      (11) Complementation of ∆froR is missing

      (12) Line 198: The growth experiment at 4 uM shows differences between WT and mutant, but at 2 uM cells showed already low fro expression due to cell death (which has not been proven by CFU counts). This discrepancy should at least be discussed.

      (13) The critical in vitro experiment is missing: does purified FroI get oxidized by HOCl and dissociated from FroR?

      (14) Lines: 350-355: The claim that the fro system is the first-line defense is unproven.

    3. Reviewer #2 (Public review):

      Summary:

      Foik et al. studied the regulation of the fro operon in response to HOCl, an oxidant derived from immune cells, especially neutrophils. They use a transcriptional fusion of YFP to the froA promotor in an mCherry expressing P. aeruginosa strain to determine fro-induction under the microscope. They use this system to study fro expression in medium, in the presence of neutrophils and macrophages, neutrophil-conditioned medium, and several chemical stimuli, including NaCl, HOCl, hydrogen peroxide, nitric acid, hydrochloric acid, and sodium hydroxide. They also use a corneal infection model to demonstrate that froA is upregulated in P. aeruginosa 20 h post infection and perform transcriptional analyses in WT and a froR mutant in response to HOCl.

      Strengths:

      Their data clearly shows that HOCl is a strong inducer of the fro Operon. Addition of HOCl-quenching chemicals together with HOCl abrogates the response. They also show that a froR mutant is more susceptible to HOCl than WT. Their transcriptomic data reveals genes under control of the FroR/FroI sigma factor/anti sigma factor system.

      Weaknesses:

      Although the presented evidence is mostly solid, some of their findings need to be evaluated more carefully; explaining the rationale behind some of the experiments might enhance the article; and some of the models proposed by the authors seem far-fetched, as outlined below:

      Unexpected outcomes and open questions for future research:

      (1) As outlined above, HOCl seems to be the main inducer of the fro operon. Interestingly, during interaction with immune cells, macrophages and neutrophils seem to induce a reporter gene under fro control in a similar manner, although macrophages are generally thought to produce less HOCl, when compared to neutrophils. May be this view needs to be revised, or another reactive species, produced by macrophages, can activate the fro operon as well.

      (2) HOCl is typically unstable in the presence of biomolecules. Nevertheless, medium conditioned by activated neutrophils is a strong inducer of the fro operon. The medium used by the authors for this experiment contains taurine, and, as the authors acknowledge, this taurine will likely react with HOCl to form the more stable taurine N-chloramine. Similarly, the MinA bacterial medium used to treat P. aeruginosa with HOCl directly also contains ammonium ions at mM concentrations, which could potentially react with HOCl to form monochloramine. It could be speculated that taurine N-chloramine and other chloramines are as effective as HOCl in activating the fro-operon.

      (3) The fro operon was originally described to be activated by shear stress ("flow-regulated operon"). How shear stress and HOCl-stress are related, or if fro activation by both stimuli is a coincidence, remains unclear. The authors propose a model, in which flow transports oxidizing molecules, which ultimately activate the fro operon. However, the initial work by Sanfilippo et al. (2019, Nat Microbiol) used plain LB medium in a fluidic chamber to induce the shear stress, which should be free of oxidants, and certainly of HOCl.

      Comments on revised version:

      The authors have addressed my concerns appropriately.

    4. Author response:

      The following is the authors’ response to the original reviews.

      We greatly appreciate the efforts of the reviewers, which have provided insightful and helpful comments to improve the manuscript. The feedback touches upon a number of topics, focusing on clarification or justification of experimental techniques and on understanding the mechanism by which P. aeruginosa detects HOCl. All reviewers raised the issue of how HOCl activates fro expression, including whether free or protein-bound methionine, cysteine, or other HOCl byproducts induce this expression. For the upcoming revision, we plan to perform experiments that address this issue and will discuss potential mechanistic models in light of the new data. In addition, we plan to perform additional experiments to address a reviewer’s concerns regarding the dependence of the fro response on HOCl production by neutrophils. The revision will correct imprecise statements pointed out by reviewers, and address all remaining issues requiring clarification or further discussion, including the range of HOCl sensitivity, relationship between HOCl and flow sensitivity, and justification for testing the fro response to nitric acid.

      We have completed a number of experiments and responded thoroughly to reviewer comments below. We thank the reviewers again for their details comments, which suggested additional experiments and interpretations that have resulted in significant additional insight into the potential mechanism of HOCl sensing and its relevance with neutrophils.

      Reviewer #1 (Public review):

      Summary:

      Foik et al. report that hypochlorous acid, a reactive chlorine species generated during host defense, activates the transcription of the froABCD in P. aeruginosa. This gene cluster had previously been associated with a potential role during the flow of fluids and appears to be regulated by the sigma factor FroR and its antisigma factor FroI. In the present study, the authors show that froABCD is expressed both in neutrophils and macrophages, which they claim is likely a result of HOCl but not H2O2 production. Fro expression is also induced in a murine model of corneal infection, which is characterized by immune cell invasion. Expression of the fro system can be quenched by several antioxidants, such as methionine, cysteine, and others. FroR-deficient cells that lack froABCD expression during HOCl stress appear more sensitive to the oxidant.

      Strengths:

      The authors provide a number of data supporting their claim that transcription of the froABCD system is induced by reactive chlorine species. This was shown by RNAseq, qRT-PCR, and through microscopy using a transcriptional reporter fusion. Likewise, elevated expression of froABCD was shown in vitro and in vivo, excluding potential in vitro artifacts. The manuscript, while mostly descriptive, is easy to follow, and the data were presented clearly.

      We greatly appreciate the efforts of the reviewer and thank them for their succinct summary of the manuscript.

      Weaknesses:

      (1) Lines 60-62: Some of the authors' conclusions are not supported by the data and thus appear unfounded. One example: "we determine that fro upregulation.....These data suggest a novel mechanism..." Their data do not show that MSR upregulation is a direct effect of FroABCD. Instead, it could be possible that the FroR sigma factor also controls the expression of msr genes, which would be independent of froABCD.

      We thank the reviewer for pointing out this important distinction. We have clarified in lines 63-65 in the clean version of the revision that MSR upregulation depends on FroR rather than FroABCD.

      (2) The authors show increased fro transcription both in neutrophils and macrophages; however, the two types of immune cells differ quite dramatically with respect to myeloperoxidase activation and HOCl production.

      Neither has this been discussed nor considered here.

      We agree that the distinction between the cell types is important and have added a brief description of the differences in respiratory bursts and ROS production between the two cell types and our justification for focusing on neutrophils in lines 102-103. We think it’s very interesting that Fro appears to be activated by macrophages, which are not associated with HOCl production on their own. We think that it would be interesting to identify what is inducing Fro in macrophages in future work.

      (3) With respect to the activation of fro expression upon challenge with conditioned media from stimulated neutrophils, does the conditioned media contain detectable amounts of HOCl? Do chloramines, which are byproducts of HOCl oxidation with amines, also stimulate expression?

      This is an excellent question that addresses which molecules Fro is responding to from neutrophils. We have performed additional experiments that confirm that PMA-stimulated neutrophils produce HOCl (Fig. S2) through the use of a commercial hypochlorite sensor assay, which claims high specificity for detecting HOCl. We further confirmed that this production is inhibited by pretreatment with the MPO-specific inhibitor 4ABAH. These data support the interpretation that the Fro response to stimulated neutrophils requires MPO activity, of which the major product is HOCl. We have described this in lines 117-127.

      Our data does not exclude the possibility that other MPO products could activate Fro expression. We were unable to obtain a reliable source of the major secondary MPO product, taurine chloramine, for our experiments, unfortunately. We believe that understanding the potential for secondary products to activate the response is an important and interesting question that can be explored in a future study. We have added a discussion of this in lines 367-376.

      (4) A better control to prove that this fro expression is indeed induced by HOCl in activated neutrophils would be to conduct the experiments in the presence of a myeloperoxidase inhibitor.

      We thank the reviewer for raising this point. We have performed the suggested set of experiments and found that indeed, the pre-treatment of neutrophils with MPO inhibitor 4-ABAH prior to PMA stimulation suppresses the activation of fro (Fig. 2D and Figure 2-figure supplement 1-2). The results are discussed in lines 117-127.

      (5) The work was conducted with two different P. aeruginosa strains (i.e. AL143 and PAO1F). None of the figure legends provides details on which strain was used. For instance, in line 111, the authors refer to Figure S1B for data that I thought were done with PAO1F, while in 154, data were presented in the context of the infection model, which was conducted with the other strain.

      We thank the reviewer for pointing this issue out. We have ensured that strain names appear in all the revised figure legends. To clarify, only mouse experiments and a related RT-qPCR assay used strain PAO1F due to prior IACUC approval of this strain and its use in previous publications.

      (6) It would be good if immune cell recruitment at 2hrs and 20hrs PI could be quantified.

      We previously quantified neutrophil recruitment at the site of corneal abrasion at 24 hours using the same conditions and strains (Ratitong, B. et al., J. Immun, 2022). While we do not have immune cell recruitment data for the 2 hr and 20 hr time points, the previous data show significant neutrophil recruitment near the latter time point, which is consistent with the interpretation that fro expression is activated by stimulated neutrophils. We have discussed this in lines 212-216.

      (7) The conclusions of Figure 4 are, in my opinion, weak (line 187-188; "It is possible that ....."). These antioxidants likely quench the low amounts of NaOCl directly. This would significantly reduce the NaOCl concentrations to a level that no longer activates expression of fro. There is no direct evidence provided that oxidized methionine induces fro expression. Do the authors postulate that this is free methionine, or could methionine and/or cysteine oxidation in FroR increase the binding affinity of the sigma factor to the promoter? Another possibility is that NaOCl deactivates the anti-sigma factor. None of these scenarios has been considered here.

      We acknowledge that our model of HOCl sensing was unclear and thank the reviewer for their insight. This critique is echoed by reviewer #2 in comment 3 as well. We recently found that the FroI anti-sigma factor has the highest concentration of methionine and cysteine residues of all known P. aeruginosa anti-sigma factors (Appendix 2—Table 1). Given that FroR and FroI form an extracytoplasmic function sigma – anti-sigma pair, which are associated with transducing extracellular signals to the cytoplasm, we have proposed an alternative model in which HOCl or secondary RCS molecules are detected through their oxidation of cysteine and methionine residues in FroI. This is discussed in lines 271-279 and lines 363-367.

      (8) Line 184: The reaction constants of HOCl with Cys and Met are similar.

      We thank the reviewer for pointing out this important clarification. We have revised the sentence to accurately reflect this in lines 243-244.

      (9) Treatment with 16 uM NaOCl caused a growth arrest of ~15 hrs in the WT (Figure 5A), whereas no growth at all was recorded with 7.5 uM in Figure 3A.

      We thank the reviewer for catching this. We have determined that the concentration of NaOCl in the reagent used for this particular experiment was lower than expected, thus requiring a higher concentration to achieve growth inhibition. We have repeated the experiment with new reagent and find that the results (now Figure 6A) are similar to the previous experiment but at a lower concentration of 4 micromolar, consistent with the concentration found to be sub-inhibitory in Figure 3A.

      (10) The concentration range of NaOCl causing fro expression is extremely narrow, while oxidative burst rapidly generates HOCl at much higher concentrations. This should be discussed in more detail.

      We appreciate the reviewer’s comment, which is related to reviewer #2’s comment #9. We have clarified the reported production rates of HOCl, which far surpass the bacterial MIC. After greater consideration, we believe secondary HOCl products including taurine chloramine could have a more significant role in vivo. While this molecule is less potent than HOCl, it is longer-lived, retains bactericidal activity, and retains the ability to oxidize methionine. We have discussed this in lines 377-405.

      Reviewer #1 (Recommendations for the authors):

      (1) Some statements in the text don't match the data shown in the Figures. For instance:

      (a) Figure 2B shows ~65-fold fro expression, but the text states: "...increased expression of fro expression by 30fold..."(line 88).

      The YFP/mCherry value of PMA-stimulated is 67.4 in the Figure (now Figure 2C). The fold-change is computed relative to unstimulated conditioned medium (third column, which has a value of 2.2), which is a 30-fold change. We have added a citation in the main text to the Source Data, which provides these raw values, to help clarify the computation for readers, and added in the legend that the value for unstimulated is greater than 1.

      (b) Figure 3B shows ~35-fold fro expression at 1 uM NaOCl, but the text states: "...NaOCl increased fro expression by up to 74-fold..."(line 110).

      We have clarified that the increase is relative to untreated, added that untreated value is below 1 in the caption, and provided a citation in the main text to the Source Data, which contains the raw values. The change is measured relative to untreated, for which the YFP/mCherry value is 0.48. The value at 1 uM is 35.7, giving a 74-fold change.

      (2) Line 229: While the ∆froR strain was sensitive to HOCl, the strain was tolerant". Please revise.

      We have corrected this typo (now lines 328- 329). We meant to convey that growth was not entirely inhibited in the froR strain.

      Reviewer #2 (Public review):

      Summary:

      Foik et al. studied the regulation of the fro operon in response to HOCl, an oxidant derived from immune cells, especially neutrophils. They use a transcriptional fusion of YFP to the froA promoter in an mCherry-expressing P. aeruginosa strain to determine fro-induction under the microscope. They use this system to study fro expression in medium, in the presence of neutrophils and macrophages, neutrophil-conditioned medium, and several chemical stimuli, including NaCl, HOCl, hydrogen peroxide, nitric acid, hydrochloric acid, and sodium hydroxide. They also use a corneal infection model to demonstrate that froA is upregulated in P. aeruginosa 20 h post-infection and perform transcriptional analyses in WT and a froR mutant in response to HOCl.

      Strengths:

      Their data clearly shows that HOCl is a strong inducer of the fro Operon. The addition of HOClquenching chemicals together with HOCl abrogates the response. They also show that a froR mutant is more susceptible to HOCl than WT. Their transcriptomic data reveal genes under control of the FroR/FroI sigma factor/anti sigma factor system.

      Weaknesses:

      Although the presented evidence is mostly solid, some of their findings need to be evaluated more carefully; explaining the rationale behind some of the experiments might enhance the article, and some of the models proposed by the authors seem far-fetched, as outlined below:

      We greatly appreciate the reviewer’s efforts and thank them for highlighting strengths and areas for improvement.

      (1) In line 76 the authors claim "Relative to P. aeruginosa that were incubated in host cell-free media, P. aeruginosa in close proximity to human neutrophils or that were engulfed in mouse macrophages appeared to increase fro expression (Fig. 1C)". Counting bacterial cells in Figure 1C shows that 1 in 17 bacteria (5.8%) induce the froA-promotor in media in the absence of immune cells, while 4 in 72 bacteria (only 5.5%) do the same in the presence of neutrophils. Contrary to the authors' claims, it appears that P. aeruginosa actually decreases fro-expression in close proximity to neutrophils. There is a slight increase in fro-expression in bacteria co-incubated with macrophages (3 in 21, or 14.3%). A more rigorous statistical analysis might substantiate the authors' claim, but, as is, the claim "neutrophils increase fro expression" is untenable.

      We believe the images alone do not give an adequate representation of the data and have quantified a larger portion of the data, which has been added as Figure 1D. The quantification supports the original claim that fro expression is increased during co-incubation with macrophages and neutrophils. Since there was not sufficient statistical sampling to distinguish engulfed P. aeruginosa from free ones, this part of the claim has been removed from the text (updated in lines 80-84).

      (2) The authors should explain the rationale behind some of the chemicals used. Why did they use nitric acid? Especially at these high concentrations, a strong acid such as nitric acid might have a significant influence on the medium pH. I understand that the medium is phosphate-buffered, but 25 mM nitric acid in an unbuffered medium would shift the pH well below 2. Similar considerations apply to hydrochloric acid and sodium hydroxide.

      We thank the reviewers for pointing out the need for this clarification. We have updated Figure 3D with a lower concentration of NaOH at 1 uM, which is the same concentration as NaOCl that activates fro expression. Due to the high buffering capacity of our medium, a high concentration of 6 mM NaOH was needed to induce a discernible change in pH and this high concentration of NaOH had no obvious effect on growth. Neither 1 uM nor 6 mM NaOH produced a change in fro expression, consistent with our previous findings that the effect is not due to sodium ions or higher pH. These updated findings are described in lines 181-190.

      Since the effect of chloride is already controlled for using NaCl and the concentration of HCl used was not sufficient to cause a significant change in pH in the buffered medium, we have removed the HCl group from the data.

      We have clarified that nitric acid was used because it is a strong oxidizer that is not found in neutrophils and that concentrations used were near the minimal inhibitory concentrations (lines 145-147 and lines 191-198). We acknowledge that the growth inhibition from HNO<sub>3</sub> could be due to pH or oxidation. However, since no change in fro expression was observed at concentrations approaching the inhibitory concentration, we did not address the potential effects of low pH from nitric acid on fro expression.

      (3) In line 187, the authors state that "It is possible that oxidized methionine increases fro expression" and they suggest a model to that effect in Figure 5D. It is unclear why the authors singled out methionine sulfoxide, since a number of other things get oxidized by HOCl. In line 184, the authors state, in the same vein, that "HOCl oxidizes methionine residues 100-fold more rapidly than other cellular components". The authors should state which other cellular compounds they are referring to. Certainly not cysteine and other thiols, which react equally fast and are highly abundant in the cell: P. aeruginosa contains 340 µM GSH, 140 µM CoA-SH (https://doi.org/10.1074/jbc.RA119.009934) plus free cysteine and cysteines in proteins (based on codon usage, 1.34% of amino acids in proteins are cysteine, while methionine is only slightly more present at 2.10%, although a number of starting methionines are removed from mature proteins).

      We acknowledge that our HOCl sensing model had been vague and unclear and thank the reviewer for their insight. This critique is echoed by reviewer #1 in comment 7 as well.

      Our initial suggestion that methionine sulfoxide was sensed was motivated by the observation that methionine sulfoxide reductases are upregulated by HOCl. However, we have revised this based on feedback from reviewers and further consideration of chlorine redox chemistry. Interestingly, we found that the FroI anti-sigma factor has the highest concentration of methionine and cysteine residues of all the P. aeruginosa anti-sigma factors (Appendix 2—table 1). Given that FroR and FroI form an extracytoplasmic function sigma – anti-sigma pair, which are associated with transducing extracellular signals to the cytoplasm, we propose a model in which HOCl or secondary RCS molecules are detected by their oxidation of cysteine and methionine residues in FroI. This is discussed in lines 271-278 and lines 362-366.

      (4) Overall (and this is probably not addressable with the authors' data), some very interesting questions remain unanswered: what is the molecular mechanism of fro-induction? How is the FroR/FroI system modulated by HOCl? Does the system sense free or protein-bound methionine-sulfoxide? Are certain methionine residues in these proteins directly oxidized by HOCl? Many "HOCl-sensing" proteins are also modified at cysteine residues or amino groups; could those play a role? And lastly: what is the connection between shear/fluid flow and HOCl, or are these totally separate mechanisms of fro-induction?

      We thank the reviewer for raising these excellent mechanistic questions. Issues relating to HOCl sensing are addressed in the preceding comment.

      Regarding the connection to shear sensing, Padron et al., 2023 found that the detection of flow in P. aeruginosa can be attributed to chemical transport, in particular to H<sub>2</sub>O<sub>2</sub> that was present in growth media. Based on the same principle, we expect Fro to be upregulated in flow at much lower concentrations than those observed in stationary fluids. The activation of Fro and the effects of HOCl would thus be expected to be flow-sensitive. We have commented on this important factor in the discussion in lines 406-417.

      Reviewer #2 (Recommendations for the authors):

      (1) To address 1, the authors could evaluate the microscopic images in the same manner in which they evaluated the other microscopic images, as, for example, presented in Figure 2B or Figure S1B.

      See response to Weakness point (1).

      (2) To address 2, please explain the choice of the chemicals (why nitric acid?), but also provide the pH of the media with those high concentrations of strong acids and bases, and interpret them in light of the permissible pH range for P. aeruginosa growth. More sensible controls might be a lower NaOH concentration in the range that would be reached through the amount of NaOH in the NaOCl stock at the highest NaOCl concentrations used. As for the acids, I don't see a reason to use these acids at these high concentrations. Please explain.

      See response to Weakness point (2).

      (3) To address 3: The authors could specify their methionine-sulfoxide model a bit more, so that testable hypotheses can be developed. If the authors think the FroR/FroI system senses free methionine sulfoxide, they or others could add methionine sulfoxide to the medium and check induction. If they think specific methionine residues in these proteins are oxidized, they could provide evolutionary evidence of conserved methionine residues. Or, based on a structure or structural prediction, they (or others) could mutate methionine residues, e.g., at the protein's surface or potential protein/protein-interaction sites and assess the effect on HOCl-based activation. Or they could consider other amino acids known to be highly reactive towards HOCl and mutate those in a future study.

      See response to Weakness point (3).

      Further comments:

      (4) The headline of the figure legend of Figure S1 seems incomplete. Please mention the flow experiments shown in Figure S1A.

      We have updated the title of this figure, which now appears in the eLife format as Figure 1 – figure supplement 1.

      (5) What is the difference between the data presented in Figure 4C (bars "UTR" and "NaOCl") and the same bars in Figure S2A? Is this redundant or a re-plot?

      In this revision, Figure 4C has become Figure 5B and Figure S2A is now Figure 5—figure supplement 1. Only the 1 uM NaOCl condition is replotted. We have described this in the legend for Figure 5—figure supplement 1.

      (6) Line 228: hpd is more likely a gene of the aromatic amino acid catabolism.

      We thank the reviewer for pointing this out. We have removed the ‘branched’ descriptor in this sentence, now in lines 325-326.

      (7) Line 229: "While the ΔfroR strain was sensitive to HOCl, the strain was tolerant (Figure 5A)". Please clarify. Which strain was tolerant? WT?

      We have corrected this typo (now line 328-329). We meant to convey that growth was not entirely inhibited in the froR strain.

      (8) Line 247: "Activated neutrophils produce HOCl concentrations as high as 50 µM [24]." This "50 µM" number is often quoted; the citation trail typically leads to Weiss et al. 1982 (https://doi.org/10.1172/JCI110652). However, a more factually correct statement based on that paper would be "2 x 10^6 neutrophils, activated with 30 ng/mL PMA at 37C in 1 mL of Dulbecco's buffer can produce around 50 nmol HOCl per hour". In the particular reference 24, Dybpukt et al used a methodology similar to Weiss et al., and here around 50nmol were produced by the same number of cells in 30 min in response to 100 ng/mL PMA. Please clarify accordingly.

      We thank the reviewer for bringing this to our attention. We have altered the language to indicate that the production rate is for this specific set of parameters. Related to this, Reviewer #1 Comment #10 requested a more detailed discussion of the significance of the Fro response, since it is at much lower HOCl concentration than produced by neutrophils. We have discussed this in lines 377-405.

      (9) Line 275: "which is strain PA14 strain". Please clarify.

      We have fixed this error and entered it into the Key Resource Table.

      (10) Line 311: "Cultures containing densities below 10 P. aeruginosa per frame were concentrated using a syringe filter with 0.2 or 0.8 µm pore sizes (Millipore, Burlington, MA)." Isn't that a bit concerning when testing the induction of an operon that is supposedly activated by shear through fluid flow? Did the authors convince themselves that this procedure does not induce fro?

      We do not expect the filtering procedure to cause changes in gene expression because cells are imaged immediately after filtering. Nonetheless, we performed additional experiments (Fig 3F and 2D) entirely without concentrating cells and found that YFP/mCherry levels were consistent with previous data in Figure 3. This rationale has been added to the Methods section under the “Fluorescence and Phase Contrast Microscopy” section.

    1. eLife Assessment

      This study provides important insights into the neural mechanisms linking sleep and long-term memory consolidation. By combining behavioural, genetic, imaging, and connectomic approaches in Drosophila, it identifies a target neural circuit that will be of broad interest to researchers studying sleep, memory, and neural circuits. The evidence supporting the involvement of the identified circuit in the regulation of sleep and memory is solid and represents a substantial advance in the field. Nevertheless, there is limited evidence to support the mechanistic claim that this circuit directly links sleep and memory consolidation within the available data, and some results should therefore be interpreted with appropriate caution.

    2. Reviewer #1 (Public review):

      [Editors' note: this version has been assessed by the Reviewing Editor without further input from the original reviewers. The authors have considered and discussed the comments raised in the previous round of review.]

      Summary:

      The authors aim to use state-of-the art behaviour, imaging and connectome techniques to identify the neural interaction between sleep and long-term memory consolidation in the PAM-DPM circuits, a well-known dopaminergic pathway within Drosophila Mushroom Body.

      Strengths:

      The investigation follows a logical strategy to collect huge dataset of sleep, appetitive memory and live imaging. The authors identified and showed that activation of a PAM subset: alpha-1 reduces sleep quality and memory consolidation in a starvation dependant manner. The author also convincingly demonstrated the corresponding neuronal responses of DPM neurons following PAM alpha-1 activation, and the positive role of DPM neural activity in sleep and memory consolidation. Moreover, the new data provide TRIC-LUC provided better temporal resolution of neural activity correlates for PAMalpha1-DPM inhibition. Importantly, the author demonstrated that memory loss derived from PAM alpha 1 activation can be partly restored by ectopic sleep enhancement via feeding THIP at the memory consolidation period after training.

      Weaknesses:

      Although the revised version carries arguments to satisfy the reviewers' concern, the writing is now less cohesive. Crucially an explanation however remains required for the following experimental contradiction: the central observation of the study indicates that PAM alpha1 activation cause DPM inhibition which disrupt sleep and memory consolidation. Therefore, one would expect a reduced PAMalpha1 and increased DPM activities after memory training, but the authors found the opposite is true from now enhanced TRIC-LUC dataset. The authors indicate this data reinforce the inhibitory nature of PAM-alph1-DPM, but it does not explain why such a reduced DPM activity is observed after training.

    3. Reviewer #2 (Public review):

      Summary:

      Sleep plays a critical role in memory consolidation, but the neural mechanisms underlying this relationship remain incompletely understood. The authors examined a specific subset of PAM dopaminergic neurons, PAM-α1, and DPM neurons in Drosophila. These neurons have previously been implicated in memory, and DPM neurons have also been linked to sleep. The study explores whether this circuit provides a mechanistic link between sleep and memory consolidation.

      Strengths:

      The authors report several novel findings. Brief activation or inhibition of PAM-α1 neurons, or brief inhibition of DPM neurons during the first few hours after training, impairs 24-hour LTM. Notably, these brief manipulations disrupt sleep for many hours afterward, particularly during the night. The authors further show that perturbation of PAM-α1 and DPM neurons impairs sleep and appetitive memory consolidation under starvation conditions, and that pharmacological sleep induction during the night rescues the LTM defects. Together, these findings suggest that PAM-α1 and DPM neurons are involved in sleep regulation and LTM consolidation under starvation. These are important observations that advance our understanding of the circuits regulating sleep and memory consolidation.

      Weaknesses:

      Some claims require additional evidence or clarification.

      (1) Previous studies linking impaired memory to reduced sleep have primarily examined conditions involving severe sleep deprivation. In contrast, this manuscript argues that relatively modest decreases in total sleep, accompanied by sleep fragmentation, are sufficient to impair memory consolidation. It remains unclear whether sleep fragmentation of this magnitude is itself critical for LTM consolidation. An independent method for inducing comparably mild sleep loss and fragmentation would be needed to directly test this interpretation.

      (2) It is unclear why both activation and inactivation of PAM-α1 neurons produce similar effects on sleep and memory. In addition, MB299B-labeled neurons exert stronger effects on memory than MB043B-labeled neurons, whereas MB043B-labeled neurons have stronger effects on sleep. If sleep disruption is the primary driver of impaired memory consolidation, a stronger correspondence between the sleep and memory phenotypes might be expected. The authors speculate that MB043B may affect sleep through non-PAM neurons, but without identifying the relevant neurons, this remains speculative.

      (3) The complex schematic model (Fig. 12), with parallel circuits and unidentified neuronal groups, underscores the difficulty of interpreting the current data. In the "less activity" arm of the model, distinct circuits are proposed to regulate sleep and LTM, respectively, and DPM neurons are not included. This makes it difficult to reconcile the model with the central claim that the PAM-α1-to-DPM microcircuit links sleep and LTM consolidation.

      (4) The TRIC-LUC reporter system is not ideal for resolving dynamic changes in neuronal activity. Activity-dependent Ca²⁺ signaling must first reconstitute the TRIC transcriptional system, which then drives luciferase transcription, translation, and accumulation. The original characterization of TRIC indicates that TRIC signals accumulate and decay over several hours. Thus, the kinetics of the TRIC-LUC reporter should be interpreted cautiously, particularly when inferring transient or precisely timed changes in neuronal activity.

      (5) Including data from training under fed conditions would provide a more complete understanding of state-dependent neural activity and would help distinguish starvation-specific effects from more general circuit mechanisms.

    4. Reviewer #3 (Public review):

      Summary:

      Understanding the neural circuits that link sleep and memory remains a fundamental challenge in neuroscience. In this study, Lin Yan and colleagues investigate how dopamine signaling in Drosophila regulates long-term memory (LTM) formation in the context of sleep. They identify a specific microcircuit between protocerebral anterior medial dopamine neurons (PAM-DANs) and dorsal paired medial (GABAergic DPM) neurons that modulates memory consolidation. Their findings suggest that disrupting the basal activity of PAM-α1 neurons during early consolidation impairs LTM, with particularly pronounced effects under starvation conditions. Notably, sleep fragmentation caused by this disruption can be pharmacologically rescued, restoring LTM. These results provide compelling evidence how dopamine signaling plays a crucial role in linking sleep and memory, offering new insights into the underlying mechanisms.

      Strength:

      This study presents a well-executed investigation into sleep-memory interactions, utilizing a combination of connectomics, behavioral assays, functional imaging, and pharmacological manipulations. The authors convincingly demonstrate that the PAM-α1 and DPM circuit interact, highlighting a potential mechanism by which sleep influences memory consolidation. The anatomical and functional dissection of this circuit is of high interest to the field, and the study's integration of sleep and memory processes contributes significantly to our understanding of the role of dopamine in cognitive functions. Additional experiments investigating the contribution of MBON-α1 to the circuit, connectomic analysis together with a dissection of dopamine receptor function further strengthen the proposed circuit motif and its biological relevance.

      Weaknesses:

      While the study is well designed, presents compelling findings and has been further strengthened by additional experiments, some aspects remain unclear. The role of DPM neurons in memory consolidation seems not yet fully resolved, as different genetic approaches yield variable results. Furthermore, some manipulations impair memory without affecting sleep fragmentation - or vice versa, suggesting that the observed memory deficits cannot be explained solely by impaired sleep-dependent consolidation. It would also have been interesting to discuss potential mechanisms by which dopamine receptor-mediated cAMP signaling could lead to a reduction in Ca²⁺ signals. I am confident that these questions can be addressed in future studies.

      Conclusion:

      Overall, this study provides valuable new insights into how sleep and dopaminergic circuits interact to regulate memory consolidation in Drosophila and may reveal general principles underlying the neural regulation of memory.

    5. Author response:

      The following is the authors’ response to the previous reviews

      eLife Assessment

      This study provides important insights into the neural mechanisms linking sleep and long-term memory consolidation. By combining behavioural, genetic, imaging, and connectomic approaches in Drosophila, it identifies a target neural circuit that will be of broad interest to researchers studying sleep, memory, and neural circuits. The evidence supporting the involvement of the identified circuit in the regulation of sleep and memory is solid and represents a substantial advance in the field. Nevertheless, there is limited evidence to support the mechanistic claim that this circuit directly links sleep and memory consolidation within the available data, and some results should therefore be interpreted with appropriate caution.

      We appreciate the reviewer’s careful evaluation of our manuscript, and we agree that (as with any experimental study) there is still a lot to do to fully understand the mechanisms underlying the linkage between sleep and memory consolidation.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The authors aim to use state-of-the art behaviour, imaging and connectome techniques to identify the neural interaction between sleep and long-term memory consolidation in the PAM-DPM circuits, a well-known dopaminergic pathway within Drosophila Mushroom Body.

      Strengths:

      The investigation follows a logical strategy to collect huge dataset of sleep, appetitive memory and live imaging. The authors identified and showed that activation of a PAM subset: alpha-1 reduces sleep quality and memory consolidation in a starvation dependent manner. The author also convincingly demonstrated the corresponding neuronal responses of DPM neurons following PAM alpha-1 activation, and the positive role of DPM neural activity in sleep and memory consolidation. Moreover, the new data provide TRIC-LUC provided better temporal resolution of neural activity correlates for PAMalpha1-DPM inhibition. Importantly, the author demonstrated that memory loss derived from PAM alpha 1 activation can be partly restored by ectopic sleep enhancement via feeding THIP at the memory consolidation period after training.

      Weaknesses:

      Although the revised version carries arguments to satisfy the reviewers' concern, the writing is now less cohesive. Crucially an explanation however remains required for the following experimental contradiction: the central observation of the study indicates that PAM alpha1 activation cause DPM inhibition which disrupt sleep and memory consolidation. Therefore, one would expect a reduced PAMalpha1 and increased DPM activities after memory training, but the authors found the opposite is true from now enhanced TRIC-LUC dataset. The authors indicate this data reinforce the inhibitory nature of PAM-alph1-DPM, but it does not explain why such a reduced DPM activity is observed after training.

      We thank the reviewer for their point of view. We have faithfully reported all experimental observations acquired from our enhanced TRIC-LUC dataset as objectively as possible. We note that the reviewer postulates a particular expected activity shift (reduced PAM-α1 activity and elevated DPM activity after memory training), yet it is not clear why. As our data show, this is a highly connected microcircuit and it is not easy to predict how it might change. Our data show that there is change and dismissing our empirically measured results purely based on a theoretical expectation is not justified. The brain operates as an intricately interconnected network; neural activity dynamics cannot always be simply inferred from static circuit polarity. Progress in deciphering neural circuit function relies on iterative rounds of experimental testing. Like most neuroscience investigations, the present study cannot resolve every open question, and we explicitly acknowledge several unresolved directions worthy of future exploration in the Discussion section.

      Reviewer #2 (Public review):

      Summary:

      Sleep plays a critical role in memory consolidation, but the neural mechanisms underlying this relationship remain incompletely understood. The authors examined a specific subset of PAM dopaminergic neurons, PAM-α1, and DPM neurons in Drosophila. These neurons have previously been implicated in memory, and DPM neurons have also been linked to sleep. The study explores whether this circuit provides a mechanistic link between sleep and memory consolidation.

      Strengths:

      The authors report several novel findings. Brief activation or inhibition of PAM-α1 neurons, or brief inhibition of DPM neurons during the first few hours after training, impairs 24-hour LTM. Notably, these brief manipulations disrupt sleep for many hours afterward, particularly during the night. The authors further show that perturbation of PAM-α1 and DPM neurons impairs sleep and appetitive memory consolidation under starvation conditions, and that pharmacological sleep induction during the night rescues the LTM defects. Together, these findings suggest that PAM-α1 and DPM neurons are involved in sleep regulation and LTM consolidation under starvation. These are important observations that advance our understanding of the circuits regulating sleep and memory consolidation.

      Weaknesses:

      Some claims require additional evidence or clarification.

      (1) Previous studies linking impaired memory to reduced sleep have primarily examined conditions involving severe sleep deprivation. In contrast, this manuscript argues that relatively modest decreases in total sleep, accompanied by sleep fragmentation, are sufficient to impair memory consolidation. It remains unclear whether sleep fragmentation of this magnitude is itself critical for LTM consolidation. An independent method for inducing comparably mild sleep loss and fragmentation would be needed to directly test this interpretation.

      We appreciate the reviewer’s suggestion. While alternative assays for inducing sleep loss or sleep fragmentation are indeed available, this line of investigation lies beyond the core scope of the present study. We will certainly take this valuable suggestion into consideration for the future studies.

      Regarding the question of whether sleep fragmentation of this magnitude per se is critical for long-term memory consolidation, we would like to highlight relevant published evidence. Prior work in rodents and human (Bonnet and Arand, 2003; Van Someren et al., 2015; Ramesh et al., 2012; Baud et al., 2014) and our earlier study (Liu et al., 2019) have demonstrated that alterations in sleep architecture, independent of changes in total sleep amount, can affect multiple physiological processes, including memory. Given this existing supporting evidence, we respectfully argue that this does not constitute a weakness of the present manuscript.

      (2) It is unclear why both activation and inactivation of PAM-α1 neurons produce similar effects on sleep and memory. In addition, MB299B-labeled neurons exert stronger effects on memory than MB043B-labeled neurons, whereas MB043B-labeled neurons have stronger effects on sleep. If sleep disruption is the primary driver of impaired memory consolidation, a stronger correspondence between the sleep and memory phenotypes might be expected. The authors speculate that MB043B may affect sleep through non-PAM neurons, but without identifying the relevant neurons, this remains speculative.

      The concern raised by the reviewer that certain interpretations remain speculative represents a common situation in most published research. This interesting direction warrants further investigation in future work, but falls outside the scope of the present study. In the revised manuscript, we have elaborated on the differences observed between these two GAL4 drivers. We have also conducted additional experiments to investigate a well-characterized memory-related recurrent loop of PAM-α1 neurons in sleep regulation. We respectfully note that no single study can comprehensively address all outstanding questions.

      (3) The complex schematic model (Fig. 12), with parallel circuits and unidentified neuronal groups, underscores the difficulty of interpreting the current data. In the "less activity" arm of the model, distinct circuits are proposed to regulate sleep and LTM, respectively, and DPM neurons are not included. This makes it difficult to reconcile the model with the central claim that the PAM-α1-to-DPM microcircuit links sleep and LTM consolidation.

      We appreciate this careful comment on our schematic model in Figure 12. This diagram aims to summarize the key findings obtained in the present study while also explicitly laying out unresolved questions that await future investigation. In our view, including open, outstanding questions in the working model does not undermine the interpretation of our existing experimental results. In the revised manuscript, we have modified the corresponding text to clarify this point and distinguish firmly between conclusions supported by our data and tentative components requiring follow-up validation.

      (4) The TRIC-LUC reporter system is not ideal for resolving dynamic changes in neuronal activity. Activity-dependent Ca<sup>2+</sup> signaling must first reconstitute the TRIC transcriptional system, which then drives luciferase transcription, translation, and accumulation. The original characterization of TRIC indicates that TRIC signals accumulate and decay over several hours. Thus, the kinetics of the TRIC-LUC reporter should be interpreted cautiously, particularly when inferring transient or precisely timed changes in neuronal activity.

      We fully acknowledge the inherent limitations of the TRIC-LUC reporter system, as pointed out by the reviewer. Every experimental tool comes with characteristic strengths and drawbacks. Although TRIC-LUC suffers from temporal delays, our experiment does not aim to capture acute, immediate effects; instead, it examines long-term dynamics of neuronal activity. To date, within Drosophila neurobiology, no superior technique is available for non-invasive long-term monitoring of neuronal activity in freely behaving flies. We share the hope that new tools capable of reporting neuronal activity in real time will be developed and applied, which will facilitate deeper mechanistic understanding of neuronal dynamics.

      (5) Including data from training under fed conditions would provide a more complete understanding of state-dependent neural activity and would help distinguish starvation-specific effects from more general circuit mechanisms.

      We appreciate this suggestion. First, our memory paradigm relies on reward-based associative learning, and starvation is required for flies to express robust memory, so to do this would require a completely new experimental set up. Second, our core findings demonstrate that transient perturbations of this neuronal circuit trigger sleep disturbances and memory deficits specifically under starvation conditions. Therefore, measurements of neural activity under fed conditions are not directly relevant to the central conclusions of the present study. We agree that related experiments on other behaviors under fed states constitute an interesting direction and could be pursued in future investigations.

      Reviewer #3 (Public review):

      Summary:

      Understanding the neural circuits that link sleep and memory remains a fundamental challenge in neuroscience. In this study, Lin Yan and colleagues investigate how dopamine signaling in Drosophila regulates long-term memory (LTM) formation in the context of sleep. They identify a specific microcircuit between protocerebral anterior medial dopamine neurons (PAM-DANs) and dorsal paired medial (GABAergic DPM) neurons that modulates memory consolidation. Their findings suggest that disrupting the basal activity of PAM-α1 neurons during early consolidation impairs LTM, with particularly pronounced effects under starvation conditions. Notably, sleep fragmentation caused by this disruption can be pharmacologically rescued, restoring LTM. These results provide compelling evidence how dopamine signaling plays a crucial role in linking sleep and memory, offering new insights into the underlying mechanisms.

      Strength:

      This study presents a well-executed investigation into sleep-memory interactions, utilizing a combination of connectomics, behavioral assays, functional imaging, and pharmacological manipulations. The authors convincingly demonstrate that the PAM-α1 and DPM circuit interact, highlighting a potential mechanism by which sleep influences memory consolidation. The anatomical and functional dissection of this circuit is of high interest to the field, and the study's integration of sleep and memory processes contributes significantly to our understanding of the role of dopamine in cognitive functions. Additional experiments investigating the contribution of MBON-α1 to the circuit, connectomic analysis together with a dissection of dopamine receptor function further strengthen the proposed circuit motif and its biological relevance.

      Weaknesses:

      While the study is well designed, presents compelling findings and has been further strengthened by additional experiments, some aspects remain unclear. The role of DPM neurons in memory consolidation seems not yet fully resolved, as different genetic approaches yield variable results. Furthermore, some manipulations impair memory without affecting sleep fragmentation - or vice versa, suggesting that the observed memory deficits cannot be explained solely by impaired sleep-dependent consolidation. It would also have been interesting to discuss potential mechanisms by which dopamine receptor-mediated cAMP signaling could lead to a reduction in Ca<sup>2+</sup> signals. I am confident that these questions can be addressed in future studies.

      We greatly appreciate the reviewer’s positive evaluation of our work and the thoughtful suggestions regarding future directions. As acknowledged in the manuscript, our study centers on identifying a shared circuit that coregulates sleep and memory processes. We have performed preliminary investigations of downstream circuitry, and our results indeed support the idea that sleep and memory can be modulated independently. Importantly, we have avoided drawing definitive conclusions that memory deficits arise purely from impaired sleep-dependent consolidation. Instead, we emphasize the existence of a common circuit mechanism governing both processes.

      We also thank the reviewer for drawing attention to dopamine receptor-mediated cAMP signaling and Ca<sup>2+</sup>dynamics. This observation constitutes an additional finding that requires more extensive mechanistic follow-up. Given the scope of the current work, we have not dedicated a separate discussion section to dissecting this pathway. This promising line of inquiry will be pursued in our future research.

      Conclusion:

      Overall, this study provides valuable new insights into how sleep and dopaminergic circuits interact to regulate memory consolidation in Drosophila and may reveal general principles underlying the neural regulation of memory.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      No further issue, apart from the reverse figure 13 are not found in the main text as indicated.

      We thank the reviewer for this careful check. We have performed a full-text search for Figure 13 throughout the revised manuscript and found no relevant citation. We have also carefully cross-checked all figure numbering and confirm that all figure labels are accurate in the current version.

      Reviewer #2 (Recommendations for the authors):

      In Fig. 8B, some individual GCaMP measurements show values below −100% ΔF/F₀. Under the stated definition, ΔF/F = (Fn - F0) / F0, values below −100% would require Fn to be negative. Since raw fluorescence intensity cannot be negative, values below −100% require further explanation. The authors should clarify whether Fn represents raw fluorescence or processed fluorescence, and whether the plotted traces underwent any normalization, subtraction, detrending, or transformation beyond the stated formula.

      We greatly appreciate the reviewer’s rigorous scrutiny of our data and the valuable question raised. Our fluorescence signals were calculated using the standard formula ΔF/F = (Fn − F0)/F0, where Fn = F_ROI − F_background, and all calculations were implemented accordingly. We have carefully revisited all raw imaging datasets and identified the source of the issue. During initial data processing, we retained all acquired recordings without excluding samples exhibiting focal plane drift. This drift occasionally yielded negative values for Fn (F_ROI − F_background). Beyond the five DPM cell bodies from four brains in Figure 8B highlighted by the reviewer, we further detected six additional DPM cell bodies from four brains in Figure 2A affected by the same artifact. We have now excluded these drifting preparations, regenerated all corresponding plots, and updated the statistical analyses in the revised manuscript. For transparency, we upload both the raw and processed datasets as supplementary materials to clarify this point.

    1. eLife Assessment

      This important study introduces NoSeMaze, a semi-naturalistic platform for continuous, high-dimensional tracking of social and cognitive behaviors in group-housed mice, and uses it to show that individual social rank is stable across changing social contexts. By integrating automated dominance measures, proactive social behaviors, and reinforcement-learning-based profiles, the authors demonstrate a novel framework for examining how stable individual differences shape social structure. The findings provide compelling evidence that dominance traits are stable across changing social contexts and largely independent of non-social cognitive performance, supporting the view that social rank reflects an intrinsic dimension of individuality however, the broader functional significance of dominance in this paradigm remains somewhat ambiguous, including the extent to which the measured behaviors capture how dominance operates in more naturalistic social settings. This work will be of broad relevance for behavioral neuroscience and social behavior research.

    2. Reviewer #2 (Public review):

      Summary:

      This manuscript presents the "NoSeMaze", a novel automated platform for studying social behavior and cognitive performance in group-housed male mice. The authors report that mice form robust, transitive dominance hierarchies in this environment and that individual social rank remains largely stable across multiple group compositions. They further demonstrate that social dominance and aggressive behaviors, like chasing, are partially dissociable and that dominance traits are independent of non-social cognitive performance. The study includes a genetic manipulation of oxytocin receptor expression in the anterior olfactory nucleus, which showed only transient effects on social rank.

      Strengths:

      (1) Innovative Methodology:<br /> The NoSeMaze platform is a technically elegant and conceptually well-integrated system that enables fully automated, long-term monitoring of both social and cognitive behaviors in large groups of group-housed mice. It combines tube-test-like dominance contests, voluntary chase-escape interactions, and an embedded operant olfactory discrimination task within a single, ethologically relevant environment. This modular design allows for high-throughput, minimally invasive behavioral assessment without the need for repeated handling or artificial isolation.

      (2) Experimental Scale and Rigor:<br /> The study includes 79 male mice and over 4,000 mouse-days of observation across multiple group reshufflings. The use of RFID-based identification, automated data logging, and longitudinal design enables robust quantification of individual trait stability and group-level social structure.

      (3) Multidimensional Behavioral Profiling:<br /> The integration of social (tube dominance, proactive chasing), physical (body weight), and cognitive (olfactory learning task) measures offers a rich, multi-dimensional profile of each individual mouse. The authors' finding that social dominance traits and non-social cognitive performance are largely uncorrelated reinforces emerging models of orthogonal behavioral trait axes or "animal personalities".

      (4) Clarity and Data Analysis:<br /> The analytical framework is well-suited to the study's complexity, with appropriate use of dominance metrics, mixed-effects models, and permutation tests. The analyses are clearly explained, statistically rigorous, and supported by transparent supplementary materials.

      Weaknesses:

      (1) Scope Limitations (Sex):<br /> The study is limited to male mice, which represents a common but problematic bias.

      (2) Ambiguity of Dominance as a Construct:<br /> While the study robustly quantifies social rank and hierarchy structure, the broader functional meaning of "dominance" remains unclear.

    3. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The goal of the study was to address the question of the degree to which social position in a group is a stable trait that persists across conditions. Reinwald et al. use a custom-built cage system with automated tracking and continuous testing for social dominance that does not require intervention by the experimenter. Remixing of individuals from different groups revealed that social position was rather stable and not really predictable from other measures that were taken. The authors conclude that social position is multifaceted but dependent on characteristics like personality traits.

      Strengths:

      (1) Reductionistic, highly controlled setting that allows for the control of many confounding variables.

      (2) Very interesting and important question.

      (3) Confirms the emergence of inter-individual behavior-driven differences in inbred mice in a shared environment.

      (4) Innovative paradigm and experimental setup.

      (5) Fresh perspective on an old question that makes the best use of modern technology.

      (6) Intelligent use of behavioral and cognitive covariables to generate a non-social context.

      (7) Bold and almost provocative conclusion, inviting discussion and further elaboration.

      We thank Reviewer #1 for this constructive and balanced evaluation of our work, and for highlighting both the conceptual importance of the question and the strengths of our automated, highly controlled approach.

      Weaknesses:

      (1) Reductionistic, highly controlled setting that blends out much of the complexity of social behavior in a community.

      Our goal in developing the NoSeMaze was to provide an enriched yet standardized environment that allows animals to interact freely in a complex setting where arena geometry and access contingencies are held constant across groups. As the Reviewer also notes as a strength, this highly controlled setting with minimal experimenter interference enables us to minimize confounds and provides reproducibility across rounds and groups. We now clarify this explicitly and frame the design as a trade-off between ecological complexity and experimental controllability.

      The NoSeMaze is designed as an open system that can incorporate additional configurations. In this first study with the system, we intentionally used a design that focused on single-sex groups without mating, no intruders or external threats, stable environmental conditions, and adult mice. This allowed us to establish social behavior under one defined condition. From here, future studies can add certain levels of ecological and social complexity to progressively understand their impact on specific behaviors and group dynamics.

      Accordingly, we have revised the manuscript to clarify the scope and the role of social and environmental conditions to the here observed phenomena. We also describe how future studies with systematically modified conditions can be used to understand how they change behaviors. We further replaced potentially over-broad terms (e.g., “naturalistic,” “real-world”) with more precise wording throughout.

      We modified the following sections:

      Abstract

      We removed “… within naturalistic mouse groups.” (ll. 52-54) and changed the sentence to “The approach thus enables longitudinal modeling of individuality and social position as key resilience factors.”

      We changed “… in naturalistic groups.” (l. 38) to “… in larger male mouse groups.” and “… from naturalistic tube competitions …” (ll. 41-42) to “… from incidental competitions in the integrated tube tests …”. We also deleted “naturalistic” in l. 50.

      Discussion (ll. 691-701)

      “…The NoSeMaze aims to increase environmental complexity and group dynamics while retaining experimental control, enabling longitudinal high-dimensional phenotyping of individuals. At the same time, it is a controlled laboratory group-housing habitat optimized to capture a subset of the determinants of social complexity present in natural communities. The strength of this design lies in the continuous, observer-independent observation of complex behaviors in defined environmental and social contexts. Accordingly, we interpret our findings as applying to the social contexts and environmental conditions tested here. Building on this, its modular design allows for introducing additional environmental and social factors like stressors, mating behavior, or resource competition in the future to understand their respective impact in modifying social behaviors.”

      Conclusion

      We changed “… complexity of real-world behavior …” (ll. 729-730) to “… complexity of group behavior in semi-naturalistic conditions ...”

      (2) The motivation to enter the test tube is not "trait" (or at least not solely a trait) but the basic need to reach food and water; chasing behavior would be less dependent on this stimulus.

      Tube traversals may reflect different motivations, including routine movement between compartments to access food, the water lickport, the open arena, or the housing area. Nevertheless, we do not interpret tube-entry motivation itself as a ‘trait’. Importantly, our hierarchy readout does not quantify which animals enter the tubes, nor is it confounded by tube-entry frequency itself (cf. ll. 527-532). Rather, it captures the consistent outcomes of incidental dyadic competitions once two animals meet in the tube (push vs. retreat), aggregated across many interactions. Thus, the stable signal we report lies in repeatable competition outcomes, not in traversal propensity. Consistent with this interpretation, the overall number of tube competition events was not associated with social rank, arguing against systematic competition avoidance by low-ranking animals. In contrast, chasing is a voluntarily initiated, asymmetric interaction between an initiator and a recipient and therefore adds a social-action component beyond incidental access-linked encounters, providing a complementary readout of social behavior.

      We also noted that the term ‘trait’ may not be the optimal description in this context and replaced it throughout the ms. with more precise wording, including “internalized social rank” and “propensity to chase”.

      Results

      We added the following paragraph (ll. 524-532):

      “Tube crossings in the NoSeMaze are motivated by the intent to eat, drink, sleep, or socialize. Accordingly, competitions within the tube arise by chance when two animals enter from opposite sides at the same time. Social rank therefore captures the consistent outcomes of repeated incidental competitions (push versus retreat). This measure is not confounded by differences in tube engagement, as social rank was neither associated with participation in tube competitions (ρ = 0.11, p = 0.139; Fig. 6C, Supplementary Fig. S12C) nor with the overall number of tube detections when controlling for chasing (Spearman’s partial correlation between detection count and z-scored David’s score, corrected for the fraction of active chases: ρ = 0.022, p = 0.758).”

      Discussion

      We changed the following paragraphs:

      Line 608-611

      “Importantly, participation frequency and differences in tube entry time did not confound the resulting social ranks, underscoring the robustness of the automated incidental rank assessment.”

      Lines 653-662

      “These dynamic aspects of chasing, including its asymmetric initiator–recipient structure and proactive engagement, differ from the nature of tube competitions, which are incidental encounters. Together, tube-derived social rank and chasing describe complementary dimensions of social position, and, alongside other features such as clique formation [36], contribute to describe facets of a broader multidimensional social behavior. Within this complex environment, chasing emerges as a flexible behavioral propensity that is dissociable from formal social rank. Specifically, in the NoSeMaze, chasing contributes dynamically to the maintenance, negotiation, or clarification of social hierarchy structure.”

      (3) Dominance is only one aspect of sociality, social structure is reduced to rank. The information that might lie in the chasing behavior is not optimally used to explain social behavior beyond the rank measure.

      In this manuscript, we focused on the relationship between social rank derived from incidental tube competitions and chasing. We agree that social rank derived from tube competitions captures only one aspect of social structure. We have now clarified this point more explicitly in the revised manuscript. Here, our specific aim was to determine how chasing relates to competition-derived social rank and whether it provides information beyond rank in these mouse groups.

      We also refer to another manuscript dedicated to additional aspects of social structure captured in the video data, including approach and interaction behavior as well as social clique formation. There, these measures are again examined in relation to chasing and social rank. We also plan future studies leveraging chasing in the NoSeMaze as a readout for neurophysiological investigations.

      In the present study, social rank is based on dominance and subordination in incidental competitions in the integrated tube test. We treat chasing as a distinct, volitional social dimension rather than redundant “rank information”. Importantly, our chasing analyses add beyond social rank information in three ways: (1) structural asymmetry (initiator vs recipient roles) not captured by symmetric social rank measures; (2) elite-centric reciprocal dynamics rather than broad top-down enforcement; and (3) context dependence, with social rank–chasing coupling strengthening when group transitivity is lower. We revised the Discussion/Conclusion accordingly and additionally note that other aspects of social organization (e.g., affiliative bonding and higher-order network structure such as clique/rich-club organization; Nelias et al., 2025) require complementary measures (e.g., video-derived interaction networks) and are not the scope of the present study.

      To avoid ambiguity, we also added the following operational definitions in the Introduction and Results:

      Introduction

      Lines 90-94

      “In this study, we use social hierarchy to denote the group-level structure inferred from incidental competitions in the integrated tube tests, and social rank for an individual’s level within that hierarchy. Social position serves as an umbrella term for social rank and chasing behaviors.”

      Results

      Line 306-308

      “Here, social hierarchy refers to the group-level structure reconstructed from incidental competitions in the integrated tube tests, social rank to an individual’s position within that structure, and social position to social rank together with chasing.”

      We additionally clarified the distinction between chasing and social rank in the following sections.

      Discussion (ll. 632-664)

      “In more constrained or despotic conditions, chasing can serve as a unidirectional, dominance-related behavior directed at subordinates [9,43,44]. The larger groups observed in the NoSeMaze reveal a more nuanced role for chasing behavior. Chasing levels are individually stable, but the expression and meaning of chasing are context-sensitive. Chasing was neither broadly distributed nor consistently directed down the social hierarchy. Instead, it was initiated by a small subset of individuals – primarily those occupying high competition-based social ranks – and frequently occurred reciprocally within this group, suggesting intra-elite social dynamics rather than broad dominance enforcement. Rather than solely serving to impose social hierarchy, chasing appeared to function as a means through which individuals with high social rank monitor, negotiate, and maintain their relative standing within the top tier. Notably, the identity of frequent chasers remained stable over time and persisted across changing group compositions, indicating that the propensity to initiate chases reflects a consistent individual-level tendency rather than a purely situational response.

      However, its coupling to social rank depended on the group’s hierarchy structure. In groups with less clearly defined social hierarchies (i.e., lower transitivity), active chases aligned more strongly with social rank, suggesting that mice in less structured groups rely more on proactive signaling to clarify social rank. Indeed, the top-ranked mice in these groups exhibited relatively high levels of active chases. This context-sensitivity highlights chasing’s dual role: it serves as a tool for negotiating social rank among mice at the upper end of the hierarchy, and additionally functions to establish or reinforce hierarchical clarity when social structures are ambiguous. These dynamic aspects of chasing, including its asymmetric initiator–recipient structure and proactive engagement, differ from the nature of tube competitions, which are incidental encounters. Together, tube-derived social rank and chasing describe complementary dimensions of social position, and, alongside other features such as clique formation [36], contribute to describe facets of a broader multidimensional social behavior. Within this complex environment, chasing emerges as a flexible behavioral propensity that is dissociable from formal social rank. Specifically, in the NoSeMaze, chasing contributes dynamically to the maintenance, negotiation, or clarification of social hierarchy structure. These findings reveal a novel aspect of social dynamics: chasing is not merely a dominance display but a flexible context-dependent mechanism shaped by both individual disposition and group-level social structure.”

      Conclusion (ll. 718-723)

      “Crucially, social position is not fully described by a single behavioral dimension. We focused here on two separable dimensions: competition-based social rank and proactive chasing. Future work should integrate additional dimensions of social organization, such as affiliative bonding and higher-order network measures [36], to capture additional aspects of social organization. Chasing behavior played a dual role, reflecting a stable individual propensity while also adapting to group-level structure.”

      (4) Focus on rank bears the risk of overgeneralization for readers not familiar with the context.

      As already discussed above (see also point 2 and 3), we have sharpened the framing to reduce the risk of overgeneralization. Throughout the manuscript, we now refer to tube-derived social rank explicitly as a dominance-subordination-related axis of social organization rather than a comprehensive measure of “sociality”, and we also avoid the term “traits”, but prefer the use of “internalized social rank” or “propensity to chase” when discussing stability across rounds and contexts. Also, we removed “personality-like” as the scope of the study is to set specific behaviors in relation and not to enter the field of mouse personality classification.

      We clarified the definition of social rank in this manuscript in the Introduction (ll. 90-94) and the beginning of the Results (ll. 306-308, see also point 3).

      Abstract

      Lines 41-46

      “… Across more than 4,000 mouse-days, hierarchies derived from incidental competitions in the integrated tube tests were non-despotic, transitive, and stable even when group compositions changed. This stability supports an internalized component of competition-based social rank. Chasing was also stable across contexts. Notably, chasing was concentrated among high-ranking individuals, consistent with ongoing negotiation of social rank among individuals at the upper end of the hierarchy.”

      Lines 50-54

      “In summary, high-dimensional tracking with the NoSeMaze reveals that social position in mice is multifaceted and shaped by stable dimensions of individual behavior that persist across changing social contexts. The approach thus enables longitudinal modeling of individuality and social position as key resilience factors.”

      Discussion (ll. 622-626)

      “This temporal and contextual stability supports interpreting social rank as a stable, internalized characteristic of individuals. In this sense, ‘internalized’ refers to stability across repeated rounds and reshuffled groups in this paradigm and to relatively stable tube-competition outcomes.”

      Changes in the Conclusion starting l. 709 as highlighted in answer to the above point 3.

      (5) Conclusion only valid for the reductionistic setting, in which environment, social and non-social changes only within narrow limits, and in which the mouse population does not face challenges

      Our conclusions are bounded by the conditions tested, namely an enriched but stable environment, controlled group composition and remixing, and the absence of explicit ecological challenges such as resource scarcity or predators. As described above in the answer to point 1, we now state explicitly that our conclusions apply to the social-context variation tested here (controlled group reshuffling) under otherwise stable conditions, and we frame resource-competition and environmental-stressor manipulations as future variation to test of how this manipulation affects the here described behaviors.

      We accordingly changed the Discussion as already highlighted in point 1 above and added ll. 693-704.

      (6) Animals are not naive at the beginning of the experiment, but are already several weeks old.

      Animals entered the study as adults and were continuously group-housed (3-5 mice/cage) before entering the NoSeMaze. To reduce effects of initial apparatus novelty, mice underwent two NoSeMaze habituation sessions prior to data collection (each several hours). We now clarify these points explicitly in the Methods and scope the inference accordingly. One of our future steps is to extend this framework to earlier developmental stages (e.g., adolescence/weaning) to capture full lifespan trajectories. This however first required establishing the approach in adult mice under controlled conditions as done in the present study.

      Methods (ll. 739-747)

      “A total of 79 adult male homozygous OXTRfl/fl mice (B6.129(SJL)-Oxtrtm1.1Wsy/J, RRID: IMSR_JAX:008471, Jackson Laboratory) backcrossed > F10 to C57BL/6J background (Charles River, Sulzfeld) were used for the experiments. Animals entered the NoSeMaze as adults (see Supplementary Table S2 for ages at NoSeMaze entry across rounds) and had been continuously group-housed after weaning (3–5 mice/cage), i.e., they were not developmentally or socially naïve at study onset. Of the 79 mice, 26 mice were injected six weeks before the start of the experiment with an AAV expressing Cre recombinase (rAAV1/2-CBA-Cre) into the AON pars centralis to induce bilateral OXTR deletion (OXTRΔAON). The remaining 53 animals received an AAV expressing only dTomato (rAAV1/2-CBA-dTomato). …”

      Discussion (ll. 629-631)

      “… A future direction is to extend this framework to earlier developmental stages to understand which early experiences shape later trajectories of social position.”

      In summary, this is a wonderful study, but not one that is easy to interpret. The bold conclusion is valid only within the constraints of the study, but nevertheless points in an important direction. The paradigm is clever and could be used for many interesting follow-ups.

      To define social position as a personality trait will elicit strong opposition and much debate; the nuances of the paper might be lost on many readers and call for the (re)-consideration of many concepts that are touched. I find this attitude a strength of the paper, but the approach bears the risk of misunderstanding.

      We thank the Reviewer for the helpful comments. As detailed above, we tightened terminology and framing throughout the revised manuscript by defining stability explicitly in terms of repeatability across rounds and reshuffled groups within this paradigm, and described tube-derived social rank as one dimension of social behavior. We hope these edits preserve the conceptual message while reducing the risk of misunderstanding.

      Reviewer #2 (Public review):

      Summary:

      This manuscript presents the "NoSeMaze", a novel automated platform for studying social behavior and cognitive performance in group-housed male mice. The authors report that mice form robust, transitive dominance hierarchies in this environment and that individual social rank remains largely stable across multiple group compositions. They further demonstrate that social dominance and aggressive behaviors, like chasing, are partially dissociable and that dominance traits are independent of non-social cognitive performance. The study includes a genetic manipulation of oxytocin receptor expression in the anterior olfactory nucleus, which showed only transient effects on social rank.

      Strengths:

      (1) Innovative Methodology:

      The NoSeMaze platform is a technically elegant and conceptually well-integrated system that enables fully automated, long-term monitoring of both social and cognitive behaviors in large groups of group-housed mice. It combines tube-test-like dominance contests, voluntary chase-escape interactions, and an embedded operant olfactory discrimination task within a single, ethologically relevant environment. This modular design allows for high-throughput, minimally invasive behavioral assessment without the need for repeated handling or artificial isolation.

      (2) Experimental Scale and Rigor:

      The study includes 79 male mice and over 4,000 mouse-days of observation across multiple group reshufflings. The use of RFID-based identification, automated data logging, and longitudinal design enables robust quantification of individual trait stability and group-level social structure.

      (3) Multidimensional Behavioral Profiling:

      The integration of social (tube dominance, proactive chasing), physical (body weight), and cognitive (olfactory learning task) measures offers a rich, multi-dimensional profile of each individual mouse. The authors' finding that social dominance traits and non-social cognitive performance are largely uncorrelated reinforces emerging models of orthogonal behavioral trait axes or "animal personalities".

      (4) Clarity and Data Analysis:

      The analytical framework is well-suited to the study's complexity, with appropriate use of dominance metrics, mixed-effects models, and permutation tests. The analyses are clearly explained, statistically rigorous, and supported by transparent supplementary materials.

      We thank the Reviewer for their evaluation of our manuscript. We appreciate their recognition of the NoSeMaze as an innovative and conceptually integrated platform, the scale and longitudinal rigor of the dataset, and the clarity and appropriateness of the analytical framework.

      Weaknesses:

      (1) Conceptual Novelty and Prior Work:

      While the study is carefully executed and methodologically innovative, several of its core findings reaffirm concepts already established in the literature. The emergence of stable, transitive social hierarchies, the persistence of individual differences in social behavior, and the presence of non-despotic social structures have all been previously reported in mice, including under semi-naturalistic conditions (e.g., Fan et al., 2019; Forkosh et al., 2019). Although this work extends those findings with greater behavioral resolution and scale, the manuscript would benefit from a clearer articulation of what is genuinely novel at the conceptual level, beyond the technological advance.

      We agree with the Reviewer that transitive dominance hierarchies and stable inter-individual differences have been demonstrated previously in mice, including in semi-naturalistic settings. Our intent was therefore to address two more specific conceptual questions that prior work typically has not tested in an integrated way:

      (1) Whether an individual’s social position generalizes across distinct social contexts created by systematic changes in group composition rather than reflecting stability that is only observable within a fixed group;

      (2) How the two frequently employed social dominance metrics tube competition-based social rank and chasing relate to each other in unperturbed larger groups.

      Specifically, the conceptual advance is enabled by (A) continuous, handling-free estimation of competition-based social rank from incidental tube contests over weeks, (B) a repeated group-reshuffling (“accelerated longitudinal”) design that explicitly tests whether an individual’s social rank generalizes across distinct social contexts, and (C) parallel, continuous measurement of chasing and non-social reinforcement-learning behaviors in the same individuals and environment.

      This integration also lets us dissociate dominance-related social dimensions (competition-based social rank vs chasing) and test their relation to individual styles in non-social reinforcement-learning. Importantly, these questions are addressed in larger intact societies (9–10 mice), where hierarchy structure is shaped by more complex network-level dynamics than in smaller groups.

      We revised the Introduction and Discussion to foreground these conceptual points and to position them more explicitly in the context of prior works.

      Introduction

      We adapted the following section (ll. 76-98) and integrated the suggested references:

      “… In mice, small groups tend to form highly despotic hierarchies [24-26], whereas larger groups exhibit more complex structures [27]. These patterns suggest a strong influence of emergent group-level dynamics on social structure [12,27]. Yet, animals do not enter social groups as blank slates [28-30]; stable latent factors in the individual may also contribute to hierarchy formation [13]. Thus, it remains unclear to what extent an individual’s social position is internalized and persists across different social contexts [31], or instead is primarily an emergent property of group-level dynamics. Disentangling these possibilities requires experimental conditions that allow unperturbed, continuous tracking of all individuals in sufficiently large groups, together with systematic changes of group composition to modulate social context.

      While stable hierarchies and behavioral identity domains have been described previously in semi-naturalistic settings [9-13], many studies quantify these features either within fixed group compositions or in separate assays. Here, we therefore use systematic group reshuffling in 10-member societies to directly test whether individual differences in social rank and chasing persist across distinct social groups. In this study, we use social hierarchy to denote the group-level structure inferred from incidental competitions in the integrated tube tests, and social rank for an individual’s level within that hierarchy. Social position serves as an umbrella term for social rank and chasing behaviors. By continuously measuring social rank, chasing, and reinforcement-learning behavior in the same individuals, we further test how chasing and social rank are related to each other and how they relate to individual styles in non-social reinforcement learning.

      To enable these tests, we developed the Non-invasive Sensor-rich Maze (NoSeMaze).”

      Discussion

      We added a brief framing statement at the beginning of the Discussion (ll. 577-583) to clarify the study’s conceptual novelty:

      “Ecologically enriched, yet experimentally controlled assessments allow us to study behavioral individuality and social structure in group-living animals over extended timescales. Here, we show that individual mice carry stable, individual-specific, and multi-faceted profiles of social position and cognitive styles across changing group contexts. Our approach goes beyond prior work by testing cross-context stability under repeated, systematic group reshuffling in larger mouse societies, while measuring competition outcomes, chasing, and reinforcement-learning behavior in parallel. …”

      (2) Role of OXTR Deletion:

      The inclusion of the OXTR manipulation feels somewhat disconnected from the manuscript's central aims. The effects were minimal and transient, and the authors defer full interpretation to a separate study.

      We appreciate the Reviewer’s point and agree that the OXTR<sup>ΔAON</sup> manipulation can appear secondary to the manuscript’s central aims. We included OXTR<sup>ΔAON</sup> because oxytocin-dependent social recognition memory was hypothesized originally to impact potentially also learning an individual’s position in social hierarchy networks. Even though the effects were small and transient, we nevertheless believe it is relevant to report them, also in relation to the more profound effects of the genetic manipulation reported in a related manuscript (Nelias et al., bioRxiv 2025, 10.1101/2025.08.26.672298). Therefore, we explicitly account for genotype as a covariate in the analyses such that the main conclusions do not depend on this manipulation.

      To improve coherence, we have added one sentence in the Introduction motivating why OXTR<sup>ΔAON</sup> was included, and finally more clearly signposted that deeper mechanistic interpretation is beyond the scope of the present manuscript.

      Introduction (ll. 115-120)

      We added one short sentence introducing the rationale behind the perturbation:

      “… and (3) determine whether social rank, chasing, and non-social reward-seeking behaviors represent stable individual characteristics or dynamic features across time and changing group composition. As a secondary analysis, motivated by oxytocin’s established role in social recognition memory [32,33], we also tested whether OXTR deletion in the anterior olfactory nucleus produces detectable shifts in rank dynamics.”

      Results

      Lines 161-164

      “… This manipulation was included as a secondary biological perturbation. The primary analyses and conclusions focus on the platform and cross-context stability, and genotype is treated as a covariate unless stated otherwise. …”

      Lines 451-454

      “… As a secondary analysis, we tested whether OXTR<sup>ΔAON</sup>, which impairs de novo social recognition memory required for social clique formation in this cohort [36], also affects the measures reported here. Consistent with largely internalized features, OXTR<sup>ΔAON</sup> produced only transient effects. …”

      Discussion (ll. 594-605)

      “This study focused primarily on the relation of social rank and chasing. We however also considered their relation to additional variables including the loss of oxytocin receptors in the olfactory cortex in the adult (OXTR<sup>ΔAON</sup>), involved in de novo social recognition learning. The propensity to chase was largely unaffected by OXTR<sup>ΔAON</sup>. Mice carrying OXTR<sup>ΔAON</sup> displayed a transient reduction in social rank during the first week that normalized thereafter. This transient effect contrasts to the persistent impairment by OXTR<sup>ΔAON</sup> in forming higher-order social bonds that enable membership in stable cliques, as identified by video tracking of self-paced interactions in the same cohort [36]. Together, these findings suggest that OXT-dependent olfactory learning is critical for the formation of social context-dependent higher-order bonds, but plays a limited role in shaping hierarchy-related behaviors.”

      (3) Scope Limitations (Sex and Age):

      The study is limited to male mice, and although this is acknowledged, the title and overall framing imply broader generalizability. This sex-specific focus represents a common but problematic bias. Additionally, results from the older mouse cohort are under-discussed; if age had no effect, this should be explicitly stated.

      We thank the Reviewer for this relevant point. The study is limited to male mice. We therefore revised the title and the abstract and strengthened the limitations to make the sex-specific scope explicit.

      We additionally note ongoing work extending the same framework to female groups, where we find similar hierarchy structure and cross-context stability in the NoSeMaze in preliminary unpublished data.

      Regarding age, our design included two separate cohorts of different adult age ranges (young and older adult animals), but age was not the primary experimental factor. As detailed in our response to Reviewer #3 (point 1), we now quantify age structure explicitly and test age effects using a decomposition that separates between-group age differences from within-group age variation (mean age per group and each animal’s deviation from that mean). We also recomputed stability estimates with age-adjusted ICC models, and the resulting ICCs are highly similar to the original estimates (cf. new Supplementary Table S4), indicating that the reported metrics’ stability is not driven by age differences across groups.

      Title

      We change the title from “Individual differences drive social hierarchies in mouse societies” to “Individual differences drive social hierarchies in male mouse societies”

      Abstract

      We also added male in the abstract (ll. 37-38):

      “The interaction of these behaviors in the shaping of social position in larger male mouse groups remains largely unknown.”

      Discussion (ll. 613-617)

      “… While this study focused on male mice, in which social hierarchies are best established [9], future work is needed to explore sex-specific expressions of social structure and their neurobiological underpinnings in female groups. We therefore restrict our interpretation to male mice. Critically, male social ranks were robustly maintained within the same group over time. …”

      For a more detailed integration of age in the Methods, Results, and Discussion, we kindly refer to the reply to Reviewer #3, point 1.

      (4) Ambiguity of Dominance as a Construct:

      While the study robustly quantifies social rank and hierarchy structure, the broader functional meaning of "dominance" remains unclear. As in prior work (e.g., Varholick et al., 2019), dominance rank here shows only weak associations with physical attributes (e.g., body weight), cognitive strategy, or neuromodulatory manipulation (OXTR deletion). This recurring pattern, where rank metrics are reliably established yet poorly predictive of other behavioral or biological traits, raises important questions about what such measures actually capture. In particular, it challenges the assumption that outcomes in paradigms like the tube test or chase frequency necessarily reflect dominance per se, rather than other constructs.

      We thank the Reviewer for this clarifying point. We agree that stable social rank metrics do not necessarily imply a complete or unitary measure of “dominance.” In the revised manuscript, we therefore clarified that, in this study, social rank is operationalized as consistent competitive outcomes in incidental tube-test encounters in the NoSeMaze (see also Reviewer #1, points 2-4).

      Our data indicate that this competition-based social rank is related to, but not identical with, other social behaviors such as chasing. Likewise, body weight significantly contributes to social rank, but explains only part of the variance. We therefore do not interpret weak or partial associations with other variables as invalidating the social rank measure. Rather, we interpret them as indicating that social position is multidimensional and only partially captured by any single assay.

      In independent subsequent studies that are currently in preparation or revision, we observed that heterogeneity in the neurobiology and response to challenges was best predicted by the competition-based social rank, also compared to the other behaviors assessed here. While these observations are beyond the scope of the present manuscript, they support our view that incidental tube competition and the resulting dominance-subordination structure may provide a biologically informative measure of one important dimension of social position. At the same time, the observation here and in many previous studies that factors such as body weight explain only a limited portion of the variance remains important for our understanding of these constructs.

      We have revised the manuscript accordingly to make this distinction more explicit. In particular, we now define more precisely the terms social hierarchy, social rank, and social position in the context of this manuscript (see also reply to Reviewer #1, point 3; ll. 90-94 in the Introduction and ll. 306-308 in the Results), and we use these terms more consistently throughout. We also revised the text to avoid overstating the meaning of “dominance” where the data support a more specific interpretation.

      Finally, we now emphasize more clearly both the strength and the limitation of the present approach. The NoSeMaze allows these relationships to be assessed continuously in a minimally perturbed group-housing ecology, laying ground for future incorporation of further variables and dimensions to capture how they shape social organization.

      Specifically, we added text in the Discussion and Conclusion to clarify that competition-based social rank and proactive chasing represent separable dimensions related to dominance and subordination, but do not exhaust sociality or individuality, and that future work should integrate additional measures such as affiliative behavior and higher-order network structure.

      Discussion

      Lines 624-6631:

      “… In this sense, ‘internalized’ refers to stability across repeated rounds and reshuffled groups in this paradigm and to relatively stable tube-competition outcomes. The tube-derived social rank describes here a dimension of individual social behavior. The capacity to quantify stable individual differences across changing social contexts highlights the value of the NoSeMaze for lifespan-oriented studies of behavioral individuality. A future direction is to extend this framework to earlier developmental stages to understand which early experiences shape later trajectories of social position.”

      Conclusion

      Lines 718-722:

      “Crucially, social position is not fully described by a single behavioral dimension. We focused here on two separable dimensions: competition-based social rank and proactive chasing. Future work should integrate additional dimensions of social organization, such as affiliative bonding and higher-order network measures [36], to capture additional aspects of social organization.”

      Reviewer #3 (Public review):

      Reinwald et al. present the NoSeMaze, a semi-natural behavioral system designed to track social behaviors alongside reinforcement-learning in large groups of mice. Accumulating more than 4,000 days of behavioral monitoring, the authors demonstrate that social rank (determined by tube competitions) is a stable trait across shuffled cohorts and correlated with active chasing behaviors. The system also provides a solid platform for long-term measurements of reinforcement learning, including flexibility, response adaptation, and impulsiveness. Yet, the authors show that social ranking and chasing are mostly independent of these cognitive traits, and both seem mostly independent of oxytocin signaling in the AON.

      Strengths:

      (1) The neuroethological approach for automated tracking of several mice under semi-natural conditions is still rare in social behavioral research and should be encouraged.

      (2) The assessment of dominance by two independent measures, i.e., spontaneous tube competitions and proactive chasing, is innovative and valuable.

      (3) The integration of a long-term reinforcement-learning module into the semi-natural system provides novel opportunities to combine cognitive traits into social personality assessments.

      (4) The open-source system provides a valuable resource for the scientific community.

      Limitations:

      (1) Apparent ambiguity and inconsistency in age structure and cohort participation across rounds, raising concerns about uncontrolled confounds.

      (2) Chasing behavior appears more stable than tube-test competitions (Figure 4D vs. Figure 3D), which challenges the authors' decision to treat tube competitions as the primary basis for hierarchy determination.

      We thank the Reviewer for the evaluation of our work. We have addressed the limitations raised below with additional analyses, clarifications, and corresponding manuscript revisions.

      Major concerns:

      (1) Unclear and inconsistent handling of age groups and repeated sampling. The manuscript repeatedly refers to "younger" and "older" adults, but it is unclear whether age was ever controlled for or included in models. Some mice completed only one round, others 2-5 rounds, without explanation of the criteria or balancing.

      We thank the Reviewer for this clarifying comment. We now explicitly quantify participation in the different rounds in new Supplementary Table S3. Importantly, our primary stability analyses are implemented using variance-component mixed models (REML) that naturally handle unbalanced repeated-measures data. This is explained in more detail in the Methods section (ll. 1003-1111), where we added a statement that the LMEs are well suited for unbalanced repetitions.

      To further address unbalanced participation, we additionally performed a conservative sensitivity analysis restricted to a balanced subset, including only sessions 1 and 2 and only mice with observations in both sessions. For the key social measures, ICC estimates were highly similar in the full dataset and in the balanced subset (e.g., z-scored competition David’s score, ICC across cohorts: 0.55 without age adjustment vs. 0.56 in the balanced first-two-session subset; active chasing: 0.74 vs. 0.72; being chased: 0.61 vs. 0.60; Supplementary Table S4), indicating that unbalanced participation did not inflate the stability estimates.

      We also addressed age structure explicitly. Although age was not the primary experimental factor, the inclusion of two age cohorts allowed us to assess whether the observed behaviors and their interrelations were robust across most of the adult lifespan (cf. new Figure 1). We therefore decomposed age into a between-group component (mean age per group) and a within-group component (each animal’s deviation from its group mean) and included these terms in the relevant LME models. Tube-based dominance rank (David’s score, z-scored) showed no age effect (p<sub>age, cond.</sub> = 0.63), whereas chasing metrics showed modest age associations (cf. new Supplementary Table S5). This however only indicates that chasing was associated with age to some degree. More importantly, recomputing all stability estimates using age-adjusted ICC models yielded nearly identical ICCs for the core social measures (new Supplementary Table S4), indicating that the reported stability was not affected by age differences across groups.

      In addition, we revised the study design schematic (new Fig. 1; cf. Reviewer #1, Recommendations for the authors) to depict the separate age cohorts and the reshuffling procedure more clearly.

      We adapted the following sections accordingly.

      Methods

      Lines 980-985

      “… Most mice (n = 68) participated in at least two NoSeMaze rounds with reshuffled group members. Supplementary Table S2 summarizes the number of rounds per mouse and missing data due to technical problems. We examined the stability of social and reward-seeking metrics by correlating values from the first and second round (Spearman’s correlation, cf. Fig. 5). These round-1-to-round-2 correlations use one paired observation per mouse and are therefore not inflated by mice contributing >2 rounds.”

      Lines 991-993

      “The ICC treated mouse identity as a random intercept and NoSeMaze group (i.e., round-specific social group) as a random effect, with repetition included as fixed effect (i.e., stability across groups while holding repetition means constant).”

      Lines 1003-1011

      “Variance components for ICC estimation were obtained from LMEs fit by restricted maximum likelihood (REML), which yields less biased variance-component estimates and is well suited for unbalanced repeated-measures designs (i.e., different numbers of rounds per mouse). To assess potential confounding by age structure, age was decomposed into a between-group component (group-mean age) and a within-group component (each animal’s deviation from its group mean) and included as covariates. ICCs were recomputed in age-adjusted models (see Supplementary Table S4). Finally, we performed a conservative sensitivity analysis restricted to the first two sessions per mouse and to mice with observations in both sessions (“balanced first2”), to additionally account for unbalanced participation structure.”

      Results

      Lines 167-173

      “… The study population comprised two age cohorts: younger (16-30 weeks) and older adults (55-97 weeks) (Fig. 1B). The two age cohorts were run as separate experimental series, and group reshuffling was performed within each cohort (Fig. 1C, see Supplementary Table S2). Mice lived in groups of 9-10 for multiple rounds in the NoSeMaze, with different group members in each round (Fig. 1D, see Supplementary Table S1). This allowed us to test which individual behaviors were stable across different group compositions. …”

      Lines 441-450

      “Because the number of rounds in the NoSeMaze was unbalanced between animals (Supplementary Table S3) and age varied across groups (Supplementary Table S1-2), we also performed robustness checks. Stability estimates changed only minimally when recomputed in age-adjusted ICC models (between-group mean age and within-group age deviation; see Methods) and when restricting analyses to a balanced first-two-session subset (sessions 1–2 only; mice with both sessions) (Supplementary Table S4). Mixed models indicated that some chasing and reinforcement-learning measures showed modest age- and/or session-related shifts in absolute levels (Supplementary Table S5), but importantly, these did not affect the observed stability patterns.”

      (2) Stability of chasing appears stronger than the stability of tube competitions. Figure 4D shows highly consistent chasing behavior across weeks, while Figure 3D shows weaker and more variable correlations for tube-based David scores. This is also evident from Figure 5A-B,D. Thus, it appears that chasing, which serves to quantify dominance in similar semi-natural setups, may be a more reliable and behaviorally meaningful measure of dominance than the incidental tube competitions.

      Indeed, active chasing showed higher cross-round correlations than tube-derived social rank (e.g., R1–R2 Spearman ρ = 0.75 vs 0.57; ICC<sub>across cohort</sub> 0.74 vs 0.55, see Fig. 5 and new Supplementary Table S4). Importantly, this does not contradict the central finding that tube-derived social rank is stable across time and across remixed groups. Rather, it may highlight that chasing and social rank capture different aspects of dominance-subordination-related behavior with different statistical properties. Chasing reflects an individual’s propensity to actively initiate interactions (a strongly expressed, asymmetric behavior), which can be highly consistent across contexts. By contrast, social rank is a relational measure inferred from symmetric dyadic win–loss outcomes based on incidental competitions in the integrated tube tests and can vary with the specific set of competitors and interaction opportunities in each reshuffled NoSeMaze group, while still remaining substantially stable overall.

      Accordingly, we do not interpret the higher repeatability of chasing as evidence that it is the “better” hierarchy measure. Tube competitions yield symmetric dyadic outcomes that directly support formal hierarchy reconstruction (David’s score/Elo, transitivity, steepness) and show convergent validity with traditional tube testing. Chasing, in contrast, is asymmetric and volitional and, in our data, is concentrated in the upper social ranks and modulated by group-level social hierarchy structure, consistent with rank negotiation/signaling rather than a mechanism that assigns a full ordering to all individuals. These different properties lead to very different “win-lose” relationships when comparing tube competition events to chasing events that we specifically illustrated in Supplementary Fig. S13. We therefore revised the Results and Discussion to frame tube-derived social rank and chasing as complementary social dimensions.

      Results (ll. 402-413)

      “… Specifically, stability was high for the tube competition-based David’s score (Fig. 5A, ρ = 0.57, p < 0.001), as well as for the fraction of active chases (Fig. 5B, ρ = 0.75, p < 0.001) and of times being chased (Fig. 5C, ρ = 0.54, p < 0.001). Notably, the fraction of active chases showed slightly higher across-round stability than tube-derived David’s score and the fraction of being chased. This is in line with active chases capturing an individual propensity to initiate this behavior, whereas competition-based social rank is a relational measure that is also influenced by the set of competitors and interaction opportunities in each reshuffled group. Nonetheless also social rank and being chased were overall stable. Across the full series, ICCs for these social measures were also in the good–excellent range, indicating high across-round stability (Fig. 5D, ICC = 0.548 to 0.890).”

      Discussion

      Line 653-664 (see also Reviewer #1, point 2)

      “These dynamic aspects of chasing, including its asymmetric initiator–recipient structure and proactive engagement, differ from the nature of tube competitions, which are incidental encounters. Together, tube-derived social rank and chasing describe complementary dimensions of social position, and, alongside other features such as clique formation [36], contribute to describe facets of a broader multidimensional social behavior. Within this complex environment, chasing emerges as a flexible behavioral propensity that is dissociable from formal social rank. Specifically, in the NoSeMaze, chasing contributes dynamically to the maintenance, negotiation, or clarification of social hierarchy structure.”

      (3) Unbalanced participation across rounds compromises stability analyses. Stability analyses (e.g., ICCs, round-to-round correlations) assume comparable sampling across individuals. However, some mice contribute 1 round, others 2, 3, 4, and even 5 rounds. This imbalance may inflate stability estimates or confound group reshuffling effects, and the rationale for variable participation is not explained.

      We thank the Reviewer for raising this point. Indeed, participation was unbalanced across rounds (see the same Reviewer #3, Major concerns 1 and new Supplementary Table S3), mainly due to missing data from occasional technical failures of the RFID detectors or the reinforcement learning water port during acquisition in some groups (for details, see Supplementary Table S2). We added this rationale behind variable participation to our Methods section (for details, see Major concern 1, ll. 981-982, “Supplementary Table S2 summarizes the number of rounds per mouse and missing data due to technical problems.”)

      We now quantify round participation (new Supplementary Table S3) and directly address potential bias from unequal sampling in two ways. First, the round-1-to-round-2 (R1–R2) stability correlations use one paired observation per mouse and therefore are not inflated by mice contributing more than two rounds. Second, ICCs were estimated from REML variance-component mixed models that account for unbalanced repeated-measures by design and use all available observations. To further rule out inflation from unequal sampling or non-random missingness, we additionally report a conservative sensitivity ICC restricted to a balanced subset including only each mouse’s first two observed sessions and only mice with both sessions (“first2-balanced”). Full-sample ICCs and first2-balanced ICCs were highly similar (new Supplementary Table S4), indicating that participation imbalance did not affect the stability estimates. For details on the changes made in the manuscript, see Reviewer #3, Major concern 1.

      Recommendations for the authors: 

      Editor's notes:

      Should you choose to revise your manuscript, if you have not already done so, please include full statistical reporting including exact p-values wherever possible alongside the summary statistics (test statistic and df) and, where appropriate, 95% confidence intervals. These should be reported for all key questions and not only when the p-value is less than 0.05 in the main manuscript.

      Readers would also benefit from noting that the mice were male in the abstract.

      We have revised the manuscript accordingly and now report exact p-values values wherever possible for the key results in the main text. Because full statistical reporting for every analysis in the main text would substantially interrupt readability, we provide the complete statistical details in a Supplementary Excel File (Supplementary Material – Systematic Statistical Reporting), including sample sizes, degrees of freedom, the number and type of permutation tests, exact p-values, and 95% confidence intervals. We also provide Extended Data Sheets for all linear-mixed effects models that account for covariates such as age and round of participation in the NoSeMaze (cf. Reviewer #3, Major Concern (1) for the additional analyses). To guide readers to these resources, we now explicitly refer to the Supplementary Material – Systematic Statistical Reporting at several points in the manuscript:

      Results

      Lines 271-274

      “Detailed statistical reporting for all analyses, including n, degrees of freedom, exact p-values, and 95% confidence intervals, is provided in the Supplementary Material - Systematic Statistical Reporting. Extended Data includes additional linear mixed-effects models controlling for potential confounding variables.”

      Lines 430-431

      “All statistical details are provided in the Supplementary Material – Systematic Statistical Reporting.”

      Additional references to the supplementary statistical reporting were inserted at lines 252-254, 562-563, and 567-568.

      Methods

      Lines 962-965

      “Full details on the statistical tests, including n, degrees of freedom, exact p-values, and 95% confidence intervals, are provided in the Supplementary Material – Systematic Statistical Reporting, as well as in the Extended Data for the LMEs accounting for different covariates.”

      We also revised the abstract and the title to explicitly state that the mice were male.

      Manuscript changes:

      Results, figure legends, and supplementary tables expanded to include full statistical reporting; abstract revised to specify male mice.

      Reviewer #1 (Recommendations for the authors):

      I would recommend being much more explicit about the reductionistic nature of the study and how the limitations are turned here to an advantage, while at the same time acknowledging the challenges of extrapolating beyond these boundaries. Most importantly, dominance should be positioned more clearly and cautiously within a framework of social behavior (and social structure) in general. The authors include cognitive tests, etc., to generate context, but this context is dependent on the same circumstances that possibly contribute to the social structure. The information that lies in the chasing behavior as an additional measured variable might be used better to provide more context.

      We thank the Reviewer for these recommendations. We have better clarified the framing to make explicit that the NoSeMaze is a controlled laboratory group-housing system. The specific conditions are now discussed in more detail in relation the observed social behaviors. We describe the tube-derived social rank as one dimension of social organization, and proactive chasing as another one. This study presents a first necessary step to understand their shared and distinguishing features. We also elaborated the Discussion to better clarify the trade-off between ecological complexity and experimental control, and to highlight that the value of the NoSeMaze for continuous, observer-independent phenotyping under standardized conditions. We now clearly state the importance to vary conditions in the system to see how the social behaviors and also their relation to non-social features changes depending on context conditions. For detailed changes, see our responses to Reviewer #1, points (1), (3), (4), and (5), and Reviewer #2, Weakness (4).

      Manuscript changes:

      Abstract, Introduction, Discussion, and Conclusion revised to clarify scope, construct interpretation, and the complementary roles of competition-based social rank and chasing.

      The precision of the description of the experimental design should be improved. When were the cohorts mixed, or did they stay separate? Which animals were old, which were young? This remained a bit confusing.

      We revised the presentation of the study design accordingly. Specifically, we clarified that the younger and older adult cohorts were run as separate experimental series and that group reshuffling occurred within, but not across, these cohorts. We also revised the study schematic in Fig. 1 and the corresponding description in the manuscript to more explicitly depict cohort structure, timing of NoSeMaze rounds, and between-round reshuffling (details are provided in the Supplementary Tables 1-3). For new age-related analyses and additional robustness checks, see our response to Reviewer #3, Major Concern (1).

      Manuscript changes:

      Figure 1 and its legend were revised to depict cohort structure, NoSeMaze rounds, and between-round reshuffling more explicitly. In addition, the Results subsection “Ecological longitudinal assessment in the NoSeMaze” were updated to clarify that the younger and older adult cohorts were run separately, that reshuffling occurred within but not across cohorts, and which animals belonged to each age-defined cohort. We also now cross-reference Supplementary Table S2 for round-specific age information and cohort composition.

      Results

      Lines 167-170

      “The study population comprised two age cohorts: younger (16-30 weeks) and older adults (55-97 weeks) (Fig. 1B). The two age cohorts were run as separate experimental series, and group reshuffling was performed within each cohort (see Supplementary Table S2).”

      It is also not fully clear across which groups stability measures were obtained: are these across all groups or within the subgroups with a given characteristic?

      We thank the Reviewer for highlighting this point. We now state explicitly that stability measures were computed across all eligible rounds, using mixed-effects models that account for repeated observations of the same mouse and for round-specific group membership. We also added a conservative sensitivity analysis restricted to a balanced first-two-session subset. Full and balanced-subsample stability estimates were highly similar, indicating that the main conclusions are robust to the participation structure. For details, see our response to Reviewer #3, Major Concerns (1) and (3).

      Manuscript changes:

      Methods and Results revised to clarify the level of analysis and model structure, as well as new supplementary robustness tables (Supplementary Tables 3-5) added.

      Minor point: In Figure 3, 18 groups are mentioned, but 19 are shown.

      We thank the Reviewer for this clarification. The apparent discrepancy arose because the group labels in Figure 3 follow the numbering of all experimental groups, whereas only groups with available tube-competition data are shown in this panel. Accordingly, group 16 is absent because no tube data were available for that group (see Supplementary Table S2), and groups 20 and 21 are likewise not included for the same reason. We have revised the figure legend and corresponding text to make this explicit and to avoid the impression of a numbering inconsistency.

      “Fig. 3: Social rank derived from incidental competitions in the integrated tube tests of the NoSeMaze.”

      “C, Box plots of metrics characterizing social hierarchy for 18 groups, including transitivity, steepness, stability, and uncertainty-by-repeatability (for details, see ‘Source Data’). Group labels correspond to original experimental group IDs. Only groups with available tube-competition data are shown. Therefore, numbering is non-consecutive (e.g., groups 16, 20, and 21 are absent; see Supplementary Table S2).”

      Reviewer #2 (Recommendations for the authors):

      (1) To better distinguish this study from previous literature, we recommend incorporating a more focused discussion (or adding to the intro) of how the findings advance our understanding of social hierarchy beyond prior works.

      We have sharpened the conceptual framing in both the Introduction and Discussion. In particular, we now distinguish more explicitly between the well-established observation that hierarchies form and remain stable within fixed semi-naturalistic groups, and the more specific question addressed here: whether an individual’s social position generalizes across changing social contexts created by repeated group reshuffling. We added additional references and also emphasize that the present study integrates continuous measurements of competition-based social rank, chasing, and reinforcement-learning features in the same individuals and environment. For details, see our response to Reviewer #2, Weakness (1).

      Manuscript changes:

      Introduction and Discussion revised to foreground conceptual novelty relative to prior semi-naturalistic work.

      (2) Given the weak associations between dominance rank and other traits such as body weight, cognitive performance, and oxytocin receptor manipulation, we suggest further clarifying what is being captured by these measures.

      We agree and have clarified this point throughout the manuscript. We now define tube-derived social rank explicitly as an operational measure based on repeated competitive outcomes from incidental dyadic tube tests, highlighting it as one dimension of social behavior.

      We also make clearer that proactive chasing is not redundant with rank, but instead captures a distinct, partly dissociable dominance-related interaction mode. For details, see our response to Reviewer #2, Weakness (4).

      We understand the question on the meaning of social rank as the correlations to body weight and cognitive performance are only punctual. We would like to mention here already that the competition-based social rank turns out to be a strong predictor of individual reactivity in a series of challenges. These works are in currently in preparation for publication and will make the relevance of these measures more clear. Related to this, we find in these studies that chasing and competition-based social rank predict different behavioral and neuronal aspects of individual reactivity.

      Manuscript changes:

      Abstract, Results, Discussion, and Conclusion revised to clarify construct interpretation and multidimensionality.

      (3) We recommend modifying the title and abstract to more clearly reflect the male-only design of the study. In addition, please indicate whether any age-related differences were observed. If age had no measurable effect, this should be stated explicitly to justify the combination of age groups.

      We revised the title and abstract to make the male-only design explicit. We also now report age-related analyses directly in the manuscript. Briefly, age was modeled by separating between-group age structure from within-group age variation. Some measures showed modest age associations in mean level, but most importantly, age-adjusted stability estimates for the core social metrics were highly similar to the original estimates, indicating that the reported stability is not driven by age differences across groups. For details, see our response to Reviewer #3, Major Concern (1).

      Manuscript changes:

      Title and abstract revised. Methods, Results, and supplementary robustness analyses expanded to report age effects explicitly.

    1. eLife Assessment

      This important study provides evidence that locus coeruleus activity is coordinated with heart rate during sleep, confirming previous work in mice and humans, with a possible role for sleep-dependent memory consolidation. The claims are supported by convincing evidence. This work will be of interest to neuroscientists focusing on sleep, memory, and autonomic functions.

    2. Reviewer #2 (Public review):

      Summary:

      This convincing study builds on previously published findings in both mice and humans to advance quantitative insights into the coupling between noradrenergic activity fluctuations during mouse NREM sleep and heart rate fluctuations. The work reaffirms the presence of coordinated infraslow fluctuations in sigma power and heart rate during NREM sleep and that this coordination is enabled by noradrenaline-releasing neurons in the locus coeruleus. Also supporting previously published work in mice and humans, the authors describe a link between the strength of these infraslow fluctuations and memory consolidation in mice and humans.

      Strengths:

      A major finding of this study is the mechanistic insight it provides into the regulation of the previously understudied very-low-frequency (0-0.15 Hz) component of heart rate variability, and the demonstration, through elegant optogenetic bidirectional interference, that infraslow noradrenergic fluctuations are an underlying driving force. This finding will promote recognition of heart rate variability in sleeping mice as a read-out of neuronal activity patterns that control autonomic balance.

      Another strength of the study is its translational part, whereby the sigma power-heart rate coupling in mouse is used to identify a previously unrecognized correlation between such coupling and memory consolidation in humans. This widens the applicability of heart rate variability measures, highlighting their use as biomarkers for noradrenergic fluctuations and associated sleep-dependent memory consolidation.

      Weaknesses:

      The study impresses by the thorough parallel analysis of both mouse and human correlational data between electrophysiological and fluorescent activity measures of the sleeping brain. Further work will be needed to disentangle the mechanisms by which heart rate is regulated, notably the contribution of parasympathetic and sympathetic nervous systems, to establish the very low frequency heart rate variability in mice as a novel biomarker for noradrenergic dynamics in the sleeping brain.

    3. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This study examined whether infraslow fluctuations in noradrenaline and in heart rate are coupled and how they are affected by sleep transitions. The authors used the fluorescent NA biosensor GRAB-NE2m in the medial prefrontal cortex of mice to record extracellular NA while also recording EEG and EMG during sleep-wake episodes. They also analyzed previously published human data to reproduce relationships they found between sigma power and RR intervals in mice.

      Strengths:

      This is an impressive study with significant strengths, as it involves a rich set of data that includes not only observations of associations between heart rate and noradrenergic dynamics but also optogenetic manipulation of the locus coeruleus. Human data is presented to show parallels in the association between sigma power during sleep and phasic heart-rate bursts.

      We thank Reviewer #1 for their thoughtful, detailed, and constructive evaluation of our manuscript. We appreciate their recognition of the strengths of the study, particularly the integration of noradrenergic recordings, optogenetic manipulation, and cross-species analyses. We are especially grateful for the reviewer’s careful attention to clarity, experimental interpretation, and control comparisons. The comments have helped us sharpen the framing of our hypotheses, clarify causal claims, improve statistical reporting, and better explain our closed-loop approach and heart rate analyses. We have addressed each point in detail below and believe that the revisions substantially strengthen the manuscript.

      Weaknesses:

      (1) Language could be clearer and more precise. As detailed below, in both the introduction and the discussion, the way the hypotheses and study objectives are described could use some revision to be more precise and accurate.

      Thank you for this helpful comment. We have sharpened the description of the study objectives, hypotheses, and interpretation of the findings to better distinguish between what was directly tested, what was inferred, and what remains speculative. We revised the language throughout these sections to improve clarity, accuracy, and overall readability

      (1A) In the introduction on p. 4: The overarching question is framed as "could the peripheral autonomous systems be a read-out of the central LC-NE system and thus be a biomarker of memory consolidation and LC dysfunction?" This gives the impression that the LC function would be the main influence on peripheral autonomous systems. There are, of course, many influences on peripheral autonomous systems, so it would be advisable for the authors to be more specific here about what signal(s) in particular would be predicted to be sensitive markers of LC function.

      Thank you for this important point. We agree that heart rate reflects the integrated output of multiple autonomic mechanisms and should not be interpreted as being exclusively driven by LC activity. Cardiac dynamics arise from the balance between sympathetic and parasympathetic influences, which themselves are regulated by several central and peripheral systems. In addition, recent work shows that the infraslow oscillations observed during NREM sleep are not restricted to norepinephrine alone but also involve other neuromodulatory systems, including acetylcholine and serotonin (e.g., Teng et al., PNAS 2025, Kjaerby et al., iScience, 2026). Our intention was therefore not to imply that the LC is the sole driver of peripheral autonomic dynamics. We have revised our overarching questions to make them more specific: Please see new text below:

      (Introduction, page 4/5). “The sympathetic and parasympathetic autonomic nervous system are involved in HRV, which is conventionally analyzed across three primary frequency bands: high frequency (HF) HRV, low frequency (LF) HRV, and very low frequency (VLF) HRV (Berntson et al., 1997). Due to the frequency overlap with VLF HRV, we wondered if central infraslow NE dynamics could be linked to this poorly understood HRV indicator. Furthermore, are infraslow NE fluctuations directly reflected by HRV under different physiological states or does LC–HR coupling scale differently with LC output? Specifically, if infraslow NE oscillations display faster frequencies - as occurs during sleep fragmentation - will cardiac dynamics exhibit corresponding changes? Conversely, given that stronger infraslow NE dynamics correlate with memory consolidation through their regulation of sleep spindles, could the peripheral autonomic signatures provide an accessible cross-species biomarker of spindle-dependent memory consolidation? Addressing these questions could help bridge mechanistic insights into LC-mediated sleep regulation with established HRV metrics used in human physiology.”

      (1B) In the discussion on p. 12: "In this study, we leveraged real-time measurements of mPFC NE levels and HR measurements from EMG recordings in mice to investigate the causal link between the two variables with high temporal resolution in freely moving sleeping mice, with similar inspection in humans." To test the causal link between mPFC NA levels and HR measures, the study would manipulate NA levels just in the mPFC and not elsewhere in the brain. However, in this study, the manipulation occurred in the LC, and so there would be broad cortical changes in NA levels. Thus, it could be that LC activity causes HR changes via a non-PFC pathway.

      We thank the reviewer for this important comment. Indeed, mPFC NE is merely a readout of LC activations and we expect that NE in other brain regions would show the same patterns. Indeed, mPFC NE is not expected to provide any causal link to heart rate. We have revised added a sentence to the results section and also changed the initial summary part of the discussion to reflect this better.

      (Results, page 6). “mPFC was selected as a representative cortical readout of LC-mediated norepinephrine dynamics, as infraslow NE fluctuations are coordinated across widespread brain regions.”

      (Discussion, page 14/15). “Variability in HR is a non-invasive biomarker of autonomic nervous system function and is frequently disrupted in ageing and Alzheimer’s disease. Here, by combining real-time measurements of mPFC NE dynamics with simultaneous HR recordings in freely sleeping mice, we demonstrate that HR closely tracks the infraslow phasic activity of the LC–NE system.”

      (2) Comparisons with the control condition need further development.

      (2A) While the authors did include a key YFP control condition, in the main text no direct statistical comparison between the closed-loop optogenetic stimulation (ChR2) condition and the YFP control condition was reported. (It was reported in Supplementary Figure 2c-d.) Instead, in the main text, the authors only reported that the effects of stimulation were significant in the closed-loop condition and not in the control. However, that is not the same as demonstrating that the two conditions significantly differed from each other, and it is the direct test that is important for the conclusions, so it seems important to include this result in the main presentation.

      We thank the reviewer for this important point and agree that direct statistical comparisons between ChR2 and YFP conditions are important for interpretation. These comparisons were performed and are shown in Supplementary Figure 2c–d, but we acknowledge that this was not sufficiently emphasized in the main text. We are now more clearly referring to this comparison in Result section:

      (Results, page 9). “The magnitude of pre-stimulation NE descent and post-stimulation NE ascent was reduced as the thresholds increased, indicating less pronounced NE dynamics as LC stimulation became more frequent (Fig. 2e, for direct comparison with YFP control, see Suppl. Fig. 2c-d).”

      Our rationale for prioritizing the within-animal threshold comparisons in the main figure was that the central experimental question concerned how progressive shifts in infraslow NE oscillatory frequency influence the NE–HR relationship. Because variability in viral expression levels (both NE sensor expression and LC opsin expression) introduces substantial between-animal variability, we considered within-animal comparisons across threshold conditions to provide the most informative representation of how changes in LC-driven NE dynamics alter cardiac responses. That said, we agree that highlighting the direct ChR2 versus YFP comparison is important for the overall interpretation. As mentioned, the reference to Suppl. Fig. 2c– d, where the between-group analyses are visualized are now clearly referred to.

      (2B) In addition, the authors should address the issue that the pre-stimulation NE was consistently significantly lower in the YFP condition than in the ChR2 condition (see Supplementary Figure 2c), which is a potential confound.

      We thank the reviewer for bringing up this important point. We agree that differences in stimulation timing between ChR2 and YFP animals could complicate interpretation in a closed-loop design and appreciate the opportunity to clarify this aspect of the experiment.

      In the ChR2 condition, animals were exposed to repeated optogenetic LC activation designed to mimic progressively faster infraslow NE dynamics. Such repeated stimulation is expected to produce a gradual elevation in tonic NE levels across the recording session, which explains the higher pre-stimulation baseline relative to YFP controls. We acknowledge that elevated tonic NE levels could introduce additional physiological effects. For example, higher NE tone would be expected to increase α2mediated autoinhibitory feedback on LC neurons and presynaptic NE release. Within the LC itself, we expect that optogenetic stimulation would largely override such effects due to the strong Na+-mediated depolarization induced by ChR2 activation. However, NE release in downstream regions such as the mPFC may be influenced to some extent by elevated tonic noradrenergic tone. Importantly, such feedback mechanisms would likely also occur under physiological conditions characterized by elevated LC activity, such as stress or sleep fragmentation. Because the goal of our stimulation paradigm was to model progressively faster infraslow NE dynamics under physiologically relevant conditions, we believe this feature of the manipulation may in fact increase the translational relevance of the model. We specifically address this elevation in Figure 3, where we show that very rapid stimulation regimes are accompanied by signs of compensatory cardiovascular regulation, likely reflecting baroreceptor-mediated responses to sustained increases in heart rate.

      We agree that the precise contribution of elevated tonic NE to the overall manipulation cannot be fully disentangled in the present study. We therefore avoid overinterpreting these effects and have instead added text to the manuscript acknowledging this consideration without extensive speculation.

      (Results, page 9) “Since pre-stimulation NE baseline levels progressively became higher in the ChR2 condition compared with YFP controls (Fig. 2c+e, Supplementary Fig. 3a-f), it demonstrates that they arise from stimulation-dependent modulation of noradrenergic tone rather than nonspecific signal drift. As a result, the pre-stimulation state at higher thresholds differed between ChR2 and control conditions, which should be considered when interpreting the immediate effects of laser stimulation across thresholds.”

      (2C) Direct comparison of the strengths of correlations shown in Figure 2h vs. Supplementary Figure 2f should be included. Currently, we see relatively weak correlations in both ChR2 and YFP conditions, and it is not clear if the relationships differ in the control. It seems they are still present in the control condition but weaker which would contradict the apparently broad claim on p. 7 that "No such effects were present in the control condition" (it is not entirely clear whether this claim refers to all effects discussed in the figure or just a subset - this language should be clarified).

      We thank the reviewer for this important comment and agree that the original wording could be interpreted as implying a complete absence of an NE–RR relationship in the YFP condition. To address this concern, we directly compared the strength of the NE– RR relationship between ChR2 and YFP animals. Please see Supplementary Figure 3h.

      Using a linear mixed-effects model that accounted for repeated measurements within animals, we found that the slope of the NE–RR relationship was significantly steeper in ChR2 animals than in YFP controls (all thresholds: slope difference = 1.915, p < 0.0001; thresholds −15, −10, and −5 only: slope difference = 2.367, p < 0.0001). Consistent with this result, comparison of Pearson correlations using Fisher's r-to-z transformation also indicated significantly stronger coupling in ChR2 animals than in YFP controls (see Statistics in the Supplementary File).

      These analyses demonstrate that an inverse relationship between NE and RR is present under physiological conditions in YFP animals, but that optogenetic LC activation substantially strengthens this coupling. We have revised the corresponding text to clarify this:

      (Results, Page 9): “An inverse NE–RR relationship was present in both ChR2 (Fig. 2h) and YFP (Suppl. Fig. 3f) groups but was significantly stronger in ChR2 animals than in YFP controls (Suppl. Fig. 3h).”

      (2D) Did the YFP controls vs. ChR2 animals show any differences in the number of NA states that triggered stimulation in the closed-loop system? With ChR2 animals, stimulation changes NA, which could change future triggering. In YFP animals, nothing changes NA (other than natural fluctuations), so the dynamics of stimulation timing could diverge between groups in a way that complicates interpretation. Specifically, if ChR2 stimulation raises NA and prevents future threshold crossings, ChR2 animals may end up receiving fewer subsequent stimulations than YFP animals (or a different temporal clustering). If the number or pattern of stimulation differed in two groups, it would be important to have a yoked control where matched animals get the same stimulation pattern but not triggered by their own NA.

      We thank the reviewer for this important point. We agree that differences in stimulation timing between ChR2 and YFP animals could complicate interpretation in a closed-loop design.

      Importantly, stimulation triggering was based on relative declines in NE fluorescence calculated against a rolling 2-minute baseline, rather than absolute NE levels. Thus, as tonic NE levels gradually increased in ChR2 animals, the threshold adapted accordingly, reducing the likelihood that elevated baseline NE alone would prevent future triggering. Instead, stimulation continued to occur when NE declined relative to the recent baseline, thereby preserving the infraslow closed-loop structure.

      The number of stimulation events across thresholds is already reported in the manuscript (Methods, p. 30 and corresponding figure legends), but we have now indicated more clearly in the result section where to find the information:

      (Results, Page 9). “Mean traces of NE and RR were aligned to LC stimulation onset (Fig. 2c-d, for number of laser stimulations see Fig. 2 legend or Methods).”

      (Methods, Page 31). “For the LC activation-related analysis, 108 events were found for Threshold -15 (11 of these being YFP), 260 events for Threshold -10 (55 of these YFP), 777 events for Threshold -5 (296 of these YFP), 1,444 events for Threshold 0 (510 of these YFP), and 1,148 events for Threshold 5 (377 of these YFP) across ten animals (four being YFP).”

      While the total number of events was lower in YFP animals, this is expected in part due to the smaller group size (4 YFP vs. 6 ChR2 animals included in this analysis). Furthermore, because ChR2 stimulation increased NE levels by design, more pronounced subsequent declines in NE may have modestly facilitated additional threshold crossings.

      Nevertheless, the overall temporal structure of the stimulation paradigm remained comparable across groups, and YFP animals underwent the same closed-loop stimulation protocol. Our primary comparison was mainly based on shifts in within-animal NE oscillatory frequency and how this change would impact the connection to HR. Thus, we believe the present control condition appropriately addresses the central question of whether optogenetic LC activation are able to conduct a continuum of NE oscillations.

      (3) Some more discussion/explanation of the rationale for the closed-loop approach and how it influences how we should interpret the results could be useful. For instance, currently, it is not clear whether LC stimulation needs to be timed after an NA dip to yield the effects seen.

      We thank the reviewer for pointing this out. The rationale for the closed-loop LC stimulation approach was to test whether the relationship between infraslow LC–NE dynamics and heart rate is maintained only under physiological infraslow conditions or whether it breaks down when the rhythm becomes progressively faster, as occurs during sleep fragmentation and other high-arousal states. Specifically, we asked whether heart rate continues to track LC–NE fluctuations as the infraslow rhythm shifts toward higher frequencies, thereby assessing its utility as a potential biomarker of disrupted restorative sleep.

      To address this while preserving the intrinsic temporal structure of infraslow LC activity, we implemented a closed-loop strategy in which stimulations were triggered following defined declines in the NE signal, using a rolling preceding 2-minute window as baseline. This allowed LC activation to occur during the descending phase of the endogenous infraslow cycle, maintaining its physiological phase structure while systematically increasing its effective frequency. By progressively relaxing the decline threshold, stimulations were triggered earlier in the cycle, thereby compressing the infraslow period in a controlled manner.

      Importantly, the intention was not to test whether LC stimulation specifically needs to occur after an NE dip to elicit the observed effects. Rather, triggering stimulation during the decay phase provided a way to accelerate the infraslow rhythm without disrupting sleep through indiscriminate stimulation. This enabled us to examine whether the coupling between LC–NE dynamics and heart rate remains stable under increasingly rapid infraslow regimes. Our results indicate that this relationship weakens at higher infraslow stimulation frequencies, suggesting that heart rate reliably reflects physiological LC–NE oscillations but becomes less tightly coupled when the rhythm is compressed beyond its normal range.

      We have clarified this rationale in the revised manuscript by adding the below section in the result section.

      (Result, page 8/9). “This approach enabled controlled compression of the infraslow NE cycle by triggering LC activation during the descending phase of the endogenous NE signal, thereby increasing the effective oscillatory frequency while preserving the temporal structure of physiological LC–NE dynamics. This strategy allowed us to test whether heart-rate responses continue to track LC-driven NE fluctuations as the infraslow rhythm becomes progressively faster.”

      (4) The section on heart rate decelerations is hard to follow. In particular, I was not sure how to interpret Figure 3f-j. For Figure 3f, what does the middle line represent? The laser onset or the max RR value after laser onset? What is the baseline that is used to correct the values to obtain amplitudes? If it is the whole period before the maximal RR value or the laser onset, wouldn't baseline values differ significantly across conditions and so potentially account for differences seen between conditions in the reported HR decelerations? Larger HR decelerations may be seen in conditions with higher HR simply as a regression to the mean phenomenon.

      We thank the reviewer for this feedback and agree that additional clarification of Figure 3f–j is warranted.

      For Figure 3f, the central line represents the peak RR value (maximal heart-rate deceleration) identified within the 2–7 s window following laser onset, rather than the laser onset itself. We realize this was not sufficiently clear and have revised the figure and corresponding Results text to clarify this point.

      Regarding baseline correction, RR amplitudes were calculated as the difference between the RR at peak deceleration and the mean RR during the 8–10 s period preceding the RR peak, as described in the manuscript. Thus, the baseline was defined locally for each event and was not based on the entire pre-laser period or stimulation onset. We chose this approach to account for shifts in baseline heart rate across conditions and to capture the relative magnitude of the deceleration response rather than absolute RR values.

      We appreciate the reviewer’s point regarding potential regression-to-the-mean effects, particularly in conditions with higher baseline heart rates. This is an important consideration. However, because the amplitude measure was baseline-corrected on an event-by-event basis, we believe the reported differences are unlikely to be explained solely by higher pre-stimulation heart rate. At the same time, our findings clearly show that elevated baseline heart rate influence the dynamic range of deceleration responses. Our findings show that under physiological conditions with elevated HR, larger heart-rate fluctuations would also contribute to increased HRV, which is often interpreted positively, despite potentially reflecting fragmented or dysregulated sleep states in this context.

      To improve readability, we have revised the figure and associated text.

      (Results, page 10). “To quantify the HR decelerations that happened after LC activation, we took the maximal RR value (so slowest HR) 2-7 s after LC stimulation and baseline corrected the value to the mean RR during the 8–10 s period preceding the RR peak to obtain their amplitude.”

      (5) The findings regarding LC suppression could be further clarified.

      (5A) Page 8: "observed a response in NE decline" - please be more precise. Did NE decline more or less?

      We thank the reviewer for this suggestion and agree that the original wording was imprecise. To clarify the direction and nature of the response, we have revised the text to state:

      (Results, page 11). “...we observed a gradual NE decline sustained throughout the laser period that was not observed in the YFP condition…”

      (5B) It would be helpful to also show the correlation between NE and RR in the control (YFP) condition and whether there were any differences between YFP and Arch conditions (Figure 4e).

      We thank the reviewer for this suggestion. We have now added the corresponding YFP correlation to Figure 4e. In the YFP group, the relationship between NE and RR showed a similar negative trend but did not reach statistical significance (p = 0.053). To directly assess whether the NE–RR relationship differed between Arch and YFP animals, we performed both a linear mixed-effects analysis and a Fisher r-to-z comparison.

      Neither analysis revealed a significant difference between groups. The linear mixed-effects model showed no significant Group × NE interaction (slope difference = 0.994, p = 0.51), indicating that the NE–RR coupling was not altered by LC suppression. Similarly, Fisher's r-to-z comparison found no significant difference between the correlations (p = 0.42).

      We believe this result is consistent with the relatively modest nature of the LC suppression paradigm. While Arch stimulation produced a clear reduction in NE levels, it did not induce a large shift in the overall NE–RR relationship. Instead, the data suggest that heart-rate responses remain coupled to noradrenergic fluctuations under both physiological conditions and during mild LC suppression. We have added the YFP data and clarified this interpretation in the revised manuscript.

      (Results, page 12): “A similar negative relationship was observed in YFP controls (Fig. 4e), and the strength of the NE–RR association did not differ significantly between Arch and YFP animals (Suppl. File, Statistics), suggesting that LC suppression did not substantially alter the underlying coupling between these measures.”

      (5C) This sentence took me multiple readings to understand - it would be helpful to rewrite to make it clearer: "indicating that, while HR generally did not respond strongly to LC suppression, the variability in RR responses was dependent on NE changes to the suppression (Figure 4e)."

      We agree with the reviewer that this phrasing is hard to understand and we have optimized for better clarity. Please see new version below:

      (Results, page 11/12). “Notably, despite the absence of a robust group-level HR effect, NE and RR responses remained negatively correlated across trials (Fig. 4e), indicating that HR dynamics continued to track the magnitude of noradrenergic suppression at the individual-response level”.

      (5D) The two colors in Figure 4 are similar and hard to distinguish.

      We agree that the colors are hard to separate and have altered them to make them easier to separate.

      (5E) The correlations shown in Figure 4j seem to be driven by just two of the cases. Are the effects significant when outliers are removed?

      We thank the reviewer for raising this point. To assess whether the observed correlation in Figure 4j was disproportionately driven by a small number of data points, we performed a formal outlier analysis using the ROUT method (Q = 1%). This analysis did not identify any statistical outliers in the dataset. Therefore, we did not have an objective basis for excluding any observations from the analysis.

      (5F) Page 10: Were there any differences in memory performance between the Arch and YFP conditions?

      We thank the reviewer for this question. The memory experiments were based on a previously published dataset (Kjaerby, Andersen et al., 2022), in which the primary objective was to assess the effect of LC suppression on sleep spindle dynamics and memory consolidation. In the present study, we performed an additional analysis by extracting heart-rate (RR) measures from these recordings to evaluate whether cardiac responses could serve as a biomarker of LC-mediated noradrenergic regulation.

      However, due to technical limitations (EMG recording often suffers from noise) in extracting reliable RR signals from all animals in this dataset, the number of subjects available for this secondary analysis was reduced. All these considerations are described in Methods/Mice. As a result, we were not sufficiently powered to perform a direct statistical comparison of memory performance between Arch and YFP groups based on RR measures alone. Instead, we examined whether RR responses to LC suppression predicted behavioral performance across animals. When pooling Arch and YFP conditions, we observed a correlation between the magnitude of the RR response and subsequent memory performance, suggesting that heart-rate dynamics reflect noradrenergic modulation relevant for memory consolidation.

      To avoid overinterpretation, we have therefore limited our conclusions to reporting this association rather than making direct group-level comparisons between Arch and YFP animals, and we have clarified this point in the revised manuscript.

      (Results, page 13): “Interestingly, across pooled Arch and YFP animals, larger RR increases following LC suppression were associated with better subsequent memory performance. Additionally, RR and NE responses to LC suppression were negatively correlated indicating that animals showing stronger NE reductions also exhibited larger RR changes. Together, these findings suggest that heart-rate dynamics covary with noradrenergic responses during sleep and may reflect physiological processes relevant for sleep-dependent memory consolidation.”

      (5G) Page 10: "We found a correlation between RR responses to LC suppression and sigma power, suggesting that a stronger HR reduction response is linked to higher spindle power." It should be noted in the text that the correlation was not specific to sigma (it was also seen for theta and beta, Figure 4i).

      We agree with reviewer that this should be highlighted. We have changed the sentence:

      (Results, page 12). “Furthermore, we found a correlation between RR responses to LC suppression and sigma power, suggesting that a stronger HR reduction response is linked to higher spindle power; similar correlations were also observed in the theta and beta frequency ranges (Fig. 4h-i).”

      (6) It is not clear which of the sigma power and RR interval findings do/do not exactly line up between the mice and humans. It could be helpful to have a table comparing them. For instance, was the finding in humans that pre-HRB sigma power was positively associated with slowing in heart rate after the HRB also seen in mice? Was there evidence in mice (as seen in the human sample) that sleep-dependent memory improvement was associated with pre-HRB sigma power?

      We thank the reviewer for this thoughtful comment and agree that the cross-species comparisons could be communicated more clearly. Our intention was not to imply exact one-to-one correspondence between all mouse and human findings, but rather to examine whether central–autonomic coupling surrounding phasic heart-rate events shows conserved features across species while acknowledging species-specific physiological differences.

      Importantly, the mouse and human analyses were designed to address related but not identical questions. In mice, we leveraged optogenetic LC suppression to probe a more causal relationship between noradrenergic activity, heart-rate slowing, spindle-related dynamics, and memory consolidation. Previous work using this dataset demonstrated that 2-minute LC suppression robustly enhances spindle density and that spindle enhancement correlates with improved memory performance. In the present study, we therefore asked whether heart-rate slowing covaries with this LC-mediated spindle/memory relationship, supporting HR as a potential biomarker of these restorative processes.

      By contrast, causal manipulation of LC activity is not feasible in humans. Instead, we focused on naturally occurring HRBs as putative downstream signatures of phasic LC– NE activity, motivated by our mouse findings that NE increases precede HR accelerations. We observed conserved coupling between sigma activity and HR dynamics across species, although the temporal profile differed, with sigma activity occurring closer to the HRB in humans than in mice. These temporal differences may reflect species-specific differences in cardiac and sleep physiology.

      Regarding the reviewer’s specific questions, the positive relationship between pre-HRB sigma power and post-HRB heart-rate slowing was tested in humans, where this metric showed the strongest relationship to behavioral outcome. We did not directly test the same measure in mice because, unlike humans, HR recovery following HRBs did not show a pronounced baseline shift (Fig. 5b), limiting the interpretability of this comparison. Similarly, we did not directly correlate pre-HRB sigma power with memory performance in mice, as the more causal LC suppression paradigm already demonstrated a spindle–memory relationship in this species and was the focus of our mechanistic analysis.

      To reduce confusion, we have revised the Results conclusion to make a clearer overview:

      (Results, page 14): “In conclusion, mice and humans displayed evidence of conserved autonomic-central coupling, reflected in coordinated HR and sigma power dynamics surrounding phasic cardiac events.

      However, the temporal relationship between these events differed across species, likely reflecting differences in sleep and cardiovascular physiology. In mice, causal manipulation of LC activity demonstrated that heart-rate dynamics covary with LC-mediated noradrenergic and spindle-related processes linked to memory consolidation. In humans, sigma power preceding HR bursts was associated with both post-HRB heart-rate slowing and sleep-dependent memory improvement, suggesting that autonomic–central coupling surrounding HR events may provide a translational marker of restorative sleep processes (Fig. 5n).”

      (7) Page 18: It is not clear if the sex of mice was balanced across controls and optogenetics groups.

      We thank the reviewer for this important comment and agree that the sex distribution should be reported more clearly. These experiments relied on the availability of animals from the heterozygous TH-Cre transgenic line, and given the relatively small cohort sizes, perfect balancing across sex and experimental groups was not always feasible.

      For the LC activation experiments, the sex distribution was: YFP: 2 male / 2 female; ChR2: 4 male / 2 female. For the LC suppression experiments, the distribution was: YFP: 4 female; Arch: 3 female / 1 male.

      Although the groups were not perfectly sex balanced, we had no strong reason to expect robust sex-dependent differences in the physiological effects of these optogenetic manipulations, particularly given the relatively strong and acute nature of the intervention. At the same time, we acknowledge that the present study was not powered to assess sex as a biological variable, and subtle sex-dependent effects therefore cannot be excluded. To improve transparency, we have now clarified the sex distribution in the Methods/Mice section.

      Reviewer #2 (Public review):

      Summary:

      The major part of this study reproduces previously published findings in both mice and humans and provides incremental analyses on these findings. In essence, the work reaffirms the presence of coordinated infraslow fluctuations in sigma power and heart rate during NREM sleep. It further confirms previous findings that coordination depends on noradrenaline-releasing neurons in the locus coeruleus. Also supporting previously published work in mice and humans, the authors describe a link between the strength of these infraslow fluctuations and memory consolidation in mice and humans.

      Strengths:

      The authors successfully replicate key previously reported phenomena across both mice and humans. Confirmatory studies and demonstrations of reproducibility are essential for progress in neuroscience. To maximize their value, such studies should clearly acknowledge their confirmatory nature and carefully situate what, in their view, are novel results, going beyond existing literature.

      Weaknesses:

      The authors' interpretation of their data needs to be revised. Many of their claims regarding the mechanistic basis of their findings and the predictive value of their correlative datasets are not supported by the available evidence.

      In the present manuscript, several citations of literature on the work they reproduce lack precision or completeness, which reduces transparency and obscures how the reported findings relate to previously established results.

      We thank Reviewer 2 for the thoughtful comment regarding positioning of our findings relative to the literature, and caution in mechanistic interpretation. In response, we have revised the Introduction, Results, and Discussion to more clearly acknowledge foundational studies in this area and to better clarify how the present work extends beyond them.

      We agree that prior work has demonstrated infraslow coupling between sigma activity, norepinephrine (NE) dynamics, and heart rate (HR), and has established a role for the locus coeruleus (LC) in coordinating these oscillations. However, cardiac measures in these studies were typically treated as secondary observations rather than as primary experimental targets. A central goal of the present study was therefore to provide a systematic and mechanistically grounded characterization of NE-mediated HR dynamics during sleep across multiple timescales, including infraslow oscillations, sleep–wake transitions, and causal manipulations of LC activity.

      Importantly, we also aimed to relate infraslow HR fluctuations to the very-low-frequency (VLF) component of heart rate variability (HRV), which remains comparatively under-characterized and mechanistically unresolved in the clinical HRV literature. By linking LC activity, NE dynamics, and HR fluctuations across behavioral states, our findings provide a biologically grounded framework that may help explain this component of HRV.

      A second major objective of the study was translational. Because direct LC recordings are not feasible in humans, we asked whether cardiac dynamics alone could reflect the infraslow, memory-consolidating potential of sleep and thus serve as a noninvasive biomarker. By directly manipulating LC activity and demonstrating corresponding changes in HR dynamics, our results strengthen the mechanistic rationale for using HRV—particularly its VLF component—as an accessible proxy of LC-dependent sleep physiology.

      We therefore respectfully disagree with the suggestion that the present study does not provide novel insight. Rather, the revised manuscript now more clearly emphasizes that our contribution lies in (i) systematically characterizing NE-dependent HR dynamics across sleep states, (ii) linking these dynamics to the poorly understood VLF component of HRV, and (iii) establishing a causal and translational framework for using cardiac measures as markers of LC-mediated sleep processes.

      We hope the reviewer finds that the revised Introduction and Discussion better highlight both the existing literature and the specific advances provided by the present work.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      I have been convinced by discussions about the replicability crisis that it should be a standard practice to share data upon publication in a publicly accessible online repository such as Open Science Framework or OpenNEURO. Making data available can increase the impact of the research study. Simply stating "data available upon request" as done in the current draft is not sufficient as, unfortunately, when data are not shared upon publication in a public repository, it can be impossible to gain access by request from the researchers - the majority of requests from other researchers to obtain data are not complied with (e.g., Vanpaemel, Vermorgen, Deriemaecker, & Storms, 2015; Wicherts, Bakker, & Molenaar, 2011).

      We fully agree with the reviewer about the importance of data sharing for transparency, reproducibility, and maximizing the impact of research. Consistent with these principles, we will make the human dataset publicly available on the Open Science Framework upon publication and have added the link to this repository in the manuscript under Data Availability (https://osf.io/g6emj/). With respect to the mouse dataset, we respectfully note that this dataset is currently the subject of multiple planned and ongoing analyses that extend beyond the scope of the present manuscript. Releasing these data publicly at this stage could compromise these efforts and lead to potential misinterpretation prior to completion of the full analytic pipeline. For this reason, we believe that it would not be appropriate to share the mouse data in a public repository at this time. However, we remain committed to transparency and will make the mouse data available upon reasonable request during this period, with the intention of publicly releasing the dataset once the planned analyses are complete.

      (1) p. 4: There is some orphan text at the top of the page ("marker of Alzheimer's disease. Furthermore, maintaining LC neural density prevents neurodegeneration (13). Given the reported reduction in HRV in aging and Alzheimer's disease, suppressed").

      Thank you. We have removed the orphan text.

      (2) p. 22: What does NFR refer to?

      Novel-to-familiar ratio. We have removed the abbreviation from the main text. It is now only used in the figure.

      (3) I might have missed this information, but it was not clear to me how the epochs to be analyzed were selected, and for the mice, how much wake vs. sleep they included.

      We thank the reviewer for this comment and apologize that the epoch selection criteria were not sufficiently clear. They were mentioned under Methods. For the optogenetic analyses in mice, epochs were selected based on NREM sleep including microarousals to ensure that physiological responses were evaluated within stable sleep conditions while allowing for natural brief interruptions.

      For the LC activation (ChR2) experiments, stimulation epochs were included only if NREMinclMA began at least 30 s before laser onset and continued for at least 30 s after laser onset. For the LC suppression (Arch) experiments, the criterion was similarly ≥30 s of NREMinclMA prior to laser onset, but extending ≥60 s following laser onset to accommodate the longer suppression response profile.

      Thus, analyses were intentionally restricted to sleep periods, and wakefulness was not included except where it emerged naturally as an outcome of the manipulation or transition under investigation. We have now clarified these inclusion criteria in the Methods/ Event marker selection section to improve transparency.

      (4) In the Supplement, Figure 3 appears before Figure 2.

      Thanks. We have corrected it.

      Signed, Mara Mather

      Reviewer #2 (Recommendations for the authors):

      The authors' interpretation of their data needs to be revised. Many of their claims regarding the mechanistic basis of their findings and the predictive value of their correlative datasets are not supported by the available evidence.

      There are three major directions in which this study would need to be revised.

      Part 1 - Literature citations of both mouse and human literature need to be revised, and citations placed in a manner that accurately reflects what has been previously done. In detail:

      (1.1) The statement on p. 4 regarding "the extent to which phasic infraslow NE fluctuations ...is not well understood" disregards a previous publication in which closed-loop optogenetic stimulation of LC was already shown to regulate HR variations on the infraslow time scale (10.1016/j.cub.2021.09.041).

      We thank the reviewer for this important point. The study by Osorio-Forero et al. (2021) was already cited as ref. 4 in our original manuscript and in the discussion, we specifically highlighted this study: ‘These findings build on Osorio-Forero et al. (35) as well as other studies’; however, we agree that our wording did not sufficiently emphasize its key finding that closed-loop optogenetic manipulation of LC activity can coordinate infraslow heart-rate fluctuations and spindle clustering during NREM sleep.

      We have revised the relevant paragraph in the introduction to explicitly acknowledge that ref. 4 demonstrated causal coordination between LC activity, sleep spindle dynamics, and heart-rate fluctuations. We have also refined our statement of the knowledge gap to clarify which novel questions our study addresses. We believe these revisions more accurately position our work within the existing literature and clearly distinguish our contributions from prior studies.

      We have updated the introduction in several places to address these comments:

      (Introduction, page 4). “It has previously been reported that HR fluctuates at similar infraslow frequencies as NE fluctuations and sigma power in mice (Lecci et al., 2017; Osorio-Forero et al., 2021) and that phasic HR fluctuations correlate with infraslow changes in pupil diameter (Carro-Domínguez et al., 2025), a proxy for changes in NE levels (Murphy et al., 2014; Reimer et al., 2016). Importantly, optogenetic manipulation of the LC has demonstrated that infraslow LC activity coordinates sleep spindle clustering and heart rate fluctuations during NREM sleep (Osorio-Forero et al., 2021). Together, these findings support a functional coupling between the central LC–NE system and peripheral cardiac dynamics, further supported by findings that HR increases accompany MAs during NREM sleep (Carro-Domínguez et al., 2025;

      Osorio-Forero et al., 2025).”

      (Introduction, page 4/5). “Due to the frequency overlap with VLF HRV, we wondered if central infraslow NE dynamics could be linked to this poorly understood HRV indicator. Furthermore, are infraslow NE fluctuations directly reflected by heart-rate variability under different physiological states or does LC–HR coupling scales differently with LC output? Specifically, if infraslow NE oscillations display faster frequencies - as occurs during sleep fragmentation - will cardiac dynamics exhibit corresponding changes? Conversely, given that stronger infraslow NE dynamics correlate with memory consolidation through their regulation of sleep spindles, could the peripheral autonomic signatures provide an accessible cross-species biomarker of spindle-dependent memory consolidation? Addressing these questions could help bridge mechanistic insights into LC-mediated sleep regulation with established HRV metrics used in human physiology.”

      (1.2) In this same published paper, optogenetic stimulation of LC was already used to "determine the causal relationship..." (p.6). However, the authors do not cite these data.

      We thank the reviewer for this comment. As noted in our response to Comment 1.1, ref. 4 was already cited in the Introduction, and we have now revised that section to more explicitly emphasize this paper. In addition, we have modified the wording in the Results section.

      (Results, page 8). “After finding the inverse correlation between NE and RR in natural sleep transitions, we next sought to further characterize the causal influence of LC activity on HR dynamics, using a closed-loop optogenetic approach to modulate NE oscillatory frequency during NREM sleep.”

      (1.3) The authors' speculation about LC-induced sympathetic and parasympathetic actions is premature: this study does not provide pharmacological experiments in this direction. However, two published studies implied a parasympathetic mechanism linked to infraslow fluctuations of LC activity (10.1016/j.cub.2021.09.041, 10.1016/j.cub.2017.12.049). These findings should be appropriately cited. While HR decelerations may reflect compensatory autonomic responses, there is no direct evidence in the present study that these effects are sympathetically mediated. An alternative, and equally plausible, interpretation is enhanced parasympathetic activity. This distinction is particularly important given that the observed increases in mean HR during LC stimulation cannot distinguish between reduced parasympathetic activity and increased sympathetic drive.

      We thank the reviewer for raising this important point and agree that the current study does not provide direct mechanistic evidence to disentangle sympathetic versus parasympathetic contributions to LC-mediated heart rate regulation. This was not the intention of our study; rather, the relevant discussion section was meant to provide mechanistic interpretations and hypotheses based on the observed physiology. To avoid overstating our conclusions, we have revised the wording to more clearly emphasize the speculative nature of these interpretations. Furthermore, we have added the suggested references demonstrating that muscarinic blockade reduces heart rate and pupil fluctuations during sleep, which support the possibility of a parasympathetic contribution to infraslow LC-related dynamics.

      (Discussion, page 16/17). “Prior findings demonstrate that LC is linked to the autonomic nervous system. Stimulation of LC projections decrease parasympathetic cardiac vagal activity (Wang et al., 2014) and also influences sympathetic output through direct projections to the preganglionic cells in the sympathetic nervous system (Karemaker, 2017; Nygren and Olson, 1977; Samuels and Szabadi, 2008). While the mechanisms generating VLF HRV are not well defined (Armour, 2003; Shaffer et al., 2014; Wang et al., 2014), there is a clear parasympathetic component (Taylor et al., 1998). This combined with the ability of pharmacological blockage of the parasympathetic system to block infraslow oscillations of HR (Osorio-Forero et al., 2021) and pupil diameter during sleep (Yüzgeç et al., 2018), led us to expect a slowing of HR during LC suppression due to parasympathetic disinhibition.”

      (1.4) It is not clear why prefrontal NE signals were associated with HR fluctuations. Literature evidence indicates other brain areas that are functionally more directly linked to autonomous fluctuations.

      We thank the reviewer for raising this point. We do not speculate that mPFC NE activity is causally linked to heart rate fluctuations or that the mPFC directly mediates the observed autonomic dynamics. Instead, mPFC NE signaling was used as an experimentally accessible readout of LC activity. Importantly, accumulating evidence suggests that infraslow NE oscillations are globally coordinated phenomena that are expressed across multiple brain regions during sleep, making mPFC NE a valid proxy for LC-driven neuromodulatory state dynamics. We have clarified this in the result section:

      (Results, page 6). “mPFC was selected as a cortical readout of LC-mediated norepinephrine dynamics, as infraslow NE fluctuations are coordinated across widespread brain regions.”

      (1.5.a) The lack of effect of Arch-inhibition of LC on infraslow NE signals is concerning. Prior work showed that bilateral LC inhibition does affect NE signals and also infraslow sigma power fluctuations (10.1016/j.cub.2021.09.041, 10.1038/s41593-024-01822-0). This discrepancy should be explicitly acknowledged and discussed on p.9.

      We thank the reviewer for highlighting this important point. To clarify, we did observe a robust effect of LC inhibition on noradrenergic signaling, with clear suppression of the NE signal following Arch-mediated LC inhibition (Fig. 4c), aligned to laser onset. In addition, LC suppression increased neuronal synchronization, including enhanced sigma power relative to YFP controls (Fig. 4h), consistent with previous reports showing that reduced LC activity promotes synchronized sleep-related oscillations. Thus, we do not interpret our findings as indicating an absence of LC suppression effects.

      The apparent discrepancy relates specifically to the absence of a statistically significant group-level shift in infraslow NE or HRV power during NREM sleep, rather than the efficacy of the manipulation itself. We note that our inhibition paradigm was intentionally mild, consisting of repeated 2-minute suppression periods separated by 4-minute intervals, and was designed to introduce subtle shifts within physiological ranges rather than globally reorganize infraslow sleep structure. Accordingly, we consider the immediate NE and heart-rate responses to LC inhibition to be the most sensitive physiological readouts of LC-mediated regulation in this context.

      We also note that the studies cited by the reviewer used different suppression paradigms, including more frequent manipulations. Thus, while our suppression scheme did not result in detectable infraslow reorganization, we do not believe they reflect an ineffective LC suppression as we demonstrated a clear NE reduction in response to time-locked LC suppression.

      We have added a sentence to the result section to explain the lack of effect infraslow power:

      (Results, page 12). “This likely reflects the relatively mild and intermittent LC suppression paradigm, which was designed to remain within physiological ranges and therefore did not globally reorganize infraslow sleep dynamics.”

      (1.5.b) Additionally, the authors state in the Discussion that they "find no consistent modulation in HR during LC suppression, suggesting that the LC-HR connection is more strongly associated with sympathetic activity rather than parasympathetic inhibition". However, the results primarily demonstrate an absence of modulation in mean HR, while preserving a significant relationship between RR intervals and stimulation. This indicates a modulation of heart rate variability, even in the absence of changes in average HR. Notably, such variability-related effects may fall outside the VLF range and could instead involve higher-frequency components.

      We thank the reviewer for this important clarification. We agree that our original wording may have conflated the absence of modulation in mean HR with the absence of autonomic modulation more generally. To address this point, we revised the Discussion to emphasize that LC activity may influence VLF HRV through sympathetic activation and/or indirect modulation of cardiac vagal activity, even in the absence of robust changes in average HR. We have softened our previous interpretation that the LC-HR relationship is primarily sympathetic in nature and instead discuss a more nuanced interaction between sympathetic and parasympathetic influences on HRV dynamics.

      (Discussion, page 16/17). “Prior findings demonstrate that LC is linked to the autonomic nervous system. Stimulation of LC projections decrease parasympathetic cardiac vagal activity (Wang et al., 2014) direct projections to the preganglionic cells in the sympathetic nervous system (Karemaker, 2017; Nygren and Olson, 1977; Samuels and Szabadi, 2008). While the mechanisms generating VLF HRV are not well defined (Armour, 2003; Shaffer et al., 2014; Wang et al., 2014) there is a clear parasympathetic component (Taylor et al., 1998). This combined with the ability of pharmacological blockage of the parasympathetic system to block infraslow oscillations of HR (Osorio-Forero et al., 2021) and pupil diameter during sleep (Yüzgeç et al., 2018) led us to expect a slowing of HR during LC suppression due to parasympathetic disinhibition. Interestingly, we found no consistent modulation in HR during LC suppression, suggesting that the LC-HR connection may also somehow be driven by sympathetic outflow. Previous research had indicated that LF power may also represent sympathetic activity, but this interpretation has been challenged due to the mixed contribution of both autonomic branches (Houle and Billman, 1999; Japundzic et al., 1990; Reyes et al., 2013). LF and HF ratio (LF/HF) were traditionally thought to reflect balance between sympathetic and parasympathetic activities (i.e., the sympatho-vagal balance), though currently considered as an oversimplification of non-linear integration of autonomic signals (Billman, 2013; Pagani et al., 1986). Our findings implicate the VLF may offer a precise marker for central arousal states, given its overlaps with infraslow phasic fluctuations of LC-NE levels. Although parasympathetic activity appears important for the expression of VLF oscillations, growing evidence suggests that VLF dynamics reflect broader interactions between the heart and autonomic nervous system rather than simple sympathetic or parasympathetic control alone (Armour, 2003; Shaffer et al., 2014). Within this framework, infraslow LC–NE dynamics may represent one central contributor to these slow cardiac fluctuations during sleep either directly through sympathetic activation or indirectly by inhibition of cardiac vagal activity.”

      (1.6) The relationship between sigma power fluctuations and HR is different in humans than in mice. This has been shown before (10.1126/sciadv.1602026, 10.1038/s41593025-02159-y). This work should be mentioned on p. 11.

      We already acknowledge prior studies demonstrating species differences in the relationship between sigma power fluctuations and heart rate in the discussion section. However, to accommodate the reviewer’s comment and improve clarity for the reader, we have now also added the suggested references to the Results section, where the relationship between sigma power fluctuations and HR is first discussed.

      (Results, page 14). “These temporal differences may reflect species-specific physiology differences in cardiac timescales, which has also been previously reported (Bergel et al., 2025; Carro-Domínguez et al., 2025; Lecci et al., 2017).”

      (1.7) Correlations between the strength of infraslow sigma power fluctuations and memory consolidation have been published and should be discussed (10.1126/sciadv.1602026). It is surprising to see that correlations with learning in humans are done using pre-HRB sigma peaks rather than heart rate. This is a measure that is very close to the one used by Lecci et al.; this similarity should be clearly acknowledged. The way the data are currently presented limits this manuscript's novelty, also in its translational aspect.

      We agree that the work by Lecci et al. (2017) established an important relationship between the association of infraslow sigma power fluctuations and memory consolidation, which is highly relevant to our findings.

      Importantly, the underlying infraslow fluctuations in neuromodulatory tone are increasingly recognized as key regulators of sleep spindle dynamics (sigma power), including from our own previous work demonstrating that direct manipulation of locus coeruleus–norepinephrine infraslow rhythms alters spindle organization and sleep continuity. Thus, our findings are conceptually aligned with prior studies linking sigma fluctuations to memory consolidation.

      However, we would like to clarify an important distinction in our translational approach. While Figure 4 demonstrates that heart rate dynamics during sleep can predict memory performance in mice, Figure 5 was designed to address the translational potential of these findings in humans, where direct neuromodulatory readouts are not readily accessible. Here, we deliberately focused on heart rate bursts and their associated sleep dynamics as a clinically tractable physiological measure.

      We acknowledge that the pre-HRB sigma increase may appear conceptually similar to the measure used by Lecci et al.; however, our approach is not equivalent. Rather than selecting spindle or sigma peaks themselves, we aligned analyses to heart rate accelerations and examined the robust upregulation of sigma activity preceding these events. In this framework, sigma activity serves as a physiological readout linked to autonomic dynamics, rather than being the primary anchor of analysis. We chose this measure because heart rate bursts are influenced by multiple physiological factors, and the associated sigma dynamics provided the clearest and most robust relationship with memory outcomes in the human dataset.

      We have revised the Discussion to more clearly acknowledge the similarity to prior work. We believe our findings extend prior observations by providing evidence that sleep-related heart rate fluctuations may serve as a non-invasive readout of the memory-preserving function of sleep.

      (Discussion, page 18). “Previous work in humans demonstrated that the strength of infraslow sigma oscillations correlates with sleep-dependent memory consolidation in humans (Lecci et al., 2017).”

      (1.8) Conclusions as to whether sigma fluctuations might be slightly slower and less powerful in mice are not justified. More work is required to determine which infraslow manifestations are most useful for cross-species comparisons. Moreover, little is currently known about LC activity in human sleep. A careful look into how pupil diameter correlates with sigma power should provide clues for further discussion (see Carro-Dominguez et al). A detailed study of infraslow fluctuations in human sleep should also be discussed https://doi.org/10.1101/2024.11.06.620875.

      We thank the reviewer for this thoughtful comment. We agree that our original phrasing suggesting that sigma dynamics in mice may be “slower and less powerful” than in humans was overly interpretive. We have removed this sentence from the Results section. In addition, we have expanded the Discussion to more thoroughly integrate recent human literature on infraslow sleep dynamics.

      (Discussion, page 19). “Importantly, infraslow fluctuations of sigma power in human sleep have received growing attention. Recent work demonstrates that the infraslow fluctuation of sigma power segments N2 sleep into functional phases associated with arousal and memory-related sleep markers (Dimitriades et al., 2024). Complementary findings using pupillometry show that pupil diameter fluctuates on similar infraslow timescales during NREM sleep and is inversely related to spindle clustering, providing indirect evidence that arousal-related noradrenergic dynamics shape human sleep microstructure (Carro-Domínguez et al., 2025). However, LC activity during human sleep remains inferred rather than directly measured, and systematic perturbation studies linking LC output to spindle–autonomic coupling in humans are currently lacking. Together, these observations underscore both the promise and the current limitations of cross-species comparisons of infraslow sleep dynamics.”

      (1.9) Citation of literature should be as explicit as possible. Referring to "many studies rely on plasma levels..." while including some that actually did real-time fiber photometric measures is misleading.

      We thank the reviewer for this suggestion and agree with the reviewer about the importance of accurately representing prior studies. We have now updated the discussion to clarify this.

      (Discussion, page 15). “Many studies have linked HR to NE (Fawaz and Simaan, 1963; Sundaram et al., 1991; Tanoue et al., 2022; Watson et al., 1979). Many rely on plasma levels of NE, which, while linked to central NE (Gurguis and Uhde, 1998), has low temporal resolution, making causal interpretations harder. In recent years, the use of biosensors and fibre photometry allows for very reliable estimate of the temporal dynamics of NE changes making association to HR more precise (Osorio-Forero et al., 2021).”

      (1.10) Regarding the discussion on the baroreflex: The emphasis is placed predominantly on sympathetically mediated effects. However, the description of the baroreflex loop is incomplete, as it overlooks the substantial contribution of parasympathetic modulation. In particular, heart rate adjustments within the baroreflex are primarily mediated by parasympathetic mechanisms.

      We agree with reviewer that this important notion should be added. We have rephrased the discussion as below:

      (Discussion, page 20). “These neurons suppress the activity of the rostral ventrolateral medulla, ultimately resulting in reflex parasympathetic activation with sympathetic inhibition lowering the HR (Aicher et al., 2000; Lanfranchi and Somers, 2002).”

      (1.11) Regarding the interpretation of HRV analysis, the discussion places disproportionate weight on sympathetic modulation in the interpretation of HRV metrics. This framing is inconsistent with recent conceptual clarifications, including a recent Nature Reviews Cardiology article by Menuet et al. (10.1038/s41569-02501160-z), which cautions against simplistic low-frequency/high-frequency (LF/HF) interpretations of autonomic balance.

      We thank the reviewer for this thoughtful comment. We agree that HRV frequency bands should not be interpreted as exclusive markers of specific autonomic branches. In the original manuscript, we cited Billman (2013), which challenges the validity of LF/HF as a measure of sympatho-vagal balance, to acknowledge these conceptual limitations. However, we recognize that some of our phrasing, particularly in the section discussing compensatory HR decelerations, may have implied branch-specific dominance.

      We have updated the Discussion section as follows:

      (Discussion, page 17). “Interestingly, we found no consistent modulation in HR during LC suppression, suggesting that the LC-HR connection may also somehow be driven by sympathetic outflow. Previous research had indicated that LF power may also represent sympathetic activity, but this interpretation has been challenged due to the mixed contribution of both autonomic branches (Houle and Billman, 1999; Japundzic et al., 1990; Reyes et al., 2013). LF and HF ratio (LF/HF) were traditionally thought to reflect balance between sympathetic and parasympathetic activities (i.e., the sympatho-vagal balance), though currently considered as an oversimplification of nonlinear integration of autonomic signals (Billman, 2013; Pagani et al., 1986). Our findings implicate the VLF may offer a precise marker for central arousal states, given its overlaps with infraslow phasic fluctuations of LC-NE levels. Although parasympathetic activity appears important for the expression of VLF oscillations, growing evidence suggests that VLF dynamics reflect broader interactions between the heart and autonomic nervous system rather than simple sympathetic or parasympathetic control alone (Armour, 2003; Shaffer et al., 2014). Within this framework, infraslow LC–NE dynamics may represent one central contributor to these slow cardiac fluctuations during sleep either directly through sympathetic activation or indirectly by inhibition of cardiac vagal activity.”

      We have also replaced the title of the Discussion section title “Locus-coeruleus-mediated sympathetic control drives compensatory heart rate decelerations” with “Locus-coeruleus activation drives compensatory heart rate decelerations via autonomic feedback mechanisms”

      (1.12) In particular, the manuscript attributes VLF power primarily to sympathetic activity, despite evidence that very-low-frequency RR-interval oscillations are strongly dependent on parasympathetic integrity. Notably, parasympathetic blockade has been shown to nearly abolish VLF oscillations in humans (Taylor et al., Circulation, 1998; doi:10.1161/01.CIR.98.6.547).

      We thank the reviewer for highlighting this important point. It was not our intention to imply that VLF HRV is driven primarily by sympathetic activity or to disregard the important role of parasympathetic integrity in shaping VLF oscillations. Our intention was to present the multifactorial and incompletely resolved nature of the VLF component; however, if our wording can be interpreted otherwise, we agree that clarification is warranted.

      In response, we have revised the manuscript to better reflect the current understanding of VLF physiology. Specifically, we now make clearer that, while VLF HRV remains less mechanistically defined than HF and LF HRV, substantial evidence supports a strong parasympathetic contribution to the expression of VLF oscillations. We have incorporated the reviewer-suggested references within this comment and others and clarified this point throughout the revised manuscript.

      (Discussion, page 16). “While the mechanisms generating VLF HRV are not well defined (Armour, 2003; Shaffer et al., 2014; Wang et al., 2014) there is a clear parasympathetic component (Taylor et al., 1998).”

      (Discussion, page 17). “Although parasympathetic activity appears important for the expression of VLF oscillations, growing evidence suggests that VLF dynamics reflect broader interactions between the heart and autonomic nervous system rather than simple sympathetic or parasympathetic control alone (Armour, 2003; Shaffer et al., 2014).”

      (1.13) Moreover, based on work by Armour (2003) and Kember et al. (2000, 2001), the VLF rhythm is thought to emerge from stimulation of afferent sensory neurons within the heart, further arguing against a purely sympathetic interpretation. Together, these findings indicate that the discussion overemphasizes sympathetic mechanisms and underrepresents the contribution of parasympathetic and afferent cardiac pathways to HRV, particularly in the VLF range.

      We thank the reviewer for this important point. Our response to this comment is largely aligned with our response to comment (1.12). In the revised Discussion, we now more clearly acknowledge that VLF oscillations likely arise from more complex cardioautonomic interactions than simple sympathetic and parasympathetic innervation alone. At the same time, we have intentionally avoided an extensive discussion of these mechanisms, as our experimental design does not directly address these pathways.

      Part 2 - There are a number of conceptual issues that need more careful elaboration:

      (2.1) A highly problematic point throughout this study is the choice of AUCs, notably of NE signals and RR intervals, rather than the signal amplitudes. It confuses the correlations to events of different durations, such as MAs and wakefulness. It is not possible to draw conclusions of the kind "cardiac rhythm are tightly coupled with the infraslow phasic NE ..." because such statements ignore that AUCs conflate amplitude and duration of a signal.

      We thank the reviewer for this comment and agree that the distinction between amplitude- and duration-related signal features is important. We deliberately chose AUC measures because our intention was to capture the overall physiological response over time, including both the magnitude and temporal evolution of the signal, rather than relying solely on a single peak value. In this context, AUC provides information about the shape and sustained nature of NE and RR changes within a defined time window, which we considered particularly relevant for temporally dynamic responses.

      At the same time, we appreciate the reviewer’s concern that AUC may conflate response amplitude and duration, particularly when comparing events of different lengths such as microarousals and wakefulness. To minimize this issue, the AUC windows were intentionally kept relatively short and fixed, thereby limiting the influence of prolonged wake episodes or differences in transition duration on the measure. Thus, our intention was not to quantify the total duration of awakenings, but rather the immediate physiological response profile surrounding the event.

      To address this concern more directly, we compared amplitude- and AUC-based measures across all mice. The results are now shown in Suppl. Figure 1g. We found a strong correspondence between amplitude and AUC measurements for NE signals, indicating that the observed relationships are not dependent on the choice of metric. A similar, albeit weaker, relationship was observed for RR responses. This likely reflects physiological constraints on heart-rate dynamics, where the initial heart-rate acceleration is relatively similar across vigilance-state transitions (Figure 1e), while the duration of the response differs substantially. As a result, amplitude measures may underestimate differences between transitions, whereas AUC better captures the extent of the cardiac response. Importantly, direct comparison of NE and RR amplitudes still revealed a significant relationship, supporting the overall conclusion that cardiac and noradrenergic responses are coupled. However, this relationship was weaker than that observed using AUC measures, suggesting that incorporating temporal aspects of the response captures additional biologically relevant information.

      In addition to the figure, we have revised the Results section to include these considerations:

      (Results, page 7): “These comparisons were performed using area under the curve (AUC) estimates of NE and R-R responses. Importantly, comparing AUC and peak amplitude measures showed a similar overall relationship, and the coupling between NE and RR remained significant when only peak amplitudes were considered (Suppl. Fig. 1g), indicating that the findings are not solely driven by the response duration aspect of AUC. The somewhat stronger relationship observed with AUC-based measures may reflect rapid saturation of heart-rate responses across vigilance-state transitions, making response persistence an informative component of the physiological signal.”

      (2.2) A next problematic aspect of the study is the analysis of noradrenergic signals and HR at transitions (e.g., from NREM sleep to sleep wakefulness or to microarousals). The abstract does not mention these data, leaving open how they fit into the paper's message. There are also two problems with it: a) Noradrenaline levels increase with wakefulness, as do many other neuromodulators. This is not novel. Furthermore, why use an AUC measure for 0-25 s when MAs last only 5 or 15 s? b) biosensor signal comparisons are difficult to make for state transitions, because blood flow changes and modifies the fluorescent signal.

      We thank the reviewer for these comments and appreciate the opportunity to clarify the rationale and interpretation of these analyses.

      Regarding the inclusion of vigilance-state transitions, our intention was not to claim novelty in the observation that NE levels increase during wakefulness. Rather, these analyses were included to provide an additional physiological context in which to examine the coupling between NE and heart-rate dynamics. Specifically, the transition analyses allowed us to determine whether graded changes in NE across sleep-to-wake transitions were mirrored by corresponding RR changes (Fig. 1d–e) and whether these responses covaried (Fig. 1f), thereby strengthening the overall conclusion that cardiac dynamics track noradrenergic signaling across naturally occurring sleep-state fluctuations.

      Regarding the use of AUC measures, we deliberately chose this metric because it captures the overall physiological response over a defined time window, including both magnitude and temporal evolution, rather than relying solely on a peak value. In this context, AUC was intended to reflect differences in the overall response profile, for example that NE and RR responses during microarousals may return more rapidly toward baseline than during sustained wakefulness. To minimize the influence of differing event durations, the analysis window was kept fixed (0–25 s) across all transition types. Importantly, we directly compared AUC- and amplitude-based measures and found that they produced largely similar relationships (Suppl. Fig. 1g). The coupling between NE and RR responses remained significant when only peak amplitudes were considered, indicating that the observed relationship is not solely driven by response duration. However, the relationship was somewhat stronger with AUC-based measures, likely because heart-rate responses rapidly saturate across vigilance-state transitions, making response persistence an informative component of the physiological signal. As mentioned in the previous comment, we have added a Figure and new result text to highlight this.

      Regarding the concern about blood-flow related artifacts in fluorescent biosensor signals, we agree that hemodynamic contamination is an important consideration for neuromodulator recordings and applies broadly to biosensor-based measurements, including analyses of infraslow fluctuations. To minimize this issue, ΔF/F calculations were performed using the isosbestic control channel, which serves to correct for movement- and hemodynamic-related signal fluctuations. Furthermore, hemodynamic artifacts typically occur rapidly at state transitions, whereas GRAB-NE signals display slower dynamics. If uncorrected hemodynamic contamination strongly influenced the signal, we would expect abrupt signal distortions tightly aligned to arousal onset, which was not evident in our recordings. While we cannot completely exclude residual hemodynamic influences, we do not believe they account for the graded NE responses observed across vigilance-state transitions or the corresponding relationship with RR dynamics.

      (2.3) Wordings such as 'extent of arousal' to compare MAs and wakefulness are problematic. Microarousals and wakefulness are qualitatively different behaviorally, physiologically, and in terms of neuromodulatory conditions

      We thank the reviewer for the helpful comment. Our intention was not to imply that microarousals and wakefulness are qualitatively identical states differing only in magnitude. Rather, we used arousal in the broader neurophysiological sense, referring to the degree of activation of central arousal systems, with wakefulness representing part of this continuum. However, we recognize that in the sleep field, arousal is often used more specifically to describe brief EEG desynchronization events and that we furthermore use micro-arousals as a broad term for these sleep arousals, which may make our wording confusing. To avoid ambiguity, we have revised the manuscript to replace this terminology with vigilance state, sleep–wake state, or state transitions, depending on the context.

      (2.4) To interpret linear correlations between datasets, even for the ones with Rsquare values < 0.5, as 'predictive' represents an overextended interpretation that is not supported by available evidence. This concern is aggravated due to the use of AUCs that confound amplitudes and time courses. This is in particular the case for Figure 5d, 5k, or 5m…

      We thank the reviewer for this important comment. We agree that the term predictive may overstate the interpretation of these correlations, particularly given the modest R² values in some analyses. Our intention was to highlight an association between heartrate and NE-related measures rather than imply strong predictive performance or causality. We have therefore revised the wording throughout the manuscript to avoid predictive language and instead refer to these relationships as associations or correlations.

      Regarding the use of AUC, we selected summary metrics based on the physiological characteristics of the signal of interest. In cases where responses were characterized primarily by rapid shifts to a new level, amplitude measures were used. In contrast, AUC was chosen when both the magnitude and duration of the response were considered physiologically relevant.

      Part 3. A substantial number of experimental and analytical points require clarification. Here is a list of a few examples; many observations noted here apply equally to other figure panels.

      (3.1) Many figure panels leave it open about whether averages or representative data are shown, how many animals are included, and what kind of measures are plotted.

      We thank the reviewer for this comment and agree that clarity in figure presentation is important. In response, we have carefully revised the figure legends throughout the manuscript to more explicitly state whether data shown are representative examples or group averages, clarify the number of animals included in each analysis, and specify the measures being plotted. We have also added “data are shown as mean ± SEM” where this information was previously not explicitly stated and clarified when n refers to the number of animals. We believe the revised figure legends now provide clearer guidance for interpretation, and further statistical details are available in the accompanying statistics table. It should also be noted that number of animals used for all experiments can be found in the Method section ‘Mice’.

      (3.2) In case experiments were done in a paired manner (e.g., the LC stimulations), individual data points should be shown connected for the different conditions.

      We thank the reviewer for this suggestion. While we agree that connecting individual data points is valuable for paired experimental designs, this visualization is not appropriate for the LC stimulation analyses presented here. Specifically, each condition reflects pooled stimulation events selected across animals rather than a single summary value per animal that can be directly matched across columns. As such, individual points in one condition do not map one-to-one onto points in the next condition, making connected visualizations potentially misleading.

      Importantly, although the data are presented as pooled event-level measures, the paired structure of the experiment was accounted for in the statistical analyses, such that repeated measurements within animals and the paired nature of the design were included in the relevant comparisons.

      (3.3) Numerous analyses involve heart rate measures from the neck EMG during wakefulness. However, it is not specified how, in this case, RR peaks could be detected within the high-activity EMG.

      We thank the reviewer for this question. The procedure for heart-rate detection during wakefulness is described in detail in the Heart rate detection Methods section. Briefly, RR intervals were extracted from preprocessed neck EMG recordings using a previously validated approach for mouse sleep studies. To minimize contamination from movement-related EMG activity during wakefulness, R-peaks were not detected within periods extending 100 ms before and 250 ms after detected movement, as these segments were considered too noisy for reliable peak detection. Heart-rate estimates during these excluded periods were subsequently interpolated using surrounding valid RR intervals to preserve temporal continuity and enable analysis of HR dynamics before and after movement episodes.

      (3.4) Figure 1b: PSD for RR intervals. The supplementary figure says that 11-minutelong NREMS or 5-minute-long NREMS periods were used. These are very rare events in mice, for which the average bout duration is around 2 min and the cycle length is 10 minutes. The methods do not explain how these bouts were chosen, how many of them were included, and why shorter bouts were not analyzed. Single cases or means?

      We thank the reviewer for this important point and agree that additional clarification was warranted. The selection of long NREM (including microarousals) periods was motivated by methodological considerations related to spectral analysis of very slow oscillations rather than by an assumption that these bout lengths are representative of average NREM duration in mice. Because our analysis focused on very-low-frequency dynamics (~1 cycle every 50 s), sufficiently long continuous recordings are required to reliably estimate power at these frequencies and avoid fragmentation-related edge effects that disproportionately affect shorter bouts.

      For this reason, shorter NREM episodes were excluded from the PSD analysis, as they do not provide sufficient duration to robustly capture slow-frequency components. We initially compared PSD estimates using NREM periods of at least 11 min (allowing ~10 VLF cycles) and 5 min duration to assess whether the shorter 5 min recordings introduced bias (Supplementary Fig. 1f). While longer periods resulted in higher overall power estimates, the frequency distribution remained highly similar between conditions. We therefore selected 300 s (5 min) as the inclusion criterion for the remainder of the study, as periods exceeding 10 min are uncommon in mice and would substantially limit analyses across experimental paradigms.

      To further reduce bias related to bout duration, PSD estimates were weighted by NREM episode length, as longer bouts showed systematic effects on power estimates. For Figure 1, the analysis included 144 NREM-with-MA bouts across 7 animals, and data shown represent group means rather than single examples.

      This description is also included in the Methods section/Data Analysis.

      (3.5) Figure panel 1c,d: In Panel c, what is plotted?

      As stated in the figure legend, it is the cross-correlation between NE and R-R. We have added more description in the figure legend.

      A cross-correlation between two signals? If yes, how were these signals chosen per transition? Looks rather like they plot some time course across a transition. What is time point 0?

      We thank the reviewer for pointing out that this analysis was insufficiently explained. The analysis shown represents a cross-correlation between the NE and RR signals, performed to assess their temporal relationship across different vigilance-state transitions. Specifically, for each transition type, NE and RR signals were extracted within the corresponding time windows and cross-correlated to determine the strength and timing of their interaction.

      In this context, time point 0 (lag = 0) represents perfect temporal alignment between the two signals. Positive or negative lags indicate whether changes in one signal systematically precede or follow changes in the other. Due to methodological differences between the rapid electrical heart signal and the slower fluorescent NE signal, we intentionally avoided overinterpreting fine temporal lead–lag relationships. Rather, the aim of this analysis was to characterize the overall interaction between the two signals across transitions.

      Because the relationship between NE and RR was predominantly inverse, negative cross-correlation values indicate that increases in NE are associated with decreases in RR (i.e., faster heart rate), and vice versa. Thus, this analysis served primarily to confirm and extend our other findings by quantifying the interaction between NE and heart-rate dynamics across vigilance-state transitions.

      We have revised the descriptive sentence in the Results section to improve clarity and have added a more detailed description of the signals included directly in the figure panel.

      (Results, page 7). “Here, cross-correlation analysis revealed a predominantly negative relationship between NE and RR across vigilance-state transitions, indicating that increases in NE were associated with reductions in RR (i.e., faster heart rate; Fig. 1c).”

      (Figure 1 legend). “Cross correlation (how strongly and at what temporal offset the two signals covary) between NE and RR during transitions”.

      In Panel d, what is measured here? Is the time point of the dotted line a NA trough or the moment of a transition? Show the data with connected lines. Heart rate calculation during wakefulness?

      We thank the reviewer for these questions and apologize that this was not sufficiently clear. In panel d, the dotted line indicates the NE trough, not the moment of a vigilance state transition, as specified in both the figure and figure legend. The analysis is aligned to detected NE troughs during sleep and examines the subsequent physiological dynamics, including transitions into wakefulness.

      We have updated the sentence in the result section to make it a bit more clear:

      (Results, page 7): “To explore the NE-RR relationship across sleep-wake transitions, we examined four progressive sleep-to-wake transitions using the preceding NE trough as the time stamp (time 0):…”

      Regarding the suggestion to connect data points, as noted in an earlier response, these analyses are based on event detection, where multiple events contribute from each animal. Thus, each condition represents pooled events across animals rather than a single matched value per animal, making connected-line visualizations inappropriate and potentially misleading. Importantly, the paired structure of the experimental design was accounted for in the statistical analyses.

      Heart-rate detection during wakefulness was usually not possible due to movement artefacts in the EMG. In the detection, we excluded periods with movement. Our EMG-based detection approach is described in the Heart rate detection Methods section. Briefly, movement-contaminated periods were excluded from R-peak detection (100 ms before and 250 ms after detected movement), and RR intervals were subsequently interpolated using surrounding valid values to preserve temporal continuity. Because analyses were aligned to NE troughs occurring during sleep, we limited the temporal window to avoid excessive contamination from movement-related noise associated with subsequent wakefulness. We have slightly revised the text to make these points clearer.

      Panel C is a cross correlation between NE and RR from the same traces included in the mean traces. x=0 is the NE through like the other figures.

      Panel D is the mean traces of NE during transitions with the dotted line at x=0 being the NE though

      Yes, but x = 0 for the cross-correlation is not the NE trough. It says something about how aligned NE and R-R are in time (see response further up in (3.5)).

      (3.6) The sigma band should ideally be chosen between 10-15 Hz for better consistency with the literature.

      We thank the reviewer for this suggestion and agree that consistency in the definition of frequency bands is important for comparison across studies. The sigma range used in the present manuscript was selected to match our previous publications and analyses, thereby allowing direct comparison with our earlier findings on LC-mediated regulation of sleep spindles and infraslow sleep dynamics. Maintaining the same band definition also ensured consistency across the datasets analyzed in this study.

      We acknowledge that having a shared definition of sigma power would improve alignment and facilitate comparisons across laboratories. At present, there remains some variability in the exact frequency boundaries used for spindle and sigma analyses across studies and species, although there are ongoing efforts within the sleep field to improve standardization. Importantly, we do not expect that modest adjustments of the sigma-band boundaries would materially affect the conclusions of the present study, as the spindle-related activity of interest lies well within the selected frequency range and the observed effects are broad rather than restricted to a narrow frequency bin.

      Moving forward, we aim to follow emerging consensus recommendations where appropriate. To clarify this point for readers, we have added a statement in the Methods section explaining that the sigma band was chosen to maintain consistency with our previous publications.

      (Methods: EEG power, page 29). “The sigma band was defined as 8 - 15 Hz to maintain consistency with our previous publications. Modest differences in sigma-band boundaries are not expected to affect the main conclusions.”

      (3.7) What are 'extreme LC stimulation frequencies'. The only information available is that stimulations were done for 2s at 20 Hz.

      We thank the reviewer for pointing out that this wording was unclear. By “extreme LC stimulation frequencies”, we did not refer to the within-stimulation pulse frequency (which remained constant at 20 Hz for 2 s across all conditions). Rather, we referred to the effective frequency of LC activation at the infraslow timescale, which was progressively increased through the closed-loop stimulation paradigm.

      Specifically, stimulations were triggered when NE levels crossed increasingly permissive thresholds during the descending phase of the endogenous NE signal. As thresholds increased over time (from −15 ΔF/F (%) to +5 ΔF/F (%)), stimulations occurred progressively earlier in the infraslow cycle, thereby compressing the oscillatory period and increasing the effective frequency of LC recruitment while preserving the endogenous temporal structure of NE dynamics.

      Thus, “extreme stimulation frequencies” refers to the highest rate of repeated LC activations achieved through the closed-loop paradigm, where stimulations became increasingly frequent at the infraslow level rather than changes in the 20 Hz pulse train itself. To avoid confusion, we have revised the wording throughout the manuscript to refer more explicitly to faster infraslow LC activation frequencies or increased infraslow stimulation frequency.

      We have added more information in the result section to highlight this better.

      (Results, page 8). “We employed a closed-loop paradigm, where LC stimulations (2 s 20 Hz (10 ms) blue laser pulses with a light intensity of 5 mW) were triggered when NE levels fell below increasing thresholds (-15, -10, -5, 0 and 5 ΔF/F (%), Fig. 2a-b, Methods). This approach enabled controlled compression of the infraslow NE cycle by triggering LC activation during the descending phase of the endogenous NE signal, thereby increasing the effective oscillatory frequency while preserving the temporal structure of physiological LC–NE dynamics. This strategy allowed us to test whether heart-rate responses continue to track LC-driven NE fluctuations as the infraslow rhythm becomes progressively faster.”

      (3.8) Could the discrepancy between panels 4c, left and right, be due to limited sample size? The whole figure lacks indications of sample numbers, making interpretation difficult.

      We thank the reviewer for this comment. We assume the reviewer is referring to the apparent discrepancy between the NE response and heart-rate response following LC suppression (Fig. 4c–d), where LC inhibition induced a clear reduction in NE levels, whereas mean heart-rate responses were less pronounced.

      The figure legends state sample sizes and event numbers. Specifically, for these analyses, n = 8 animals (4 Arch, 4 YFP) were included, comprising 48 Arch events and 38 YFP events, our interpretation is that this discrepancy reflects a biological observation rather than a failed manipulation. Specifically, LC suppression robustly reduced NE levels, confirming the effectiveness of the optogenetic intervention, whereas HR did not exhibit a similarly consistent group-level response. However, as highlighted by the correlation analyses, variability in RR responses remained associated with the magnitude of NE suppression, suggesting that heart-rate dynamics still reflected noradrenergic modulation at the individual-response level despite the absence of a strong mean effect.

      (3.9) It would be great if Figure 4 j could be more explicitly illustrated. For example, behavioral traces that lead to higher NFR and corresponding changes in RR AUC should be shown for animals with large and small effect sizes. The sample size seems excessively low. Can this explain the difference in slopes compared to Figure 4e, right panel?

      We thank the reviewer for this suggestion. To improve the interpretation of Figure 4j, we have now added representative examples illustrating animals with high and low memory performance and their respective RR responses following LC suppression.

      We agree that the sample size for this analysis is limited. As noted in the Methods, Figure 4j is based on a secondary analysis of a previously published dataset (Kjaerby, Andersen et al., 2022), where heart-rate measures were retrospectively extracted from EMG recordings. Due to noise-related limitations in RR detection, reliable cardiac measures could not be obtained from all animals, reducing the number of subjects available for this analysis. For this reason, we have deliberately avoided direct statistical comparisons between Arch and YFP animals and instead limited our conclusions to the observed association between RR responses and memory performance across animals.

      Regarding the difference in slope compared with Figure 4e, the two analyses are based on different levels of aggregation and address different questions. Figure 4e examines the relationship between NE and RR responses across individual LC suppression events, resulting in multiple observations per animal. In contrast, Figure 4j uses a single mean RR response and a single behavioral outcome per animal. Furthermore, Arch and YFP animals were pooled in Figure 4j to maximize statistical power and because the dataset was not sufficiently powered for direct group comparisons. Consequently, the slopes are not expected to be directly comparable between the two figures.

      (3.10) It would be important to show anatomical validation of viral expression in THcre animals and optic fiber positioning.

      We thank the reviewer for this comment. Anatomical validation of viral expression in TH-Cre animals and optic fibre placement was performed for these experiments and has been reported previously in Kjaerby et al. Nature Neuroscience paper, from which this dataset was derived. Specifically, viral targeting and fibre positioning were histologically verified as part of the original experimental validation. This is mentioned in the Method section ‘Surgery’: Viral expression and injection sites were validated through immunostaining of perfused brain slices from the experimental animals (see Kjaerby et al. (3) for more information).

    1. eLife Assessment

      This important study reveals distinct representations of task-related information in the dendrites and somata of cortical neurons during sensorimotor learning and behavioral adaptation. The evidence is compelling, combining imaging of dendritic and somatic activity during behavior to demonstrate compartment-specific encoding of sensory cues, motor actions, and corrective signals. The work will be of broad interest to neuroscientists studying dendritic computation, motor learning, and the cellular mechanisms underlying adaptive behavior.

    2. Reviewer #1 (Public review):

      [Editors' note: this version has been assessed by the Reviewing Editor without further input from the original reviewers. The authors have addressed the comments raised in the previous round of review.]

      Summary:

      In this manuscript, Scheib et al. identify distinct calcium dynamics in the somata and tuft dendrites of layer 5 pyramidal cells in mice performing a licking task. Animals are trained to lick water ports on the left or right following an acoustic cue, and can adjust their targeting when the ports are displaced. For tongue premotor cortical neurons projecting to the ventromedial thalamus, calcium transients in tuft dendrites are tightly locked to the direction-instructive cue, while somatic calcium signals are more broadly dispersed and more frequently synchronized with tongue motion and port contact. Finally, when the targets are shifted, tufts exhibit a sparse but large corrective signal on an improperly-targeted first lick, and the changes in population activity in the tufts and somata differ after adaptation to the new port locations.

      Strengths:

      In my opinion, this is a very strong manuscript which reports several novel and significant observations, contains high-quality data and (for the most part) reasonable analyses, and is clear and well-written. Most prior studies of cortical sensorimotor processing have measured the output of neurons using extracellular recording - an approach which obscures potentially important signaling differences between neuronal compartments. This study leverages cutting-edge imaging techniques in mice to document large, time-dependent differences between calcium signals at cortical somata and tuft dendrites. This phenomenon could have major implications at the cellular level for synaptic plasticity, and at the systems and behavioral levels for motor adaptation.

      Weaknesses:

      At a conceptual level, the authors may wish to elaborate a bit on what sensorimotor computation they think the circuit is implementing, and how their results help explain this implementation. Several possibilities are raised: tuft activation could "prime" the pyramidal cells in advance of movement initiation (line 319ff), or could track errors to engage plasticity (line 351ff) and solve the credit assignment problem (line 362ff). It might be helpful to make one of these proposals more concrete with a computational model, but this is not strictly necessary. [The authors explain that they will address this with modeling work in subsequent research.]

    3. Reviewer #2 (Public review):

      Summary:

      The authors set out to compare functional encoding in the tuft dendrites and somata of a specific cortical cell type during motor planning and learning.

      Strengths:

      The investigation of a specific projection type (L5 ET) is a strength that aids reproducibility and interpretation. The elegant approach to increasing the depth of field of dendritic imaging is another strength. The data analyses are largely clear in their methods, scope, and interpretation. The writing is extremely clear and appropriately referenced, with an excellent Introduction, in particular.

      Weaknesses:

      This work is largely observational, describing signals that might reflect computational transformations and/or instruct plasticity, but those possibilities have not yet been deeply investigated. The manuscript does a good job of laying out these as future directions.

    4. Reviewer #3 (Public review):

      Summary:

      This article by Scheib et al. investigates how layer 5 extratelencephalic (ET) neurons in the frontal cortex encode sensorimotor information during motor learning, focusing on differences between their apical tuft dendrites and somas. The authors alternated recordings among these ET neuronal compartments in the mouse anterior lateral motor cortex (ALM) during a cued directional licking task with a target port shift. They found that while tuft dendrites predominantly encode sensory cues, with a subset selectively active during corrective actions, somatic activity was more strongly associated with action timing. Additionally, learning induced divergent plasticity: tuft dendrites increased their selectivity but decreased response gain, maintaining stable net selectivity, whereas somas showed increased net selectivity early in learning. Together, these findings reveal distinct sensorimotor representations and learning-related plasticity in dendritic and somatic compartments, providing insight into how compartment-specific activity in the frontal cortex may contribute to motor skill acquisition.

      Strengths:

      The authors developed an innovative imaging approach and a comprehensive data analysis pipeline to address a knowledge gap in the literature. By alternating imaging of dendritic tufts and somas in the same animals, they compare compartment-specific activity during motor learning and identify distinct encoding of task variables and learning-related plasticity across these compartments. Interestingly, a subset of dendritic tufts shows activity associated with corrective actions. The findings are discussed in the context of current theories of dendritic computation, credit assignment, and motor learning, providing a useful foundation for future mechanistic studies.

      Weaknesses:

      No major weaknesses were identified.

    5. Author Response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      This important study reveals distinct representations of task-related information in the dendrites and somata of cortical neurons during sensorimotor learning and behavioral adaptation. The evidence is compelling, combining simultaneous imaging of dendritic and somatic activity during behavior to demonstrate compartment-specific encoding of sensory cues, motor actions, and corrective signals. The work will be of broad interest to neuroscientists studying dendritic computation, motor learning, and the cellular mechanisms underlying adaptive behavior.

      Thank you for this excellent summary. We recommend one change: removing the word “simultaneous”. It could perhaps be replaced with “concurrent” or simply omitted. Tuft dendrites and somata were imaged on alternating days, and most readers will probably interpret “simultaneous” as implying a faster, interleaved sampling rate.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      In this manuscript, Scheib et al. identify distinct calcium dynamics in the somata and tuft dendrites of layer 5 pyramidal cells in mice performing a licking task. Animals are trained to lick water ports on the left or right following an acoustic cue, and can adjust their targeting when the ports are displaced. For tongue premotor cortical neurons projecting to the ventromedial thalamus, calcium transients in tuft dendrites are tightly locked to the direction-instructive cue, while somatic calcium signals are more broadly dispersed and more frequently synchronized with tongue motion and port contact. Finally, when the targets are shifted, tufts exhibit a sparse but large corrective signal on an improperly-targeted first lick, and the changes in population activity in the tufts and somata differ after adaptation to the new port locations.

      Strengths:

      (1.1) In my opinion, this is a very strong manuscript which reports several novel and significant observations, contains high-quality data and (for the most part) reasonable analyses, and is clear and well-written. Most prior studies of cortical sensorimotor processing have measured the output of neurons using extracellular recording - an approach which obscures potentially important signaling differences between neuronal compartments. This study leverages cutting-edge imaging techniques in mice to document large, time-dependent differences between calcium signals at cortical somata and tuft dendrites. This phenomenon could have major implications at the cellular level for synaptic plasticity, and at the systems and behavioral levels for motor adaptation. As described below, I have only one major technical concern (which should be addressable with additional analysis), along with several relatively minor suggestions for improving the manuscript.

      We thank the reviewer for their insightful summary of the significance of the differences that we identified in the task-related activity of tuft dendrites and somata.

      Weaknesses:

      (1.2) At a conceptual level, the authors may wish to elaborate a bit on what sensorimotor computation they think the circuit is implementing, and how their results help explain this implementation. Several possibilities are raised: tuft activation could "prime" the pyramidal cells in advance of movement initiation (line 319ff), or could track errors to engage plasticity (line 351ff) and solve the credit assignment problem (line 362ff). It might be helpful to make one of these proposals more concrete with a computational model, but this is not strictly necessary.

      We thank the reviewer for this feedback. We absolutely agree that detailed computational models of each proposed computation will be very valuable and constitute an important follow-up to this work. We hope to collaborate with theorists to take that next step. Each possible computation noted by the reviewer reflects distinct differences that we observed in the task-related activity of tuft dendrites and somata. They are not mutually exclusive hypotheses to explain the same phenomenon. As such, we think they are best addressed independently in future modeling work. By making the data and a concise description of the main findings available immediately, we hope to allow computational experts in each of these areas to take advantage of the results of this study without delay.

      (1.3) My only major technical concern relates to the analyses in Figures 4F-H, 5G-I, and 6H-K (c.f. equations 2-5). Typically, one identifies population-level factors by projecting neural activity onto fixed dimensions of interest; this makes it possible to see how activity evolves over time along interpretable coordinates. Here, however, the coding directions are redefined at each time point, so the "choice" activity at time t is actually a different signal from the "choice" activity at t+1. This procedure is a bit like comparing the activity of one neuron at one time point with the activity of a different neuron at a later time point. It also makes the physiological interpretation more complicated: if the dimensions are fixed, one can see how a downstream neuron could "read out" the signal by computing a weighted sum of the activity of upstream neurons, but it is harder to see how this could happen if the weights are always rotating.

      We thank the reviewer for raising this point. We agree that our use of projections along coding directions (CDs) defined at each time point is a less conventional use of coding directions, although nearly identical calculations have been previously used to assess population-level selectivity and code stability in this task (Chen et al., 2017; Yang et al., 2022). As noted in the article, given low numbers of error trials and high trial-to-trial variability, we found that estimating the selectivity of individual ROIs for these task-dimensions was not robust and was subject to overfitting. Cross-validated projections at each timepoint provided a far more robust measure of population selectivity. Furthermore, we were able to orthogonalize stimulus, choice and outcome CDs to better identify distinct encoding of each task-variable. Finally, because the primary goal of the study was to identify any differences between tuft dendrite and somatic encoding, we think that calculating the population selectivity at each timepoint gives readers a less biased view of the selectivity of the two compartments, whereas calculating a CD over a single arbitrary time window could conflate differences in dynamics with differences in selectivity.

      We agree that calculating the CD at each timepoint makes it hard to see where the code is stable and where it is rotating, and thus how a downstream neuron might “read out” the signal. To provide this information, we have added new panels to the supplement showing the correlation of CDs across time (Figure 4 - figure supplement 1B,D). We also now provide this information for CR-CA in Figure 5—figure supplement 2A (the plots previously presented in 2A were the correlations of CR with CA, rather than CR-CA; an error that has been fixed). The following changes were also made to the Results section to clarify this issue:

      “From the linear model, we calculated coding directions (CDs) at each timepoint that maximally separated Stimulus, Choice, and Outcome activity (Figure 4F; Figure 4—figure supplement 1A) and estimated the direction and selectivity along each dimension across time (Figure 4—figure supplement 1B-E; see Methods). Allowing CDs to rotate in time (see Figure 4—figure supplement 1B,D), although unconventional, ensured that comparisons of population selectivity across the two compartments were not biased by the selection of an arbitrary CD time window.”

      We also identified a mistake in the description of CD orthogonalization in the Methods, which has been corrected as follows:

      “For each timepoint, each selectivity CD was then orthogonalized with respect to the other two selectivity CDs by a QR decomposition in which that selectivity CD was last in the order.”

      (1.4) A few comments on the behavioral task and results. After the port shift, the error rate is quite high, and doesn't diminish much between the early and late epochs (approximately 42% and 38% error rate, respectively; Figure 1I). That is, mice do not seem to fully master the task. Clearly, animals do alter their aim, but even this does not seem to change much between early and late periods (Figure 1J). I recommend that the authors show the behavioral data at a finer level of granularity (e.g., by plotting the change in exit trajectory on all individual trials across sessions, with a loess fit) to allow an assessment of the adaptation rate and when adaptation saturates. It would also be more conventional to refer to the behavioral changes as "motor adaptation," instead of "skill learning." (The latter would be appropriate if the port offset were randomized across trials, and animals received two separate cues for direction and offset, but I suspect this task would be too difficult for mice to learn.)

      We agree with the reviewer that by the end of the late period, performance on the right side (Figure 1I) has still not returned to pre-shift levels. This may reflect mice not fully mastering the task, as the reviewer suggests, or it may reflect that after the shift, the right port is substantially more difficult to reach than the left port. Unfortunately, because of high animal-to-animal and lick-to-lick variability, plotting the post-shift lick angle at a finer level of granularity is not statistically informative.

      With regard to the nature of the learning in our task, we selected “skill learning” as the best description of the motor learning task based on distinctions between adaptation and the learning of motor skills by Krakauer et al., 2019 and Heald et al., 2021. Conceptually, the difference is whether an existing motor controller memory is simply updated with new parameters, or whether the motor context has changed sufficiently that a distinct motor controller memory (which can still use parts of previous memories) is formed. In our task, after the port shift the left port forms an obstacle to reaching the right port. This obstacle was simply not present before the shift. Before the shift, ports were approximately equidistant from the mouth and easily avoided given the port separation and tongue width. Thus, avoiding an obstacle would presumably not be part of the initial motor controller memory and a distinct memory would need to be constructed.

      We agree with the reviewer, however, that given that we do not have fine-timescale dynamics of behavioral changes in response to the shift, and did not conduct other experiments (such as returning the ports to their original location) that would typically be conducted to identify “adaptation-like” or “skill learning-like” dynamics, we cannot empirically distinguish between the two. We now clarify in “Study limitations” that we call the studied behavior “skill learning” based on the nature of the task, but that our behavioral analysis cannot distinguish between adaptation and skill learning:

      “We refer to the behavioral paradigm as motor “skill learning” strictly based on the nature of the task. After the shift, mice must avoid a new obstacle close to the mouth (i.e., the left port), which we assume requires the formation of a distinct motor controller memory and therefore would be considered skill learning (Krakauer et al., 2019). However, we did not confirm that the mice exhibited specific behavioral characteristics of skill learning and it is possible that other kinds of motor learning (e.g., motor adaptation) were dominant.”

      (1.5) This is perhaps a semantic point, but it might not be entirely accurate to refer to the activity evoked by the directional cue as "sensory." Typically, a "sensory" response should encode some feature of a stimulus - in this case, the frequency of a tone. Here, it seems likely that the cue-aligned activity reflects the instructed lick direction, rather than the auditory information per se. (Presumably, these premotor neurons do not have well-behaved auditory tuning curves.) By comparison, in macaques performing center-out reach tasks, activity in dorsal premotor cortex rapidly ramps up following a visual cue instructing the direction of an upcoming reach, but one usually wouldn't refer to this activity as "visual" or "sensory" (though this is sometimes done). I suggest the authors either use "Instruction" or similar (e.g., in Figure 4F), or clarify in the text whether they think the activity is a genuine auditory response or something else.

      We understand how this could cause confusion. “Sensory” was meant to denote the nature of the differences in external events between the trial types used to calculate selectivity, not to imply that the activity was necessarily selective for detailed features of the cues outside the context of the task. Previous work in ALM cortex has labeled this selectivity direction as “stimulus” (Yang et al., 2022; Chen et al., 2024) to better emphasize that it is simply defined by the external cue. Where appropriate, we have revised the article to use “stimulus” or “instructional cues” in place of “sensory” for clarity and to better conform with convention.

      Reviewer #2 (Public review):

      Summary:

      The authors set out to compare functional encoding in the tuft dendrites and somata of a specific cortical cell type during motor planning and learning.

      Strengths:

      (2.1) The investigation of a specific projection type (L5 ET) is a strength that aids reproducibility and interpretation. The elegant approach to increasing the depth of field of dendritic imaging is another strength. The data analyses are largely clear in their methods, scope, and interpretation. The writing is extremely clear and appropriately referenced, with an excellent Introduction, in particular.

      We thank the reviewer for their appreciation of the study design, imaging methods, and scholarship of the article.

      Weaknesses:

      (2.2) It is not obvious whether the selected labeling strategy avoids labeling Layer 6 CT neurons, which would contaminate dendritic recordings. The images provided suggest enrichment in L5, but a discussion of this important potential caveat is warranted, especially since within-cell comparisons of apical dendrites to somata were not performed.

      We thank the reviewer for emphasizing the need to discuss this potential issue. For the following reasons, it is likely that the vast majority of dendrites we imaged in layer 1 originated from layer 5 ET neurons. First, as the reviewer notes, the provided images suggest enrichment in layer 5. This enrichment likely reflects the fact that most L6 CT neurons in motor and premotor cortex send denser projections to other thalamic nuclei than to VM thalamus (Winnebust et al., 2019, Cell), where we targeted our retrograde-Cre injections. Second, L6 CT neurons are predominantly untufted (Ledergerber and Larkum, 2010, J. Neurosci.), including in motor and premotor cortex (Peng et al., 2021, Nature; Ichikawa, 2025, Front. Neuroanat.). A recently identified subclass of L6 CT neurons in secondary motor cortex has dense projections to VM thalamus, but this class also appears to extend minimal dendrites into L1 (Li et al., 2024, bioRxiv). Nonetheless, we did not label post-hoc tissue collected from imaged mice with markers of precise laminar boundaries, and thus cannot definitively rule out the possibility that dendrites from a subclass of L6 CT neurons with tuft dendrites were also imaged. We have added the following paragraph to the “Study limitations” section to make readers aware of these issues:

      “L5 ET neurons in premotor cortex elaborate extensive tuft dendrites in L1, whereas Layer 6 (L6) corticothalamic (CT) neurons are predominantly untufted (Jiang et al., 2020; Peng et al., 2021). Thus, although we cannot rule out the possibility that dendrites from a subclass of L6 CT neurons were also sampled, it is likely that the vast majority of dendrites we recorded in L1 originated from L5 ET neurons.”

      (2.3) The application of DeepInterpolation to dendritic data appears to be novel, and little detail or vetting is provided. The reader is left guessing: Was the model retrained or fine-tuned on dendritic data? How does the denoising affect the resulting segmentation and activity traces? Is denoising necessary for this workflow?

      We thank the reviewer for requesting this useful additional information.

      In all cases, the model was retrained for each dendritic or somatic imaging session. Denoising improved segmentation consistency, as measured by comparing segmentations of individual sessions from the same animal. This is now specified in the Methods as follows:

      “The DeepInterpolation model was trained on each imaging session prior to denoising of that session. Denoising prior to NMF-based segmentation resulted in more robust and consistent dendrite segmentation than NMF-based segmentation without prior denoising (0.79 +/- 0.01 ⍴ vs. 0.46 +/- 0.01 ⍴; mean of the max Spearman correlation of components across sessions; random subsample of N = 3 mice, 15 sessions, 400 components).”

      With regard to how denoising impacts activity traces, examples were shown in Figure 2I, K. To provide more quantitative information to the reader, we calculated estimates of the power and reliability of the spectral content of dendrite activity traces extracted with or without denoising. These data are now shown in the new panel, Figure 2 - figure supplement 2G. The power spectral density of the denoised activity and the estimated reliable power spectral density of the raw traces match up to approximately 2.6 Hz (Figure 2 - figure supplement 2G), which is not far from the bandwidth of GCaMP8m, given its estimated combined rise and decay (Figure 2 - figure supplement 3B, C). Some frequencies beyond this point have been suppressed beyond what would be expected due to photon shot noise (as estimated by the replicate coherence-weighted PSD, or “recoverable” PSD). Further characterization of the precise nature of the suppressed high-frequency information – which could be suppressed artifacts (e.g., fast brain motion) or lost signal detail (i.e., GCaMP8m rise kinetics) – is beyond the scope of this paper.

      Details of the PSD calculations have been added to the Methods, and the following statement has been added to the Results: “Power spectral density of the denoised traces and the coherence-weighted power spectral density of the raw traces match up to approximately 2.6 Hz (Figure 2 - figure supplement 2G; Methods), which is not far from the bandwidth of GCaMP8m, given its estimated combined rise and decay (Figure 2 - figure supplement 3B, C).”

      (2.4) The activity patterns of the recorded cells appear to lack the characteristic ramping during the delay epoch previously reported in both calcium imaging and electrophysiology studies. Given that a major contribution to the significance of the work is to constrain models of ALM function, a discussion of how the data aligns with previous measurements in the same circuit would improve the work.

      Preparatory selectivity and ramping activity can be seen in Figure 3H, Figure 6I, and Figure 5 – figure supplement 1B. We note that in ALM cortex, the ramping mode explains a minority of the total variance (~17%, Yang et al., 2022), but it can appear particularly prominent in projections along certain fixed CDs.

      (2.5) It would be very informative to compare differences in signals between dendrites and somata of the same cells. Consistently tracing dendrites to their respective somata would assuage worries of potential contamination from dendrites of deeper cells and enable more direct comparisons of signal transformations between dendrites and somata. It would be good to understand the relationship between dendritic calcium signals and backpropagating action potentials in this task. The authors detect less frequent calcium events in tufts versus somata; is this due to selective backpropagation of action potentials? The dynamics of this process were recently investigated by Adam Cohen's group in vivo and in vitro, and measurements in the present settings could be compared to such work.

      We agree with the reviewer that being able to compare differences in signals between the dendrites and somata of the same cells would be very valuable. However, reliable tracing of tuft dendrites to somata from in vivo 2P anatomical imaging requires extremely sparse labeling, such that very few neurons are recorded per animal (Kerlin et al., 2019, eLife; Otor et al., 2022, Science). As stated in the “Study limitations” section of the Discussion, we suspected (correctly) that some task-related selectivity (i.e., selectivity for corrective action) would be sparsely represented in the dendrites, and thus adopted a labeling and image processing strategy that allowed us to record from many dendrites per animal. This strategy necessarily comes at the expense of generating a labeling density that precludes reliable tracing of tuft dendrites to their respective somata based on 2P morphology alone. As discussed in our response to reviewer comment 2.2 and a new paragraph of “Study limitations,” substantial contamination of the dendrite recordings by dendrites of L6 CT neurons is highly unlikely. Future studies could use simultaneous functional imaging across large volumes combined with activity-based segmentation or post-hoc high-resolution imaging of tissue sections registered to in vivo 2P imaging to accomplish both high-throughput dendritic imaging and reliable tracing.

      We thank the reviewer for pointing out that we could discuss selective backpropagation as a potential mechanism more explicitly. Our results are consistent with previous studies of L5 tufts in vivo (Francioni et al., 2019, eLife), including in ALM cortex (Maristany de las Casas et al., 2026, Science), that reported that rates of multi-branch calcium transients in the tuft dendrites of L5 neurons are lower than somatic spike rates. As discussed in “Study limitations,” there is not a clear approach in our data to determine the precise nature of the events underlying the calcium transients we measured in the tuft dendrites. Selective backpropagation of action potentials is certainly one possibility and we agree that recent research from Dr. Adam Cohen’s group should be discussed. We have added the following to the Discussion:

      “Based on previous calcium imaging of L5 tufts in ALM cortex of mice engaged in similar tasks (Kerlin et al., 2019; Maristany De Las Casas et al., 2026), we suspect that most of the activity we measured was coincident with global tuft or hemi-tree events, as well as somatic spiking. Recent in vivo voltage imaging in the hippocampus has also indicated that most spikes in distal dendrites start as bAPs that have been selectively amplified (Wu et al., 2026; Lee et al., 2026).”

      (2.6) The Coding Direction analyses presented in this work, while consistent with previous literature on population codes in ALM, are at odds with the nature of the measurements here. The changes in representation that occur between the dendrites and soma of an individual cell are probably best thought of in terms of the dynamics of signals themselves within individual neurons, rather than in the information encoded across a population.

      We thank the reviewer for giving us the opportunity to clarify this issue. As noted in the article, given low numbers of error trials and high trial-to-trial variability, we found that estimating the selectivity of individual ROIs for these task dimensions was not robust and was subject to overfitting. Cross-validated projections at each timepoint provided a far more robust measure of population selectivity. Furthermore, we were able to orthogonalize stimulus, choice and outcome CDs to better identify distinct encoding of each task variable in the population activity. Thus, the analyses are not at odds with the nature of the measurements in the study.

      Nevertheless, it is true that by recalculating the CD at each timepoint, our selectivity projections do not provide the same information as conventional projections along a fixed CD, which can indicate where the selectivity code is stable and where it is changing. To provide this information we have added new panels to the supplement showing the correlation of selectivity CDs across time (Figure 4 - figure supplement 1B, D).

      (2.7) This work is largely observational, describing signals that might reflect computational transformations and/or instruct plasticity, but those possibilities have not yet been deeply investigated. The manuscript does a good job of laying out these as future directions.

      We agree with the reviewer. As noted by the reviewer in comment (2.1), we combined a number of approaches in an innovative manner to explore how tuft dendrite activity differs from somatic activity at the population level during motor learning. These measurements provide the necessary foundation for future mechanistic studies and we think it is appropriate to share them at this stage of investigation and in the format of this article.

      Reviewer #3 (Public review):

      Summary:

      This article by Scheib et al. investigates how layer 5 extratelencephalic (ET) neurons in the frontal cortex encode sensorimotor information during motor learning, focusing on differences between their apical tuft dendrites and somas. The authors alternated recordings among these ET neuronal compartments in the mouse anterior lateral motor cortex (ALM) during a cued directional licking task with a target port shift. They found that while tuft dendrites predominantly encode sensory cues, with a subset selectively active during corrective actions, somatic activity was more strongly associated with action timing. Additionally, learning induced divergent plasticity: tuft dendrites increased their selectivity but decreased response gain, maintaining stable net selectivity, whereas somas showed increased net selectivity early in learning. Together, these findings reveal distinct sensorimotor representations and learning-related plasticity in dendritic and somatic compartments, providing insight into how compartment-specific activity in the frontal cortex may contribute to motor skill acquisition.

      Strengths:

      The authors developed an innovative imaging approach and a comprehensive data analysis pipeline to address a knowledge gap in the literature. By alternating imaging of dendritic tufts and somas in the same animals, they compare compartment-specific activity during motor learning and identify distinct encoding of task variables and learning-related plasticity across these compartments. Interestingly, a subset of dendritic tufts shows activity associated with corrective actions. The findings are discussed in the context of current theories of dendritic computation, credit assignment, and motor learning, providing a useful foundation for future mechanistic studies.

      We thank the reviewer for highlighting interesting findings in the paper and their assessment that it provides a “useful foundation for future mechanistic studies”.

      Weaknesses:

      No major weaknesses were identified.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      A very minor suggestion: it would be useful to mention the model organism in the abstract or title.

      (1.6) Thank you for catching this. We have added the model organism to the abstract as follows:

      “Using longitudinal two-photon calcium imaging, we investigated sensorimotor encoding in the apical tuft dendrites and somata of L5 extratelencephalic (ET) neurons in the frontal cortex of mice during learning of a discrete change to a cued dexterous action.”

      Reviewer #3 (Recommendations for the authors):

      Major:

      (3.1) Lines 197-199: It is unclear why the authors conclude that somas have stronger representations of choice and task outcome. In Figure 4G, there is no significant difference between dendrites and somas for Choice or Outcome coding selectivity. The differences in the Sensory/Choice and Sensory/Outcome ratios shown in Figure 4H,I are likely explained by stronger Sensory selectivity in dendrites (Fig 4g), rather than by stronger Choice or Outcome encoding in somas.

      We agree with the reviewer’s interpretation of the data. The statement at 197 - 199 was meant to reflect relative selectivity, but it was imprecise. We have replaced that sentence with the following, more precise sentence:

      “Somatic activity also encoded these features, but the representation of the stimulus was weaker – and the representation of action timing was stronger – than in the tuft dendrites.”

      (3.2) Figure 4B: The authors realign FL-associated IRFs to GO-cue timing using the mean FL latency for each trial type and animal. Because FL timing is jittered across trials and may differ between CL and CR trials, this could smear the realigned traces and complicate the interpretation of contact-associated activity. The authors should consider using trial-by-trial FL timing for realignment or quantify the impact of FL-timing variability on the resulting traces.

      We aligned average GO- and contact-IRFs in Figure 4B so that comparisons of their magnitudes could be drawn from the same time window.

      With regard to jitter across trials, we think the reviewer may have misinterpreted how the mean IRFs in Figure 4B are calculated. The contact-IRF, by its nature, is calculated once per animal and trial type with respect to FL timing and shifted once based on mean FL latency. There is no smearing due to trial-to-trial FL timing.

      With regard to systematic differences in FL timing across animals and CR vs. CL, the reviewer is correct that this could – in theory – smear the realigned mean contact-IRF shown in Figure 4B. However, differences in mean FL latency across animals and trial-types are small compared with the long-timescale contact-IRFs. Thus, the non-realigned (i.e., always FL-aligned) mean contact-IRF looks nearly identical to Figure 4B just globally offset in time, as shown in Author response image 1:

      Author response image 1.

      Since this is nearly identical to data already presented in Figure 4B, we do not think it is necessary to include it in the revised article. However, we have added the following to the Methods:

      “Population averages of contact-IRFs that were not shifted prior to averaging were nearly identical (excluding the overall temporal shift; data not shown), indicating that pooling of mean IRFs across animals and trial types produces minimal smearing of the final population IRF.”

      (3.3) Figure 5A: Are CA trials specific to motor learning, or do they reflect a corrective lick toward the alternative port after an unrewarded lick? An analysis of the second lick on left-error trials or pre-shift right-error trials could help distinguish whether correction licking reflects a general decision change after failed reward, or a motor-command correction specific to post-shift motor learning. The authors should also report the prevalence of CA versus AP trials and clarify whether these trial types are behaviorally distinct.

      We thank the reviewer for highlighting the need to emphasize that CA trials reflect a distinct behavior related to reaching the displaced port.

      By definition, CA trials started as Motor Error trials and thus reflected a corrective lick toward the same port after an unrewarded lick. Almost all first contact licks on Motor Error trials were well outside the distribution of correct left licks both pre- and post-shift (Figure 1 - figure supplement 1B,D), consistent with the interpretation of this first lick as directed toward the right port. Thus, we see no evidence suggesting that CA trials involve a decision change. CA trials are exceedingly rare pre-shift, because Motor Error trials are rare pre-shift (Figure 1I, only ~5% of all right trials).

      With regard to other error types before the shift, most expert-trained mice did not immediately sample the other port with a second lick after an unrewarded lick. They usually either stopped licking immediately or licked the unrewarded port multiple times before switching ports. When port switches occurred pre-shift, timing was highly variable across mice and trials. Even on rewarded trials, some mice would “check” the unrewarded port after consuming the reward, as can be seen in Figure 3I, J. All of these behaviors are clearly distinct from the stereotyped second lick that occurred on CA trials after the shift. We agree that the prevalence of CA and AP trials, as well as the prevalence of immediate port alternation, should be reported, and we have added that information to the article as follows:

      “On Correction Attempted (CA) trials, the first lick made contact with the incorrect port, and the mouse chose to direct a second lick toward the correct port (Figure 5A; prevalence: 54% of motor error trials). We interpreted these licks as a corrective action, because the tongue exit angle shifted further toward the correct target (Figure 5B). Abandoned Port (AP) trials were the same as CA trials, except the mouse either did not make a second attempt or the second lick was directed toward the incorrect port (Figure 5A; prevalence: 46% of motor error trials).”

      (3.4) The classification of pre-shift errors into motor and decision errors is not clear. If error-trial exit angles follow a unimodal distribution (Figure 1- Figure Supplement 1C), then the distinction between motor and decision errors may not be behaviorally well separated. The authors should explain how these categories are validated and whether conclusions depending on this classification are robust to alternative definitions.

      We do not conclude that motor errors and decision errors are distinguishable pre-shift. Pre-shift licks were classified into motor error and decision error categories only to demonstrate that the boundary we established for classifying post-shift licks classifies extremely few (~5%, Figure 1I) pre-shift licks as motor errors. No conclusions were drawn from comparisons between pre-shift licks classified as decision errors and those classified as motor errors. The categorization is defined by the distribution of exit angles pre-shift and validated by the bimodal distribution of exit angles on error trials post-shift. To improve clarity regarding our classification of pre-shift errors, we have added the following to the Results:

      “Exit angles after the shift exhibited a bimodal distribution across error trials (Figure 1G,H; Figure 1—figure supplement 1C,D), supporting this distinction in error type. The frequency of licks classified as motor errors on right-cued trials increased significantly after the shift (median pre-shift 0.06, median post-shift 0.42, p < 0.001; Figure 1I; Figure 1—figure supplement 1C,D), reflecting the new challenge of avoiding the left lickport. In contrast to after the shift, exit angles on error trials before the shift were unimodal (Figure 1—figure supplement 1C). These errors were classified based on the fixed CB in order to demonstrate that very few pre-shift licks qualify as motor errors (Figure 1I), and not to suggest that tongue trajectories before the shift are behaviorally well-separated.”

      (3.5) Figure 1- Figure Supplement 1D, post-shift decision errors: Are these truly decision errors? The lick angles appear similar to those observed before the shift, suggesting that these trials may reflect execution of a "default" or "uncertain" lick trajectory rather than an incorrect choice under the new contingency.

      The post-shift exit angles on right-cued decision error trials (Figure1 - figure supplement 1D, grey) are similar to the lick angles on correct left-cued trials pre-shift (Figure 1 - figure supplement 1A, red) and clearly different from the correct right-cued trials pre-shift (Figure 1 figure supplement 1A, blue). Thus, to the extent that the animal’s intention can be measured from lick trajectory, it was targeting the incorrect (left) port. It is also true that it may still target the previous location of the left port (a “default” left trajectory), but because the decision error makes precise targeting irrelevant to the task outcome (it is easy to reach the left port after the shift), we do not designate it as a joint decision error and motor error. As to whether the deliberative process leading to this action is somehow cognitively distinct from other behaviors typically labeled as decision errors or incorrect choices, we cannot say.

      Minor:

      (3.6) Vocabulary consistency: soma vs somata.

      When data are shown for, or derived from, multiple somata, we use “somata”. When data are shown for an individual soma (such as in a panel with data from a single example soma), we use “soma.” We could not find any use of “somas,” which would indeed be inconsistent.

      (3.7) Figure 1C: I am not sure why the lick trajectories do not depict the tongue exiting the mouse. What time window is shown? Why does it look like the trajectories are shifted to the left?

      We thank the reviewer for identifying this issue. The definition of the location labeled “mouth” was accidentally omitted. The lick trajectories in Figure 1C do depict the tongue tip once it became visible to the cameras. Jaw opening and shifting partly determined the location where the tongue became visible in the videography. These movements varied from mouse to mouse and trial to trial, so exit angle was measured from the approximate midpoint between the temporomandibular joints, which is the grey point in 1C. We have fixed the captions and Methods to precisely define this location. With regard to the appearance of a slight leftward shift in the trajectories, this reflects how the tongue exits the mouth and how the tongue tip curves downward as the tongue approaches the port.

      (3.8) Figure 1- Figure supplement 1: it could ease the comparisons to report population statistics, such as median, from panel A to panel B and D, population statistics from B to D.

      Thank you. We have added these statistics to the Figure 1 - figure supplement 1 caption.

      (3.9) Choice boundary (CB) should be defined in line 100, not 110.

      Thank you. We have fixed this.

      (3.10) Line 109: claim not supported by referenced figure (Figure 1 - Figure Supplement 1). Lick angle histogram to the right port, pre-shift does not overlap substantially with lick angle to the left port, post-shift.

      We thank the reviewer for the opportunity to clarify this. We agree that Figure 1 - figure supplement 1 is not sufficient to support the claim. First, we want to make clear that Figure 1 - figure supplement 1 does not contradict the claim. The new location of the left port can obstruct the tongue during right-cued licks, regardless of the distributions of left licks pre- or post-shift. Second, to confirm that the new location of the left port would obstruct a substantial fraction of pre-shift right-cued lick trajectories, we measured the minimum distance between tongue trajectories and the post-shift location of the left port. Of pre-shift right-cued exit trajectories, 30 +/- 5% came within 1.25 mm – half of the combined tongue width (1.5 mm) and port width (1 mm) – of the port center.

      To make this claim more precise, we have changed the statement as follows:

      “Thus, on right-cued trials, mice continuing to follow the pre-shift motor plan would be biased to more frequently contact the new left port location (Figure 1E,F; 30 +/- 5% of pre-shift trajectories came within a tongue-width of the new location) and receive punishment (i.e., timeout).

      (3.11) Line 113: claim not supported by referenced figure. Figure 1G does not display error trials.

      We have changed the line to refer to “both correct and error trials”, such that reference to Figure 1G is also appropriate.

      (3.12) Figure 2 - Figure Supplementary 3 & method: how is noise estimated?

      Thank you. The following has been added to the Methods:

      “For Figure 2 - figure supplement 3, noise was estimated as the square-root of the geometric mean of the Welch power spectrum in a high-frequency band (0.25–0.5 times the frame rate; Giovannucci et al., 2019).”

      (3.13) Figure 2D: Was imaging during the shift epoch always performed in dendrites? If so, could the imaging schedule bias comparisons between dendritic and somatic activity during learning, especially given that mice show behavioral learning between early and late post-shift sessions (Figure 1J)?

      No, imaging during the shift was not always performed in the dendrites. The following has been added to the Methods to make clear that the post-shift data reflect dendritic and somatic imaging conducted on the day of the shift with roughly similar frequency:

      “For Figure 5 and Figure 6, which make comparisons between dendritic and somatic activity during the post-shift period, 67% of animals providing somatic data (4 of 6 mice) underwent somatic imaging on the day of the shift and 80% of mice providing dendritic data (8 of 10 mice) underwent dendritic imaging on the day of the shift.”

      (3.14) Lines 163-164, "we observed that the onset of tuft activity was consistently time-locked to the GO cue (vertical green line; Figure 3B). This was in contrast to somatic activity, which had more variable timing (Figure 3E)." The authors cite panels B and E in support of this point, but these appear to be example ROIs. It would be helpful to clarify how representative these examples are, since the corresponding population summaries in panels G and H do not make the effect immediately apparent.

      These examples are representative, as supported by the population summary of activity time-locked to the GO-cue versus port contact in Figure 4B.

      (3.15) Figure 4B: It could be useful to add the lick traces here as well. To allow the reader to have an idea of contact timing with respect to the Go cue and compare the sustain response with the licking pattern.

      We understand how this could be helpful. However, since these exact traces are already present in Figure 3I,J, we think that adding them to Figure 4 is unnecessary and would add complexity to an already very busy figure.

      (3.16) Figure 5E: Why are the imaging sessions labeled 0 and +1 rather than 0 and +2? Are the dendritic and somatic imaging not alternated?

      Yes, imaging was not alternated for 3 of the 22 mice. We have clarified this in the Methods, as follows:

      “Somatic and dendritic imaging sessions alternated every other day (19 of 22 mice), except for 3 mice in which only one compartment was imaged daily (dendrite-only: 2 mice, soma-only: 1 mouse). The exceptions were due to brain curvature or the angle of the coverslip with respect to the brain, such that only one compartment could be imaged and the other compartment was underneath skull regrowth or dural thickening that made high-quality imaging impossible.”

      (3.17) Figure 6C, legend: I suppose the authors meant "remapping", not "Post-shit SI distribution" for the description of the right column.

      Thank you. We have fixed this label.

    1. eLife Assessment

      This study provides a useful analysis of the effects of arginine vasopressin (AVP) on islet cells in pancreatic tissue slices, using sophisticated spatiotemporal calcium recordings to show that AVP affects α and β cells differently depending on glucose concentration. The calcium imaging, analytical approaches, and V1b receptor-targeted peptide ligands are strengths of the work. However, the reviewers were concerned that the proposed mechanistic model is not sufficiently supported by the data. Characterisation of β-cell responses remains incomplete, and potential off-target effects and limited receptor specificity raise alternative explanations, including indirect effects mediated through α cells. The study would have been strengthened by signalling pathway analyses, genetic validation (e.g. β-cell-specific V1bR deletion), or selective V1b receptor silencing. The RNAscope data included in the revision indicate broader expression patterns but do not clearly establish receptor localisation within specific endocrine populations.

    2. Reviewer #1 (Public review):

      Summary:

      The paper investigates how AVP modulates pancreatic alpha and beta cell activity using acute mouse pancreatic tissue slices, calcium imaging, hormone secretion assays, RNAscope, and newly synthesized receptor-selective ligands. The Authors report that AVP regulates islet cell activity in a glucose- and state-dependent manner, with maximal effects occurring within physiological AVP concentrations and a bell-shaped concentration-response profile. They conclude that V1b receptors are the principal mediators of these effects and propose that IP3 receptor-dependent signaling underlies the observed nonlinear responses.

      Strengths:

      The use of fresh pancreatic tissue slices preserves islet architecture and cell-cell interactions, providing a physiologically relevant experimental model compared with isolated islets or immortalized cell lines.

      The combination of live calcium imaging, hormone secretion measurements, RNAscope, and pharmacological characterization of newly synthesized receptor-selective ligands represents a technically comprehensive experimental approach that addresses AVP signaling from multiple complementary perspectives.

      Weaknesses:

      (1) The central mechanistic model of the manuscript is not supported by the experimental data. Although the Authors repeatedly attribute the observed bell-shaped responses to IP3 Receptor activation and inactivation, no direct mechanistic evidence is provided to implicate IP3 receptors. Experiments assessing IP3 receptor function using genetic manipulation and direct measurements of IP3 signaling are necessary before such mechanistic conclusions can be drawn.

      (2) The Authors should directly demonstrate V1b receptor expression in β cells using complementary approaches, since the RNAscope data indicate broader expression but do not convincingly establish receptor localization within specific endocrine populations.

      (3) In my opinion, the central conclusion that V1b receptors are the predominant mediators of the observed effects is insufficiently supported because definitive loss-of-function experiments are lacking. Genetic deletion or selective silencing of V1b receptors should be provided to validate the proposed mechanism.

      (4) The heterogeneous responses observed among islets substantially weaken the proposed mechanistic model. Data should be provided to identify the determinants responsible for activation, absence of response, or inhibition in individual islets.

      (5) Please explain why the marked changes in alpha-cell calcium activity were not accompanied by corresponding alterations in glucagon secretion. This apparent discrepancy requires additional experimental evidence.

      (6) The Authors need to provide stronger evidence linking the observed calcium dynamics with insulin secretion, since calcium measurements alone cannot establish the proposed functional consequences.

      (7) Proper assays should be provided to assess whether the newly synthesized ligands exhibit comparable selectivity and efficacy at murine receptors rather than relying primarily on pharmacological characterization performed using human receptor-expressing cell lines.

      (8) The proposed absence of V1a receptor involvement is based primarily on pharmacological inhibition. Independent experimental approaches should be provided to exclude a contribution of this receptor subtype.

      (9) They must provide additional quantitative analyses demonstrating that the reported bell-shaped concentration-response relationship is robust across individual experiments rather than reflecting substantial biological variability.

      (10) The Authors should include experiments evaluating endogenous AVP signaling under more physiological conditions instead of relying predominantly on exogenous agonist administration.

      (11) I believe the role of forskolin deserves further clarification because many conclusions were obtained under cAMP-permissive conditions that may substantially influence AVP responses. Additional experiments without pharmacological cAMP stimulation should be presented.

      (12) Please clarify how beta cells and alpha cells were identified exclusively from functional activity patterns during calcium imaging and provide independent validation of cell identity within the analyzed recordings.

      (13) In my opinion, the manuscript relies heavily on changes in intracellular calcium activity as a surrogate for endocrine function, whereas the secretion data do not consistently support the proposed functional conclusions. Additional evidence is needed to establish a direct relationship between the observed calcium dynamics and hormone release.

      (14) The Authors should better reconcile their findings with previous reports showing minimal or absent AVP receptor expression in β cells and explain how the current data resolve these discrepancies rather than adding another possible interpretation.

    3. Reviewer #2 (Public review):

      Summary:

      In this paper Drs. Kercmar, Murko and Bombek make a series of observations related to the role of AVP in pancreatic islets. They use the pancreatic slice preparation that their group is well known for. The observations on the slide physiology are technically impressive. However, I am not convinced by the conclusions of this manuscript for a number of reasons. At the core of my concern is perhaps that this manuscript appears to be motivated to resolve 'controversies' surrounding the actions of AVP on insulin and glucagon secretion. This manuscript adds more observations, but these do not move the field forward in improving or solidifying our mechanistic understanding of AVP actions on islets. A major claim in this manuscript is the beta cell expression of the V1b Receptor for AVP, but the evidence presented in this paper fall short of supporting this claim. Observations on the activation of calcium in alpha cells via V1b receptor align with prior observations to this effect and can explain the effects of beta cell calcium and insulin secretion better than an explanation where beta cells express functional V1BR, for which direct evidence is lacking.

      I have focused my main concerns below. I hope the authors will consider these suggestions carefully - please be assured that they were made with the intent to support the authors and increase the impact of this work.

      Strengths:

      The main strength of this paper is the technical sophistication of the approach and the analysis and representation of the calcium traces from alpha and beta cells.

      Weaknesses:

      (1) There are excellent data that indicate that the actions of AVP are mediated via V1bR on alpha cells and that V1bR is 1) not expressed by beta cells and 2) does not activate beta cell calcium at all at 10 nM - which is the same concentration used in this paper (Figure 4G) for peak alpha cell Ca2+ activation (see https://doi.org/10.1016/j.cmet.2017.03.017; cited as ref 30 in the current manuscript). Any published stimulatory actions of AVP on insulin secretion can be explained by the potentiating effects of glucagon, released in response to AVP stimulation of alpha cells.

      (2) The RNAscope data offered in the revision as a second line of evidence for the expression of the V1bR in beta cells do not convince. I applaud the authors for trying as these are hard experiments to do well, as evidenced from the Gcg RNAscope signal that is not at all concentrated in the islet periphery, and in fact both color puncta occur outside of the islet at similar density. Absent a convincing concentration of Gcg signal (which is a very abundant transcript in alpha cells), it is hard to depend on these results. They certainly do not substitute experiments to determine cell autonomous activation of isolated beta cells by AVP. Claim of beta cell expression of V1br, require a more direct demonstration by staining (if appropriate antibodies exist), by beta cell-specific deletion of V1br, or by documenting the direct calcium activation in isolated beta cells in the absence of alpha cells. This should include a demonstration of Gaq-dependence in isolated beta cells.

      (3) We know from bulk RNAseq data on purified alpha, beta, and delta cells from both the Huising and Gribble groups that there is no expression of V2a. I will point you to the data from the Huising lab website published almost a decade ago (http://dx.doi.org/10.1016/j.molmet.2016.04.007) - which is publicly available and can be used to generate figures (https://huisinglab.com/data-ghrelin-ucsc/index.html). They indicate the absence of expression of not only AVP2 receptors anywhere in the islet - but the lack of expression of V1bra, V1brb, and Oxtr in beta cells. These AVP/OXT receptor expression data are largely and helpfully confirmed by the efforts in this paper that involved the generation of the V1aR agonist and V2R antagonist.

      (4) Importantly, the lack of V1br from beta cells does not invalidate observations that AVP affects calcium in beta cells, but it does indicate that these effects are mediated 1) indirectly, downstream of alpha cell V1br or 2) via an unknown off-target mechanism (less likely). The different peak efficacies in Figure 4G would also suggest they are not mediated by the same receptor. The recent work by Huixia Ren and colleagues (PMID: 41916313) that demonstrates that glucagon accelerates the frequency of beta cell calcium is in line with such a scenario.

      (5) The use of forskolin across almost all traces complicates the interpretation of the results. The design does not account for the elevation of cAMP in alpha cells and subsequent release of glucagon - particularly upon co-stimulation with AVP which permits glucagon release by activating a calcium response in alpha cells. This glucagon then could activate beta cells. If resolving the mechanism of action is the goal, often less is more. The activation of Gaq-mediated calcium is not cAMP dependent (although the downstream hormone secretion clearly often is). As was shown, AVP does not activate calcium in beta cells in the absence of cAMP. The experiments should have been completed in the absence of cAMP/forskolin, which would likely have had different outcomes on the beta cell responses and to the hormone secretion.

      (6) It is motivated by a desire to 'study the AVP dependence of both alpha and beta cells at the same time'. As best as I can determine, the design choice to conduct most studies under sustained forskolin stimulation is related to the permissive actions of AVP on hormone secretion in response to cAMP-generating stimuli. The permissive actions by AVP that are cited are on hormone secretion - which in many cell types requires activation of both calcium and cAMP signaling. Whether the activation of V1br and subsequent calcium responsive is permitted by cAMP is unclear. I believe the argument the authors are making here is that the activation of beta cell calcium by AVP is permitted by forskolin. i.e. the cAMP stimulated by it in beta cells.

      (7) Figure 9 suggests a pharmacological activation of beta cell V1bR in the low pM range. How do the authors reconcile this compare with the apparent absence of an effect of AVP stimulation at low pM to low nM doses in beta cells (Figure 5A). I note that there are changes over time with sustained beta cell stimulation with 8 mM glucose, but these changes are relatively subtle, gradual and quite likely represent the progression of calcium behaviors that would have occurred under sustained glucose irrespective of these very low AVP concentrations. I will note that the Kd of the V1bR for AVP is around 1 nM, with tracer displacement starting around 100 pM according to the data in figure 6B, which is hard to reconcile with changes in beta cell calcium by AVP doses that start 10-100-fold lower than this dose at 1 and 10 pM (Figure 9).

    4. Reviewer #3 (Public review):

      Summary:

      This work aims to better understand the role of arginine vasopressin (AVP) in the control of islet hormone secretion. This builds on previous literature in this area reporting on the actions of AVP to stimulate islet hormones. The gap in literature being addressed by these studies is primarily focused on the glucose-dependency of AVP on both insulin and glucagon secretion. A secondary objective is to explore the role of individual receptors with the use of newly generated peptides and existing tools. The methods include the use of Ca2+ imaging in pancreas slices from mice, with additional outcomes including insulin secretion in some areas. The conclusions presented are that AVP acts through V1b receptors in both alpha- and beta-cells, that this activity occurs in the high cAMP environment, and is glucose dependent.

      Strengths:

      The area of research is emerging with plenty of room for new contributions. The concept of AVP stimulating islet hormone secretion is important and deserving of further insight. The use of pancreas tissue to image primary cells makes the experiments physiologically relevant. The advancement of novel tools in this area should be helpful to other groups investigating the actions of AVP.

      Comments on revised version:

      Overall, the authors have modified their conclusions to more accurately capture the results of this manuscript. They also add the significant limitations outlined in the review process to the Discussion. With the addition of new data, however, a few concerns have emerged.

      (1) The rational for showing somatostatin staining in Figure 1 a is unclear. It also does not appear to be the same region of interest as in panel B.

      (2) It is difficult to assess the success of the RNAscope with the representative image used in Figure 1b. It is surprising how low the Gcg signal is in this image, suggesting some optimization is required. Additionally, how many mice were used (biological replicates) and how many V1b receptor+ and Gcg+ cells per mouse were quantified for the RNAscope images? This must be indicated in the methods section and should be sufficiently powered to make a conclusion.

      (3) While the addition of insulin and glucagon secretions with AVP ramp provide a functional output to the calcium imaging, it is unclear why the measurements are sometimes log10 transformed (Figure 5 K and L) but not always (Figure 5E). It is difficult to interpret negative glucagon values. What is the functional output of the dose-dependent calcium response to AVP in alpha cells if it is not glucagon?

      (4) Finally, the highlights section has not been refined to the revised interpretations of the manuscript.

    5. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      This study presents a useful finding on the effects of arginine vasopressin (AVP) on islet cells in pancreatic tissue slices, using technically sophisticated spatio-temporal calcium recordings to confirm that AVP influences α and β cells differently depending on glucose concentrations. While the study's methods - particularly the calcium imaging techniques and peptide ligand design targeting V1b receptors - are strong, the reviewers were concerned about several aspects of the experimental design. However, the results on βcell responses are incomplete and insufficient to support the manuscript's claims, especially due to the high variability of islet responses and lack of mechanistic and functional (hormone release) data. There are also concerns about the possibility of off-target effects and incomplete receptor specificity, noting that the study would have been significantly strengthened by inclusion of signaling pathway interrogation, hormone output assays, genetic validation (e.g., β cell-specific deletion of V1br), and receptor localization, although the work will still be of interest to researchers studying islet physiology in the context of health and diabetes.

      We sincerely thank the reviewers and editors for their thorough evaluation of our manuscript and their recognition of its technical strengths, including the advanced spatio-temporal calcium imaging and the rational design of selective V1b receptor ligands. We appreciate their acknowledgement of the study’s relevance for understanding AVP effects in a physiologically intact islet context and their positive assessment of our methodological rigour and innovation. The reviewers’ constructive feedback has helped us clarify the boundaries and intent of our study, which focuses on the glucose- and context-dependent modulation of α- and β-cell activity, rather than exhaustive molecular dissection.

      While the reviewers rightly emphasize the importance of receptor specificity and downstream signaling validation, we respectfully suggest that some of their concerns may reflect a lingering bias toward reductionist frameworks. Our interpretation is rooted in the emerging understanding that β-cell behaviour is largely defined by dynamic intercellular interactions within the islet collective, rather than by static gene expression or receptor localization alone (Jin et al., 2025; Korošak et al., 2021; Rutter et al., 2024). Recent studies have demonstrated that roles such as “leader” or “hub” β cells are transient and emergent, governed more by timing, environment, and local network structure than by fixed molecular identity (Postic et al., 2023; Gosak et al., 2018).

      This has profound implications for how we interpret cell responsiveness to agents like AVP: what appears as biological variability may in fact reflect context-sensitive transitions within a non-linear, self-organizing system (Stožer et al., 2021). Hence, we chose to focus on functional collective dynamics using intact pancreatic slices, rather than isolated cell models which fail to preserve the essential network architecture of islets. Although the addition of genetic models or isolated receptor measurements would strengthen receptor-specific conclusions, we argue that such approaches alone cannot resolve the physiological complexity of a system where function arises from cell–cell communication and spatiotemporal context.

      Indeed, the lack of direct correlation between receptor transcript abundance and functional outcomes has been noted in prior studies, reinforcing the view that function cannot be strictly predicted by molecular presence (Rutter et al., 2024). As articulated in our manuscript, the islet behaves as a sensory collective (Fancher & Mugler, 2017), where emergent patterns— not static cell identity—determine behaviour. This perspective aligns with broader shifts in biology away from strict genetic determinism toward causal emergence and collective agency (Ball, 2023; Levin, 2021).

      We therefore believe our study contributes not only new pharmacological insights but also a conceptual reframing of how AVP responses should be interpreted in a complex organ like the pancreas. We have added new data addressing reviewer suggestions—such as glucagon secretion assays, clarifications on the role of forskolin, and an analysis of event timing—that further support our conclusions. We also expanded the discussion on how islet variability is functionally meaningful, not just noise, and explained why β-cell responses to AVP must be interpreted within this probabilistic framework.

      We agree that future work should include receptor-specific knockouts and more direct signaling pathway assays, but these would need to be designed with careful consideration of the islet’s dynamic topology and the emergent nature of β-cell roles. In this light, we see our study not as the final word, but as a necessary systems-level foundation for more targeted interventions. We thank the reviewers again for their careful critiques and hope that our response clarifies both the rationale and scope of our work. Our revisions aim to enhance the paper’s clarity while maintaining its commitment to an integrative, physiology-rooted approach.

      We thank the reviewers and editors for their thoughtful and constructive assessment of our work. We are especially grateful for their recognition of the study’s technical strengths, including the use of spatio-temporal calcium imaging in intact pancreatic tissue and the strategic development of receptor-selective peptide ligands. We also appreciate their acknowledgement that our study contributes to the understanding of glucose-dependent AVP effects in islet physiology. The reviewers’ concerns regarding variability, receptor specificity, and functional validation helped us further clarify the scope and context of our study.

      We respectfully submit that some reservations stem from a reductionist framing that may not fully account for the collective behaviour of islets. As we and others have shown, β-cell function arises from emergent, self-organizing network dynamics, not just from static gene expression or receptor abundance (Jin et al., 2025; Korošak et al., 2021; Postic et al., 2023). In this view, pharmacological heterogeneity across islets is not simply noise or experimental inconsistency, but a signature of dynamic attractor states within the islet network (Stožer et al., 2021). Because an islet functions as a coupled system, most response variability originates from its emergent collective behavior, which eclipses variability in receptor expression or metabolic state.

      For this reason, even single-islet receptor quantification or ATP measurements would provide limited explanatory power: it is the state of the network—not absolute receptor levels—that determines whether a perturbation elicits activation or inhibition. As we illustrate in our graphical abstract, a single islet tested repeatedly under identical glucose conditions can yield divergent responses, simply because it occupies different dynamic states. These findings are in line with systems biology and network science approaches, which have revealed that cell function, especially in the β-cell collective, cannot be fully understood through reductionist parameters alone (Gosak et al., 2018; Ball, 2023).

      We have included glucagon secretion assays and new analyses to address key reviewer suggestions. Still, we chose not to pursue extensive knockouts or cAMP imaging, as these would require a different experimental scope and could risk disrupting the very dynamics we aim to understand. Likewise, while direct measurements of V1bR or IP3R expression would add molecular detail, they are not definitive without network context. The bell-shaped AVP dose-response curve and its explanation through IP3R inactivation are supported by prior studies; we invoke this mechanism not speculatively, but because it provides the most parsimonious explanation for the glucose-dependent shift in β-cell responsiveness.

      We also clarify that our study does not aim to resolve every mechanistic detail, but rather to offer a systems-level insight into how AVP modulates islet dynamics across varying glucose and cAMP contexts. The implications extend beyond AVP pharmacology, suggesting that perturbations to β-cell function must be understood within a probabilistic, state-dependent framework (Fancher & Mugler, 2017). This resonates with emerging concepts in cell physiology that emphasize causal emergence and local agency over static molecular determinism (Levin, 2021; Rutter et al., 2024).

      In summary, we see our work as part of a necessary shift in perspective—from linear receptor-function models to context-sensitive dynamic systems. We are grateful for the opportunity to revise our manuscript in response to insightful feedback and hope our clarifications and new data will strengthen its impact for the islet research community.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The authors confirmed earlier findings that AVP influences α and β cells differently, depending on glucose concentrations. At substimulatory glucose levels, AVP combined with forskolin - an activator of cAMP -did not significantly stimulate β cells, although it did activate α cells. Once glucose was raised to stimulatory levels, β cells became active, and α cell activity declined, indicating glucose's suppressive effect on α cells and permissive effect on β cells. Under physiological glucose levels (8-9 mM), forskolin enhanced β-cell calcium oscillations, and AVP further modulated this activity. However, AVP's effect on β cells was variable across islets and did not significantly alter AUC measurements (a combined indicator of oscillation frequency and duration). In α cells, forskolin and AVP led to increased activity even at high glucose levels, suggesting that α cells remain responsive despite expected suppression by insulin and glucose.

      Experiments with physiological concentrations of epinephrine suggest that AVP does not operate via Gs-coupled V2 receptors in β cells, as AVP could not counteract epinephrine's inhibitory effects. Instead, epinephrine reduced β cell activity while increasing α cell activity through different G-proteincoupled mechanisms. These results emphasize that AVP can potentiate αcell activation and has a nuanced, context-dependent effect on β cells.

      The most robust activation of both α and β cells by AVP occurred within its physiological osmo-regulatory range (~10-100 pM), confirming that AVP exerts bell-shaped concentration-dependent effects on β cells. At low concentrations, AVP increased β cell calcium oscillation frequency and reduced "halfwidths"; high concentrations eventually suppressed β cell activity, mimicking the muscarinic signaling. In α cells, higher AVP concentrations were required for peak activation, which was not blunted by receptor inactivation within physiological ranges.

      Attempting to further dissect the role of specific AVP receptors, the authors designed and tested peptide ligands selective for V1b receptors. These included a selective V1b agonist; a V1b agonist with antagonist properties at V1a and oxytocin receptors; and a selective V1a antagonist. In pancreatic slices, these peptides seem to replicate AVP's effects on Ca<sup>2+</sup> signaling, although responses were highly variable, with some islets showing increased activity and others no change or suppression. The variability was partly attributed to islet-specific baseline activity, and the authors conclude that AVP and V1b receptor agonists can modulate β cell activity in a statedependent manner, stimulating insulin secretion in quiescent cells and inhibiting it in already active cells.

      We applaud the reviewer to capture the essence of work in their introduction.

      Strengths:

      Overall, the study is technically advanced and provides useful pharmacological tools. However, the conclusions are limited by a lack of direct mechanistic and functional data. Addressing these gaps through a combination of signaling pathway interrogation, functional hormone output, genetic validation, and receptor localization would strengthen the conclusions and reduce the current (interpretive) ambiguity.

      Thank you!

      Weaknesses:

      (1) The study is entirely based on pharmacological tools. Without genetic models, off-target effects or incomplete specificity of the peptides cannot be fully ruled out.

      We partially agree with this comment and acknowledge that genetic models would provide a valuable complementary approach to address possible off-target effects or incomplete peptide specificity. However, genetic models also have important limitations, particularly when the aim is to resolve subtle, population-level physiological differences in beta cell activity. We therefore used pharmacological tools at different concentrations to test whether the observed effects were concentration-dependent and consistent with the expected receptor-mediated actions. An advantage of the pancreatic slice preparation is that it preserves much of the native tissue environment and allows pharmacological manipulation within concentration ranges closer to in vivo efficacy, thereby reducing the likelihood of nonspecific effects. To compensate for the lack of genetic models, we now emphasize the collective activity analysis as an additional strength of the study and have clarified this limitation in the revised manuscript.

      (2) Despite multiple claims about β cell activation or inhibition, the functional output - insulin secretion - is weakly assessed, and only in limited conditions. This aspect makes it very hard to correlate calcium dynamics with physiological outcomes.

      We agree that the functional output needed stronger support and have therefore expanded the hormone secretion experiments. While the effects of AVP and its analogues were tested during a stable plateau phase in the Ca<sup>2+</sup> imaging experiments, this phase provides only a narrow dynamic range for insulin release measurements in mouse slices. We therefore added a sequence of stimulations on the same slices, using 8 mM glucose and 500 nM forskolin, with glucose lowered to a non-stimulatory range between different AVP concentrations. These new experiments better define how AVP-dependent changes in Ca<sup>2+</sup> dynamics translate into insulin secretion under conditions with a broader secretory dynamic range. The new insulin and glucagon secretion data have now been added to the manuscript as Figure 5, and the text has been revised accordingly.

      (3) Insulin and glucagon secretion assays should be provided; the authors should measure hormone release in parallel with Ca2+ imaging, using perifusion assays, especially during AVP ramp and peptide ligand applications.

      We added insulin and glucagon secretion assays for AVP ramp to Figure 5.

      Additionally, there is no standardization of the metabolic state of islets. The authors should consider measuring islet NAD(P)H autofluorescence or mitochondrial potential (e.g., using TMRE) to control for metabolic variability that may affect responsiveness.

      We agree that standardization of the metabolic state of the islets would further strengthen the interpretation of the responsiveness data. We attempted to address this experimentally, but the results were inconclusive and therefore not included in the manuscript. Based on our previous unpublished observations, NAD(P)H levels appear to be significantly higher and less variable in islets within tissue slices than in isolated islets, suggesting that the slice preparation may better preserve the native metabolic state. However, we acknowledge that this remains an important limitation and we now indicate that in the manuscript. Additional experiments will be required to establish a robust and standardized approach, for example by combining NAD(P)H autofluorescence and/or mitochondrial potential measurements with Ca<sup>2+</sup> imaging.

      (4) There is a high degree of variability in response to AVP and V1b agonists across islets (activation, no effect, inhibition). Surprisingly, the authors do not fully explore the cause of this heterogeneity (whether it is due to receptor expression differences, metabolic state, experimental variability, or other conditions).

      This is a well-taken point and has indeed been one of the major bottlenecks in interpreting the results of this study. We agree that the variability in responses to AVP and V1b agonists may reflect several factors, including receptor expression, metabolic state, experimental conditions, and differences in the functional state of individual islets. However, our data also suggest that the beta cell population within an islet should be considered as a dynamic, non-linear system, in which even small differences in initial conditions or collective state can result in qualitatively different outcomes, including activation, no apparent effect, or inhibition. In this framework, the response to AVP is not determined by receptor expression alone, but by the current physiological context of the islet network. This is also why we believe that pharmacological tests are most informative when interpreted within a defined functional state rather than as isolated receptor-specific readouts. As indicated in the graphical abstract, apparently similar islets may occupy different dynamic states and therefore respond differently to the same Gq/PLC/IP3R stimulus. We have now expanded the discussion to make this interpretation more explicit and to acknowledge that receptor expression, metabolic variability, and experimental factors remain possible contributors that will require further targeted studies.

      The following text has been added to expand the discussion:

      “The heterogeneous responses observed across different islets, where some showed increased activity while others showed no detectable change or inhibition, could intuitively be attributed to variability in V1b receptor expression or signaling capacity among β cells. Such an explanation would be consistent with differences in receptor density, coupling efficiency to Gq proteins, or downstream signaling components such as PLC or IP<sub>3</sub> receptors. However, our data suggest that receptor-level variability alone is unlikely to fully explain the observed response spectrum, and that the current functional state of the islet collective must also be considered. The islet behaves as a non-linear dynamic system in which the same molecular perturbation can produce different functional outcomes depending on the current state of the β-cell collective. In such systems, cells or cell populations do not occupy a single deterministic activity state, but rather move within a landscape of possible states, with perturbations shifting the probability distribution of transitions between them. This concept is well established in dynamical systems approaches to biological cell-state transitions, where attractor landscapes, noise, and signaling inputs determine the probability of moving between alternative functional states rather than enforcing a single fixed output.

      In this framework, AVP and V1b receptor-selective agonists may reshape the probability landscape of β-cell activity. Depending on the initial metabolic, electrical, and Ca<sup>2+</sup>-handling state of the islet, the same stimulus may increase oscillation frequency, produce little detectable effect, or shift the system toward reduced activity or functional inactivation. This interpretation is also consistent with studies of pancreatic islet dynamics showing that βcell Ca<sup>2+</sup> activity emerges from coupled electrical, metabolic, and network interactions rather than from the properties of individual cells alone. Thus, molecular variability in V1b receptor expression or signaling capacity may contribute to the heterogeneous responses, but it is unlikely to determine them without considering the collective dynamic state of the islet.”

      (5) There is no validation of V1b receptor expression at the protein or mRNA level in α or β cells using in situ hybridization, immunohistochemistry, or spatial transcriptomics.

      We agree with the reviewer that spatial validation of V1b receptor expression is important for interpreting the cellular targets of AVP signaling in the islet. We have therefore added RNAscope in situ hybridization data to the revised manuscript to assess V1b receptor mRNA expression within the pancreas and islet. These new data show a broader expression pattern of V1b receptor transcripts within the islet than originally assumed, suggesting that AVP signaling may not be restricted to a single endocrine cell population. At the same time, the RNAscope analysis confirms previous reports of higher AVP receptor expression in glucagon-positive alpha cells. We have added these results to Figure 1 and revised the corresponding Results and Discussion sections to clarify that the observed functional responses may reflect both direct effects on beta cells and indirect intra-islet effects mediated through alpha-cell signaling.

      (6) AVP effects are described in terms of permissive or antagonistic effects on cAMP (especially in relation to epinephrine), but direct measurements of cAMP in α and β cells are not shown, weakening these conclusions. The authors should use Epac-based cAMP FRET sensors in α and β cells to monitor the interaction between AVP, forskolin, and epinephrine more conclusively.

      We agree that direct measurements of cAMP dynamics in alpha and beta cells would provide a more conclusive assessment of the interaction between AVP, forskolin, and epinephrine signaling. We attempted to address this experimentally; however, within the time domain of the Ca<sup>2+</sup> oscillations analyzed here, the temporal resolution and robustness of currently available cAMP readouts were not sufficient to resolve these interactions reliably. Even at slower time scales, cAMP sensor signals can be difficult to interpret quantitatively and may be overinterpreted if not tightly linked to the functional readout. We have therefore moderated the wording of the manuscript and now describe the proposed permissive or antagonistic interaction between AVP/V1b and cAMP-dependent signaling as an interpretation supported by the pharmacological Ca<sup>2+</sup> response patterns, rather than as a directly demonstrated cAMP mechanism. We now explicitly acknowledge in the limitations that most experiments were performed under cAMP-permissive conditions, which increases sensitivity for detecting AVP-dependent modulation but complicates the separation of direct beta cell effects from intra-islet interactions. Future studies using optimized cell-type-specific Epac-based sensors will be required to resolve this interaction.

      (7) Single-islet transcriptomics or proteomics (also to clarify variability) should be provided to analyze receptor expression variability across islets to correlate with response phenotypes (activation vs inhibition). Alternatively, the authors could perform calcium imaging with simultaneous insulin granule tracking or ATP levels to assess islet functional states.

      We agree that single-islet transcriptomics, proteomics, or simultaneous metabolic readouts could provide useful complementary information, particularly for describing molecular variability across islets. However, we do not think that differences in receptor expression or ATP levels alone are sufficient to explain the diversity of response phenotypes observed here. Our interpretation is that the beta cell population behaves as a collective dynamic system, in which the same input can lead to different outcomes depending on the current state of the network and its local physiological context. In such a system, AVP/V1b signaling does not necessarily impose a single deterministic response, but changes the probability distribution of accessible states, including activation, inhibition, or no detectable response. Theoretically and partially confirmed by the preliminary data, even the same islet exposed repeatedly under apparently identical conditions could be expected to display different responses if it occupies a different position within this dynamic state space at the time of stimulation. This concept is summarized in the graphical abstract and is central to our interpretation of the pharmacological data. We have therefore clarified in the Discussion that receptor expression, ATP levels, and other molecular parameters may modulate the response landscape, but are unlikely to fully define the observed functional phenotype without considering the collective dynamics of the islet.

      Added to Discussion section: “In this framework, AVP and V1b receptorselective agonists may reshape the probability landscape of β-cell activity. Depending on the initial metabolic, electrical, and Ca<sup>2+</sup>-handling state of the islet, the same stimulus may increase oscillation frequency, produce little detectable effect, or shift the system toward reduced activity or functional inactivation. This interpretation is also consistent with studies of pancreatic islet dynamics showing that β cell Ca<sup>2+</sup> activity emerges from coupled electrical, metabolic, and network interactions rather than from the properties of individual cells alone (63). Thus, molecular variability in V1b receptor expression or signaling capacity may contribute to the heterogeneous responses, but it is unlikely to determine them without considering the collective dynamic state of the islet.”

      (8) While the study implies AVP acts through V1b receptors on β cells, the signaling downstream (e.g., PLC activation, IP3R isoforms involved) is simply inferred but not directly shown.

      We agree that downstream signaling was not directly resolved at the level of PLC activation or specific IP3R isoforms. However, we did not infer Gq/PLC/IP3R involvement solely from AVP pharmacology, but used ACh as an independent Gq-coupled receptor reference stimulus in the same pancreatic slice preparation. With ACh concentration ramps, we could reproduce both activation and inactivation patterns observed with AVP/V1b stimulation, supporting the interpretation that these responses arise from modulation of the Gq-dependent Ca<sup>2+</sup> signaling axis.

      In addition, in prelilminary expriments we could observe that inhibition of Gq activity with YM254890, as well as interference with IP3R-dependent signaling using Xestospongin C, diminished the response, although not completely. This incomplete suppression is important, because it suggests that beta cell Ca<sup>2+</sup> homeostasis and collective islet activity are not controlled by a single linear pathway, but by partially redundant and context-dependent mechanisms. We have therefore revised the manuscript to state more cautiously that our data support the involvement of Gq/PLC/IP3R-dependent signaling, while acknowledging that direct measurements of PLC activity and IP3R isoform-specific contributions remain outside the scope of the present study.

      (9) The interpretation that IP3R inactivation (mentioned in the title!) underlies the bell-shaped AVP effect is just hypothetical, without direct measurements. Assays in β (and/or α)-cell-specific V1b KO mice and IP3R KO mice must be provided to support these speculations.

      We agree that the involvement of IP3R-dependent signaling should be stated with appropriate caution. However, the concept of IP3R inactivation as a mechanism contributing to bell-shaped Gq-dependent Ca<sup>2+</sup> responses is not purely hypothetical, since IP3R inactivation has been directly demonstrated in previous studies and provides a parsimonious explanation for the shift from activation to suppression at higher AVP concentrations. In the present study, this interpretation is further supported by the glucose dependence of the AVP concentration-response relationship, where different stimulatory glucose conditions shift the apparent efficacy peak.

      We also agree that cell-specific V1b receptor and IP3R knockout experiments would be valuable future approaches. In this respect, we have obtained preliminary results from a small sample of IP3R triple-knockout mice, which cannot yet be fully included because they are part of an ongoing collaboration. In these experiments, supraphysiological AVP concentrations did not produce the IP3R-like beta-cell response pattern observed in controls, namely reduced halfwidth and increased frequency, whereas alpha cell stimulation was preserved similarly to WT slices.

      At the same time, we believe that definitive knockout experiments must be carefully designed, because the beta cell population behaves as a dynamic collective system in which the response to AVP depends on the current functional state of the islet, glucose context, and intercellular coupling. We therefore now present IP3R inactivation as a strongly supported mechanistic interpretation rather than as a directly proven mechanism in this study, and we explicitly acknowledge that cell-specific V1b and IP3R genetic models will be required to fully resolve this pathway.

      Reviewer #2 (Public review):

      Summary:

      In this paper, Drs. Kercmar, Murko, and Bombek make a series of observations related to the role of AVP in pancreatic islets. They use the pancreatic slice preparation that their group is well known for. The observations on the slide physiology are technically impressive. However, I am not convinced by the conclusions of this manuscript for a number of reasons. At the core of my concern is perhaps that this manuscript appears to be motivated to resolve 'controversies' surrounding the actions of AVP on insulin and glucagon secretion. This manuscript adds more observations, but these do not move the field forward in improving or solidifying our mechanistic understanding of AVP actions on islets. A major claim in this manuscript is the beta cell expression of the V1b Receptor for AVP, but the evidence presented in this paper falls short of supporting this claim.

      Observations on the activation of calcium in alpha cells via V1b receptor align with prior observations of this effect.

      I have focused my main concerns below. I hope the authors will consider these suggestions carefully - please be assured that they were made with the intent to support the authors and increase the impact of this work.

      We thank the reviewer for their detailed input and support to increase the impact of our work and our understanding of important cellular processes overall. We have considered their suggestions carefully to further expand the strenghts of our approach and analysis.

      Strengths:

      The main strength of this paper is the technical sophistication of the approach and the analysis and representation of the calcium traces from alpha and beta cells.

      Thank you!

      Weaknesses:

      (1) The introduction is long and summarizes a substantive body of literature on AVP actions on insulin secretion in vivo. There are a number of possible explanations for these observations that do not directly target islet cells. If the goal is to resolve the mechanistic basis of AVP action on alpha and beta cells, the more limited number of papers that describe direct islet effects is more helpful. There are excellent data that indicate that the actions of AVP are mediated via V1bR on alpha cells and that V1bR is a) not expressed by beta cells and b) does not activate beta cell calcium at all at 10 nM - which is the same concentration used in this paper (Figure 4G) for peak alpha cell Ca2+ activation (see https://doi.org/10.1016/j.cmet.2017.03.017; cited as ref 30 in the current manuscript).

      We thank the reviewer for this important comment and agree that the literature on AVP actions in vivo is complex, with several possible sites of action outside the islet. We have therefore revised the Introduction to make the rationale more focused and to better separate systemic effects of AVP from studies addressing direct actions on pancreatic islet cells. At the same time, we chose not to restrict the Introduction only to the alpha cell V1bR literature, because one of the aims of the manuscript is precisely to address why AVP effects on insulin secretion have remained difficult to interpret across experimental contexts.

      Our results fully confirm a central aspect of the study cited by the reviewer, namely that V1bR activation robustly stimulates alpha cell Ca<sup>2+</sup> activity under non-stimulatory glucose conditions, and that 10 nM AVP does not produce a uniform activation of beta cell Ca<sup>2+</sup> activity. In fact, in a substantial fraction of beta cell populations, 10 nM AVP failed to activate oscillations, consistent with the view that alpha cells are the more sensitive and more direct cellular target of AVP/V1bR signaling. However, we do not think that the available transcriptomic evidence is sufficient to categorically exclude V1bR expression or functional relevance in beta cells. Re-analysis of the published dataset, together with more recent datasets and our newly added RNAscope data, supports a higher relative expression of V1bR transcripts in alpha than in beta cells, but does not justify treating beta (or non-alpha) cell expression as absent.

      We have therefore revised the manuscript to avoid overstating beta cell V1bR expression as a major isolated claim. Instead, we now present the data as evidence that AVP/V1bR signaling acts most prominently through alpha cells, while beta cell responses emerge in a concentration-, glucose-, and statedependent manner within the intact islet. This interpretation is consistent with the reviewer’s concern that 10 nM AVP preferentially activates alpha cells, but it also accommodates our observation that beta cell collective activity can be modulated under defined pharmacological and metabolic conditions. We believe that this is an important distinction, because the absence of a uniform beta cell Ca<sup>2+</sup> activation at one AVP concentration does not exclude beta cell modulation by AVP/V1bR signaling within the intact islet network. The Introduction and Discussion have been revised accordingly to clarify that our study does not simply challenge the alpha cell V1bR model, but expands it by examining how AVP-dependent alpha cell activation, possibly lower beta-cell receptor expression, and collective beta cell dynamics interact in the native pancreatic slice preparation.

      (2) We know from bulk RNAseq data on purified alpha, beta, and delta cells from both the Huising and Gribble groups that there is no expression of V2a. I will point you to the data from the Huising lab website published almost a decade ago (http://dx.doi.org/10.1016/j.molmet.2016.04.007) - which is publicly available and can be used to generate figures (https://huisinglab.com/dataghrelin-ucsc/index.html). They indicate the absence of expression of not only AVP2 receptors anywhere in the islet, but also the lack of expression of V1bra, V1brb, and Oxtr in beta cells. Instead of the detailed list of expression of these 4 receptors elsewhere in the body, it would be more directly relevant to set up their pancreatic slice experiments to summarize the known expression in pancreatic islets that is publicly available. It would also have helped ground the efforts that involved the generation of the V1aR agonist and V2R antagonist, which confirm these known AVP/OXT receptor expression patterns.

      We thank the reviewer for pointing us more directly to the publicly available islet expression datasets. We agree that the expression of AVP/OXT receptors in purified alpha, beta, and delta cells provides an important reference frame for interpreting our pharmacological data, and we have revised the manuscript to summarize these islet-specific datasets more directly rather than emphasizing receptor expression in other organs. These data support the absence or very low expression of V2 receptors in islet endocrine cells and confirm that V1b receptor expression is substantially enriched in alpha cells compared with beta cells.

      At the same time, as outlined in our response above, we do not think that the currently available transcriptomic datasets are sufficient to categorically exclude low-level V1bR transcript expression or functional relevance in beta cells within the intact islet. For this reason, we added independent RNAscope validation to assess V1bR transcripts in the pancreatic slice preparation. These data confirm stronger V1bR expression in glucagon-positive alpha cells, while also showing a broader expression pattern within the islet and pancreas.

      We have also revised the rationale for the pharmacological experiments using V1aR- and V2R-directed tools. We now present these experiments not as evidence for unexpected receptor expression, but as functional controls that are consistent with the known AVP/OXT receptor expression patterns in pancreatic islets. This better aligns the manuscript with the existing transcriptomic literature while preserving the main physiological question of the study: how AVP/V1bR-dependent signaling reshapes alpha-cell activity and beta-cell collective dynamics in intact pancreatic tissue.

      (3) Importantly, the lack of V1br from beta cells does not invalidate observations that AVP affects calcium in beta cells, but it does indicate that these effects are mediated a) indirectly, downstream of alpha cell V1br or b) via an unknown off-target mechanism (less likely). The different peak efficacies in Figure 4G would also suggest that they are not mediated by the same receptor.

      We agree with the reviewer that the absence or very low abundance of V1bR transcripts in beta cells in published transcriptomic datasets would not invalidate the observation that AVP modulates beta-cell Ca<sup>2+</sup> activity. It does, however, raise the important question of whether this modulation is mediated indirectly through alpha-cell V1bR activation, through V1bR expression in beta cells that is difficult to resolve transcriptomically, or through another mechanism. To address this more directly, we have now added RNAscope data, which confirm relatively stronger V1bR transcript enrichment in glucagon-positive alpha cells, but also show a broader V1bR transcript signal within the islet and pancreas. Thus, while our data support alpha cells as the dominant V1bR-positive endocrine population, they do not support a strict absence of V1bR-associated signaling capacity in the beta cell compartment.

      We also agree that different peak efficacies in alpha and beta cells could be interpreted as evidence for distinct receptors or indirect mechanisms. However, we favor a different interpretation: the apparent efficacy of AVP depends strongly on the physiological state in which the cells are tested. This is particularly evident in beta cells, where the AVP efficacy peak shifts with glucose concentration, suggesting that the beta-cell response is shaped by the metabolic and Ca<sup>2+</sup>-handling context rather than by receptor occupancy alone. In this framework, the same V1bR/Gq-dependent input can generate different downstream Ca<sup>2+</sup> outcomes in alpha and beta cells because the two cell types operate in different dynamic regimes.

      We have therefore revised the manuscript to acknowledge this dilemma more explicitly. We now state that beta cell effects of AVP could include indirect alpha cell-dependent components, but given the magnitude and statedependence of the beta cell Ca<sup>2+</sup> response it is unlikely to be driven by alpha cell activation. Instead, our preferred interpretation is that AVP/V1bR signaling acts within the intact islet as a context-dependent perturbation of the collective beta cell Ca<sup>2+</sup> system, with IP3R-dependent mechanisms being modulated by glucose-dependent changes in beta cell excitability and intracellular Ca<sup>2+</sup> handling.

      (4) The rationale for the use of forskolin across almost all traces is unclear. It is motivated by a desire to 'study the AVP dependence of both alpha and beta cells at the same time'. As best as I can determine, the design choice to conduct all studies under sustained forskolin stimulation is related to the permissive actions of AVP on hormone secretion in response to cAMPgenerating stimuli. The permissive actions by AVP that are cited are on hormone secretion, which in many cell types requires activation of both calcium and cAMP signaling. Whether the activation of V1br and subsequent calcium response is permitted by cAMP is unclear. I believe the argument the authors are making here is that the activation of beta cell calcium by AVP is permitted by forskolin. i.e., the cAMP stimulated by it in beta cells. However, the design does not account for the elevation of cAMP in alpha cells and subsequent release of glucagon, particularly upon co-stimulation with AVP, which permits glucagon release by activating a calcium response in alpha cells. This glucagon could then activate beta cells. If resolving the mechanism of action is the goal, often less is more. The activation of Gaq-mediated calcium is not cAMP dependent (although the downstream hormone secretion clearly often is). As was shown, AVP does not activate calcium in beta cells in the absence of cAMP. The experiments in Figures 1, 2, and 4 should have been completed in the absence of cAMP first.

      We agree with the reviewer that the use of forskolin needs to be explained more clearly, and we have revised the manuscript accordingly. Our rationale was based on the established permissive role of cAMP in AVP-dependent endocrine responses, but we acknowledge that this does not necessarily imply that the upstream V1bR/Gq-mediated Ca<sup>2+</sup> response itself is cAMP-dependent. The reviewer is also correct that forskolin elevates cAMP broadly and therefore may affect both alpha and beta cells, including the possibility that AVP-enhanced alpha cell activation and glucagon release secondarily influence beta cell activity.

      In fact, our initial experiments were performed without forskolin and revealed an important difficulty: stimulatory glucose alone can increase cAMP levels to a variable extent, as also supported by our previous work on epinephrine signaling, thereby shifting the apparent peak efficacy of AVP stimulation. Thus, forskolin was originally used to reduce this variability and create a more defined cAMP-permissive background in which alpha and beta cell responses could be compared in the same slice. However, we agree that this design works against isolation of beta cell-autonomous AVP effects.

      Within a scope of another study we have done an independent series of more focused experiments using GLP-1 receptor stimulation, which preferentially increases cAMP signaling in beta cells compared with the broad cAMP elevation produced by forskolin. We have clarified that the modulation of the AVP-dependent pathway by GLP-1 and related ligands at largely supports beta cell-autonomous AVP effects. It is part of ongoing work and will be reported independently, because a full mechanistic dissection of cAMP–AVP interactions goes far beyond the scope of the present study.

      (5) It is unexpected that epinephrine in Figure 2 does not activate the alpha cell calcium? A recent paper from the same group (Sluga et al) shows robust calcium activation in alpha cells in a similar prep by 1 nM epinephrine, which is similar to the dose used here.

      We thank the reviewer for pointing this out, but we would like to clarify that epinephrine did significantly activate alpha-cell Ca<sup>2+</sup> activity in our experiments, as shown in Fig. 3F. This result is consistent with our previous study by Sluga et al., where low nanomolar epinephrine robustly activated alpha cell Ca<sup>2+</sup> signals in the pancreatic slice preparation. The main point of the present comparison was therefore not that epinephrine is inactive in alpha cells, but that AVP produces a substantially stronger and reproducible alpha cell Ca<sup>2+</sup> response under comparable experimental conditions. This is also consistent with the data of van der Meulen et al., supporting the view that AVP/V1bR signaling is a particularly potent activator of alpha cell activity. We have revised the text to make this comparison clearer and to avoid the impression that epinephrine failed to activate alpha cells in our preparation.

      (6) Figure 8 suggests a pharmacological activation of beta cell V1bR in the low pM range. How do the authors reconcile this comparison with the apparent absence of an effect of AVP stimulation at low pM to low nM doses in beta cells (Figure 4A)? I note that there are changes over time with sustained beta cell stimulation with 8 mM glucose, but these changes are relatively subtle, gradual, and quite likely represent the progression of calcium behaviors that would have occurred under sustained glucose, irrespective of these very low AVP concentrations. I will note that the Kd of the V1bR for AVP is around 1 nM, with tracer displacement starting around 100 pM according to the data in figure 5B, which is hard to reconcile with changes in beta cell calcium by AVP doses that start 10-100-fold lower than this dose at 1 and 10 pM (Figure 8).

      We agree that the interpretation of low-pM AVP effects requires caution, particularly when compared with reported V1bR binding affinities. The apparent discrepancy between Fig. 4A and Fig. 8 most likely reflects differences in experimental design, stimulation context, and readout sensitivity. In Fig. 4A, we assessed acute AVP effects under conditions in which beta cell Ca<sup>2+</sup> responses are relatively threshold-dependent and where low AVP concentrations produced little or no activation. In contrast, Fig. 8 analyzes prolonged beta cell population dynamics during sustained stimulation with 8 mM glucose, a physiological stimulatory context in which even weak modulatory inputs may become detectable at the level of collective Ca<sup>2+</sup> activity.

      Importantly, the strongest and statistically significant effect was observed at 100 pM AVP, while lower pM concentrations showed only a trend. We therefore do not interpret the low-pM range as evidence for robust direct pharmacological activation of beta cell V1bR. Rather, these data suggest that AVP may exert permissive or modulatory effects within an already active beta cell network, where glucose-dependent excitability, receptor-effector coupling, and Ca<sup>2+</sup> amplification mechanisms can enhance the apparent efficacy of weak inputs. This interpretation is consistent with the known permissive role of AVP in endocrine responses, where AVP may not act as a primary activator alone but can increase the efficacy of other physiological stimuli.

      We have also clarified that sustained 8 mM glucose alone does not account for these effects, since Suppl. Fig. 1 shows no comparable time-dependent progression of Ca<sup>2+</sup> behavior under sustained glucose stimulation alone. Thus, we now present the low-concentration AVP effects as subtle, contextdependent modulation within the physiological stimulatory range, rather than as evidence for direct beta cell activation at concentrations below the expected receptor affinity range.

      Reviewer #3 (Public review):

      Summary:

      This work aims to better understand the role of arginine vasopressin (AVP) in the control of islet hormone secretion. This builds on previous literature in this area reporting on the actions of AVP to stimulate islet hormones. The gap in literature being addressed by these studies is primarily focused on the glucose-dependency of AVP on both insulin and glucagon secretion. A secondary objective is to explore the role of individual receptors with the use of newly generated peptides and existing tools. The methods include the use of Ca2+ imaging in pancreas slices from mice, with additional outcomes including insulin secretion in some areas. The conclusions presented are that AVP acts through V1b receptors in both alpha- and beta-cells, that this activity occurs in the high cAMP environment, and is glucose dependent.

      Strengths:

      The area of research is emerging with plenty of room for new contributions. The concept of AVP stimulating islet hormone secretion is important and deserving of further insight. The use of pancreas tissue to image primary cells makes the experiments physiologically relevant. The advancement of novel tools in this area should be helpful to other groups investigating the actions of AVP.

      We would like to thank the reviewer for recognizing the potential of our emerging area of research.

      Weaknesses:

      The conclusions are only modestly supported by the data and lack experimental depth and rigor. The rationale for only conducting studies at high cAMP conditions is not entirely clear and limits the conclusions that can be made. The use of Ca2+ is helpful, but it is a surrogate for hormone secretion. Additional measurements of hormone secretion are needed to enhance the robustness of these conclusions. Consideration of paracrine effects between alpha- and beta-cells is only superficially made and is likely essential in the context of the experimental design. For instance, there is clear literature that alpha-cells secrete several factors that work in paracrine interactions on beta-cells and autocrine actions back on alpha-cells. Conducting these studies in a high cAMP context only completely overlooks these interactions, skewing the interpretations made by the investigators. Finally, the clarity of the experiments and results could be significantly enhanced.

      We thank the reviewer for this balanced assessment and for emphasizing several issues that are central to the interpretation of our study. We agree and now explicitly state in the Limitations, that Ca<sup>2+</sup> oscillations are a surrogate readout for hormone secretion and that currently used stimulation protocols are not optimized to directly quantify the relationship between Ca<sup>2+</sup> dynamics and secretory output. To address this limitation, we have now expanded the functional part of the study by adding complete insulin and glucagon secretion measurements during AVP concentration ramps. These new data provide a stronger functional framework for interpreting the Ca<sup>2+</sup> imaging results, while also clarifying that Ca<sup>2+</sup> activity and secretion cannot be assumed to correlate linearly under all stimulation protocols.

      We have also revised the rationale for the high-cAMP experimental condition. The original aim was to reduce variability arising from glucose-dependent endogenous cAMP signaling and to study alpha and beta cell responses in a common permissive background. However, we agree that broad forskolin stimulation complicates the interpretation of cell-autonomous versus paracrine mechanisms. Independent experiments within a scope of another study demonstrate that using GLP-1 co-stimulation, which provides a more beta-cell-oriented cAMP-permissive condition and supports the interpretation that AVP can modulate beta-cell collective activity in a manner that is not solely secondary to alpha-cell activation.

      We fully agree that paracrine interactions within the islet are physiologically important and must be considered, particularly in intact pancreatic slices. Nevertheless, the rapid onset of the AVP effects observed in beta cell Ca<sup>2+</sup> activity argues against a mechanism mediated predominantly by slower indirect paracrine loops. High AVP concentrations, as shown in Fig. 5, significantly shorten the intervals between Ca<sup>2+</sup> events in both alpha and beta cells, but that the activity of the two cell populations remains largely noncoordinated. This temporal dissociation does not exclude paracrine modulation altogether, but it argues against a simple alpha-cell-driven explanation for the beta-cell response.

      We have revised the manuscript to state these points more clearly and to moderate conclusions where the data support modulation rather than definitive cell-autonomous receptor action. We believe that the added secretion experiments and alpha/beta event-timing analysis substantially strengthen the physiological interpretation of the study, and we thank the reviewer for raising these issues.

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      (1) The paragraph discussing the benefits of slice physiology over islets is not reflective of how most - if not all of your colleagues who do islet experiments conduct these. Many labs have reported for years high-quality GSIS experiments, synchronous calcium responses, and a plethora of studies detailing the mechanism of hormone and neurotransmitter actions using islet models, and have done so well. Slice physiology is a unique and helpful model that can have advantages over other models. This particular reviewer uses both models in their lab and each has benefits and - inevitably - drawbacks. Many of the possible drawbacks cited for islet studies apply equally to slices, including the possibility of altered gene expression, lack of innervation, and circulation. Added drawbacks are the exposure to higher levels of pancreatic enzymes from the slice, which require co-culture with enzyme inhibitors.

      We agree with the reviewer and have revised these limitations accordingly. Our intention was not to imply that isolated islet preparations are generally inferior, since they have provided a highly productive and rigorous experimental platform for GSIS, synchronized Ca<sup>2+</sup> dynamics, and mechanistic studies of hormonal and neurotransmitter regulation with standardized protocols with all their positive and negative sides. We modified the presentation of pancreatic slices as a complementary model with specific advantages, particularly preservation of local tissue architecture, while also acknowledging their limitations. The revised text therefore avoids a comparative hierarchy between slices and isolated islets and instead emphasizes that both models have distinct strengths and drawbacks depending on the experimental question.

      (2) If you want to demonstrate direct actions on beta cells, deconstructing the islet would be a better way to go. Less complicated, not more. Dissociated beta cells, instead of slices, were used just to prove or disprove the hypothesis of direct beta cell effects of AVP.

      We agree that dissociated beta cells can be a useful reductionist model to test whether AVP is capable of acting directly on individual beta cells. However, this approach would also remove the collective beta-cell activity that is central to the physiological question addressed in the present study. Since our data indicate that AVP effects emerge within the intact islet as rapid and extensive changes in coordinated Ca<sup>2+</sup> dynamics, dissociation would not necessarily provide a more informative model for understanding these responses. The fast onset and magnitude of the beta cell response argue against a predominantly indirect non-autonomous mechanism, and this interpretation is further supported by the GLP-1 co-stimulation experiments, which are more consistent with beta cell-autonomous modulation. We have therefore clarified in the revised manuscript that dissociated-cell experiments would be valuable for a narrowly defined receptor-cell autonomy question, but would not resolve the collective islet dynamics that are the focus of this work.

      (3) If you want to sustain the claim of beta cell expression of V1br, you would have to demonstrate this far more directly by staining (if appropriate antibodies exist), by beta cell-specific deletion of V1br, or by highly selective, well-validated pharmacology. This should include a demonstration of Gaqdependence in isolated beta cells.

      We have added RNAscope in situ hybridization data to the revised manuscript to assess V1b receptor mRNA expression within the pancreas and islet. These new data show a broader expression pattern of V1b receptor transcripts within the islet than originally assumed, suggesting that AVP signaling may not be restricted to a single endocrine cell population. At the same time, the RNAscope analysis confirms previous reports of higher AVP receptor expression in glucagon-positive alpha cells. We have added these results to Figure 1 and revised the corresponding Results and Discussion sections to clarify that the observed functional responses may reflect both direct effects on beta cells and indirect intra-islet effects mediated through alpha-cell signaling.

      Minor

      (1) O'Carroll et al. should be cited in the context of islet permissive actions of AVP/cAMP. PMID: 18434353, although that paper offers no evidence that the AVP-dependent potentiation of insulin release is mediated directly by beta cells. It does confirm dependence on PKC.

      We agree and have added O’Carroll et al. in the revised manuscript in the context of AVP/cAMP-dependent permissive actions on islet hormone secretion. We therefore use it as support for the broader concept of AVP-dependent amplification of secretion in a permissive signaling context, rather than as direct evidence for beta-cell-autonomous V1bR signaling.

      (2) Figure 2 E-H, glucose concentration mislabeled.

      We thank the reviewer for pointing this out. The glucose concentration label has been clarified: panels E–H show pooled data from separate experiments performed at 8 mM glucose, whereas panel D shows a representative experiment performed at 9 mM glucose.

      (3) The insulin secretion in 4E is difficult to interpret without a low-glucose control. If this is hard to do in a slice preparation, a separate static islet secretion experiment would help here. The possibility that the inhibition of insulin secretion traces back to the activation of delta cells by AVP could be considered - I struggle to come up with a plausible mechanistic explanation why AVP (which activates calcium in alpha and in beta cells in the presence of 8 mM G plus forskolin according to your data) would inhibit insulin secretion.

      We agree that the original insulin secretion experiment was difficult to interpret without a clearer low-glucose reference condition. To address this, we have added new insulin release experiments in which glucose was lowered to a non-stimulatory range between AVP concentrations, followed by sequential stimulation with 8 mM glucose and 500 nM forskolin in the same slices. These new data provide a broader dynamic range for assessing insulin secretion and allow a more direct comparison between AVP-dependent Ca<sup>2+</sup> modulation and secretory output.

      We also agree that AVP-dependent inhibition of insulin secretion requires careful interpretation. One possible explanation is not simply activation of delta cells, but a failure of the beta cell collective to maintain coordinated activity at very high AVP concentrations. In the Ca<sup>2+</sup> imaging data, high AVP concentrations increase activity in many beta cells, but numerous cells within the islet fail to keep pace with the collective oscillatory rhythm, leading to fragmented and less synchronized population activity. Thus, despite increased frequency of Ca<sup>2+</sup> oscillations in the islet, the integrated beta cell output may become less efficient for insulin secretion. We have added this interpretation to the revised manuscript and now discuss delta cell activation as a possible contributing mechanism, but not as the primary explanation supported by our current data.

    1. eLife Assessment

      In this important work, the authors develop methods to forecast epidemic growth from viral sequencing data alone. The evidence for the usefulness of the approach is solid, but some justifications and methodological details are incomplete. This study should be of broad interest to the community interested in viral dynamics and epidemiology.

    2. Reviewer #1 (Public review):

      Summary:

      This paper develops a formalism for quantifying epidemic dynamics in terms of relative fitnesses of circulating variants, uses the formalism to elucidate fundamental tradeoffs of epidemics driven by variants with increased transmissibility versus immune escape capability, shows the formalism implies a natural quantity measuring the impact of selection on epidemic growth, and demonstrates that the formalism enables a decomposition of epidemic dynamics into circulation among different immunity groups. The relative fitness formalism enables these analyses to be performed with genetic sequence data only, a major benefit of the model given the relatively high availability of sequence data compared to other data streams such as case counts and titers.

      Strengths:

      Linking epidemic dynamics to pathogen evolution is a fundamental problem in studies of antigenically variable pathogens, with models of epidemic dynamics and immune-driven evolution going back decades in applications to respiratory pathogens such as influenza. The COVID-19 pandemic heightened the urgency for developing methods for quantifying epidemic growth in contexts where novel variants emerge, leading to differential susceptibility among individuals with diverse exposure histories with implications for vaccination strategies. Real-world data streams such as case counts and immunological measurements have a variety of shortcomings that pose major challenges for quantitative models aiming to inform policy. In recent years, genetic sequencing data has become widely available for pathogens including SARS-CoV-2 and influenza, allowing tracking of pathogen evolution at unprecedented detail in real time, yet biases in the collection of sequence data across different populations make connections between absolute epidemic size and variant frequencies from sequence data not immediately transparent.

      This paper's contributions are exciting because they demonstrate new ways to link pathogen evolution and epidemic dynamics using very accessible data. From a theoretical perspective, the model is appealing because of its simple derivation in terms of compartmental models of epidemics, which are standard in the literature, and its clear extension to populations with heterogeneous immune histories. The latter extension leads directly to new methods for inferring immune groups with differential susceptibility to antigenically distinct variants in populations with heterogeneous immune histories without access to immunological data such as titers, an important advance given the wide applicability of quantification of antigenic relationships among variants in real populations.

      Weaknesses:

      While the demonstrated methods for forecasting short-term epidemic growth and for quantifying population immunity using sequence data are exciting as proofs of principle, the validation and statistical support provided in the analyses have drawbacks that are not fully addressed in the manuscript, weakening the evidence for the usefulness of the methods in their current form.

      The analyses forecasting epidemic growth using Gaussian process models are justified using Pearson correlation coefficients whose values are extremely low for the test data period. The explanation given for this is that the case data used to validate the predictions has worse ascertainment over time, but it is not shown directly that the model may be working well despite the low correlations. Whereas, by eye, the predicted epidemic growth curves appear to capture features of the observed epidemic growth curves, the computed metrics don't support the claim of success of the predictions. Additionally, nearly all the model fits lack estimates of uncertainty, so it is not possible to discern the significance of departures between the model and data, or subtle differences in relative fitness calculations across geographies.

      The analysis of latent pseudo-immune components also suffers drawbacks that render it more of an interesting proof of principle than a convincing tool for prediction at this point. In particular, in figures S18 and S19, metrics meant to quantify the statistical significance of the results show no difference from null models computed by permuting variants and their escape vectors, yet no interpretation is given for the lack of significance. Moreover, the model fits relating titer distance to pseudo escape distance seem unsuccessful for JN.1 infection and XBB infection histories, which is not adequately accounted for in the text, which cites just "weaker correlations" in these cohorts.

      In several instances, the evidence for the new data analyses is weakened by a lack of clarity in the presentation of the technical details of the methods. For example, in the discussion of the Gaussian process models, it was not clear what features of the problem inform the choice of kernel (Matern 5/2), which hyperparameters were used, and how novel this use of Gaussian processes is. In the section describing methods for predicting epidemic growth rate from selective pressure, the discussion of the gradient boosting regressor model provided no intuition as to why this method performed better than the others tested or whether this was particularly important to the conclusions, and the lack of discussion of uncertainty or variability in the model predictions makes it difficult to assess the significance of the time series estimates alone. In the discussion of the latent immune factor model, the mismatch between the notation used in Equation 5 compared to that in Equation 18 made the derivations more difficult to follow. Subsequently, the explanation of the fitting of the pseudo-immune model left out details, such as an explicit definition of distance in pseudo-escape space, to what extent the group-level mean aggregated titer measurement captured features of the titer data (despite ignoring interindividual variability), and a thorough discussion of the successes and shortcomings of the fits in different scenarios. More explicit presentation of the mathematical choices going into the methods, sources and quantification of uncertainty, and cases where the model performs well or poorly could significantly bolster the case for the usefulness of sequence data in quantitatively predicting epidemic growth and antigenic relationships among variants in practice, in more general settings than those carried out here.

    3. Reviewer #2 (Public review):

      Summary:

      The authors first introduce a framework to understand how different phenotypic drivers of viral evolution, i.e., changes in transmissibility versus immune escape, complicate epidemic forecasting using only genetic data. To overcome these complications, they advance an evolutionary "selective pressure" metric to predict population-wide epidemic growth from genetic data alone. Separately, they introduce a latent space model to infer a "pseudo" population immune structure from geographic variation in viral lineage dynamics, and find that the inferred pseudo-structure predicts human serological data.

      Strengths:

      This paper begins with a useful pedagogical exposition on the connection between fitness-driven frequency dynamics and underlying mechanisms of viral-immune co-evolution. A major contribution of this paper - a method to infer variant-specific escape properties from geographically non-uniform variant frequency dynamics alone - is an interesting and potentially timely one, given the advance of sequencing-based surveillance.

      Weaknesses:

      The logical flow of the pedagogy part of the text works against the reader, which is problematic since it motivates the rest of the text. Moreover, some important modelling choices and procedures, particularly with respect to the selective pressure metric, are only cursorily described in the methods section. The lack of explanation and detail, especially relative to more simple choices that are seemingly motivated by the authors' own theory, makes it difficult to understand and therefore assess their validity and/or necessity.

    4. Reviewer #3 (Public review):

      Summary:

      This study introduces a new analytical framework to analyze how viral variant frequencies change over time and in different locations. Two examples are given that demonstrate where this approach can be useful and where other approaches can be ambiguous in characterizing novel variants. The authors then demonstrate that the spatiotemporal dynamics of variant frequencies can be used to predict future epidemic growth rates and to investigate how variants differ in immune escape.

      Strengths:

      (1) Examples are provided that make the study accessible for a general audience.

      (2) The authors demonstrate that their approach is predictive both of overall epidemic growth rates and immunological distance between variants.

      (3) The approach introduced in this study can be readily applied to current and future epidemiological challenges that are similar to SARS-CoV-2 with respect to the relative evolutionary timescales wherever there is spatiotemporal heterogeneity in the susceptible population.

      Weaknesses:

      (1) The authors conclude their abstract claiming that their method provides an early signal of epidemic growth. Can this be quantified? Could the authors perform retrospective analyses for sequences available through various cutoff times, identify how early significant new variants are detected, and compare this to other detection methods?

      (2) Analysis depicted in Figure 4 and Figure S9 could be explored further than speculatively attributing weak correlation to declining reporting rates for US states. Exploring how correlation between data and prediction varies over time during the test period might identify periods/events that explain weak correlation overall. The authors could explore predicting growth rates for estimated state prevalences rather than reported cases.

    1. eLife Assessment

      The paper provides a valuable, foundational dataset that will undoubtedly provide substantial value to the cancer research community. The aggregation and harmonization of a broad, rich, multi-omic dataset is convincing and will potentially support further downstream research on GIST, which currently has poor supporting datasets. Support for the accompanying claims is more incomplete, however: the biological demonstration rests on a single gene tested by one approach, several analyses lack sample sizes and multiple-testing correction, the claims made for the LLM assistant are not backed by direct evaluation, and the curated data matrices would ideally be deposited independently of the web.

    2. Reviewer #1 (Public review):<br /> <br /> Summary:

      This tumour type is missing from the big pan-cancer databases, so none of the popular online analysis tools works for it. That's a real gap, and it's the right one to go after. The authors build an online resource that gathers the scattered public molecular datasets for this disease, adds three of their own patient cohorts, ties everything to clinical data, and exposes interactive tools, downloads, and programmatic access so other people can build on it. To show what it does, they take one gene through the whole platform - clinical, gene-expression, protein, single-cell, immune, and drug-response and then test that gene in cell lines. So there are really two things on offer here: a resource and a practical example of using it. They land very differently.

      Strengths:

      The resource is the real contribution, and it's done with care. It covers 37 centres and nearly 2,000 samples across five kinds of molecular data, and the authors are honest about provenance: how they screened datasets in or out, where they recorded the diagnostic codes, and why they dropped ambiguous mixed-tumour collections. The key methodological decision is the right one; every analysis runs inside its own cohort, and the cross-cohort views are explicitly "for looking, not for combining." That's exactly how you should treat heterogeneous public data, and they say so plainly instead of quietly pooling everything. Their three pathologist-confirmed cohorts add genuine independent material, so this isn't a re-skin of data that already existed. And because the code and a public access point are actually available, the reuse claim holds.

      The example is internally consistent, which is what makes it persuasive. The gene reads higher in higher-risk patients across several independent cohorts and in their own protein data, tracks with the disease spreading and recurring, and lines up with worse survival. The single-cell data put it in the dividing cells; the pathway analysis points to proliferation. Three independent data types landing on the same proliferation story are the strongest part of the biology.

      Weaknesses:

      The honest problem is that the entire biological story rests on one gene, tested one way. The lab work is two cell lines with the gene knocked down, showing less growth and migration: there is no rescue to confirm the effect is real, no second gene to show the approach generalises, nothing in a living animal. That earns the modest claim: the resource can point you at a candidate worth testing. It does not earn the headline claim that the platform reliably generates good target hypotheses, because we only ever watch it succeed once. One example illustrates a workflow; it doesn't establish a method.

      Some of the statistics won't survive scrutiny. The clearest case is a perfect separation between treatment-resistant and treatment-sensitive cases from a single immune cell population, reported with no error bars, no check for information leakage, and apparently from very few samples. A perfect result in that setting is almost always overfitting or a small-sample artefact, not a strong classifier. The same pattern shows up elsewhere: small groups, p-values with no effect sizes or error bars, and no correction for the enormous number of features and cohorts being tested across the whole platform. Separately, one drug result is a correlation against a predicted sensitivity score from a model.

      The AI assistant gets far more weight than the evidence supports. Credit where due: the authors are clear and consistent that it only helps interpret and navigate, and never touches the data, the statistics, or the results. That's the correct line to draw, and they hold it. But the assistant itself is never tested, no accuracy numbers, no benchmark, no error analysis, no described way for a human to check what it produces. Calling it something that "fundamentally transforms the user experience" is an assertion, not a finding. And since even the literature feature is admitted not to be a proper systematic review, the prominence of the artificial-intelligence framing runs ahead of what's been shown.

    3. Reviewer #2 (Public review):

      Summary:

      dbGIST appears to be the first dedicated multi-omics resource worldwide that is specifically focused on GIST.

      Strengths:

      The main value of the paper is not simply that the authors collected datasets, but that they built a usable resource around them, with cohort-aware analyses, curated clinical labels, interactive visualizations, downloadable results, selected API access, and an optional LLM-assisted interface. The work is solid, and the database is likely to be useful for GIST researchers interested in target discovery, cross-dataset validation, drug-response hypotheses, and translational follow-up.

      The MCM7 analysis is a reasonable use case. It shows how a user can start from one candidate gene and then move across transcriptomic, proteomic, clinical, single-cell, immune-related, drug-response, and experimental evidence. I do not see this as the main discovery of the paper, but rather as a practical demonstration of what the database can do. That is appropriate for a resource manuscript.

      Weaknesses:

      (1) The authors should make the organization of the platform a little easier to follow. The manuscript refers to five primary omics layers, six omics-focused pages, and eight analytical modules. This structure is understandable after reading the relevant sections, but it may not be immediately obvious to readers. A brief clarification of how the omics layers, web pages, and analytical modules relate to each other would help.

      (2) Since dbGIST is a live web resource, the authors should provide a clear versioning statement. The manuscript should indicate which version of the database corresponds to the analyses and figures reported in the paper, and how future updates will be distinguished from the version evaluated here. This is a small point, but it matters for reproducibility.

      (3) The API function is a strength of the resource, but it is still described rather generally. The authors should give more concrete documentation of what can be accessed through the API, what inputs are required, and what type of output is returned. This could be placed in the supplementary materials. It would make the database more useful for computational users.

      (4) The manuscript should clarify the status of downloadable data. It is clear that figures, source-data tables, and selected derived outputs are available, but it is less clear whether the full processed matrices used internally by the platform are downloadable or only maintained for deployment. This distinction should be stated plainly.

      (5) The statistical reporting in the MCM7 clinical-association analyses needs a little more care. Several p-values are shown across different cohorts and clinical variables. The authors should state whether these are nominal p-values or adjusted p-values. If they are nominal, that is acceptable for a resource demonstration, but the exploratory nature of the analyses should be made clear.

      (6) The ROC analyses for imatinib response should include sample sizes, and confidence intervals for AUC values would be useful if available. Some of the AUC values are high, and without group sizes, it is difficult to judge how stable those estimates are. The authors should avoid implying that these ROC results are validated predictive models.

      (7) The interpretation of MCM7 should be slightly more cautious. MCM7 is a well-known DNA replication and cell-cycle gene, and the single-cell analyses seem to support its association with proliferative cell states. This is biologically consistent, but it also means that MCM7 expression should not be presented as tumour-cell-specific without qualification. The manuscript should frame it mainly as a proliferation-associated signal in the current analysis.

      (8) The drug-response section would benefit from a clearer explanation of the response metric. The authors report correlations between MCM7 expression and predicted response to C6-ceramide, but readers need to know whether the predicted value represents IC50, AUC, sensitivity score, or another metric. The direction of interpretation should also be made explicit, since a negative correlation can mean different things depending on the scoring system.

      (9) The single-cell annotation would be more convincing if the authors provided a compact marker-gene summary for the major cell types in each single-cell cohort. The current description of annotation by source labels, marker inspection, and manual curation is reasonable, but users of the database would benefit from seeing the marker evidence behind the labels.

      (10) The experimental validation section should include a few routine details that are currently not easy to find. The siRNA sequences or target regions, number of biological replicates, statistical tests for the CCK-8 and wound-healing assays, and details of wound-closure quantification should be reported. These additions would make the in vitro part more reproducible.

      (11) The wound-healing result should be interpreted with caution. Since MCM7 knockdown reduces proliferation, reduced wound closure could reflect changes in proliferation, migration, or both. Unless proliferation was controlled during the wound-healing assay, the authors should avoid describing this result as purely migratory.

      (12) The LLM-related claims should remain conservative. The assistant is a useful feature for navigation, plain-language explanation, and user support, especially for clinicians or wet-lab researchers. However, the strongest statements about the LLM transforming interpretation or automating analysis should be toned down. The important point is that the LLM layer helps users interact with the resource, while the numerical analyses come from predefined dbGIST modules.

    4. Reviewer #3 (Public review):

      Summary:

      The dbGist dataset/tool would provide substantial value to the cancer research community.

      Strengths:

      The manuscript presents dbGIST, a dedicated GIST-focused multiomics resource integrating data from 37 centers and ~2k samples across genomics, transcriptomics, proteomics, phosphoproteomics, and single-cell transcriptomics. Given that GIST is virtually absent from major cancer genomics consortia (TCGA, ICGC), this resource fills a genuine gap and represents a valuable contribution to the GIST research community.

      (1) The MCM7 case study effectively demonstrates the platform's utility, linking a resource-derived candidate to survival outcomes.

      (2) The LLM-assisted interface (dbGIST Assistant) is a reasonable addition for accessibility, lowering the barrier for clinicians and wet-lab researchers, who may not always have the skill set required for proper data analysis, especially for a rich and wide dataset like the dataset in question.

      Weaknesses:

      (1) Data deposition (major):

      While the manuscript references public accessions for raw source datasets and provides a GitHub repository for code, it remains unclear where the **curated, harmonized data matrices** - which represent the core value-add of this work - are independently deposited. Access to these processed data appears to depend entirely on the dbGIST web interface and API. The authors should deposit the harmonized matrices in a persistent, general-purpose repository to ensure long-term availability independent of the web platform.

      (2) LLM agent capabilities underspecified:

      The manuscript would benefit from a clearer description of the assistant's capabilities and boundaries. Specifically, what tools or actions are available to the LLM agent? Can it execute code against the underlying data, trigger analytical modules programmatically, or is it limited to natural-language explanation of pre-computed results? Clarifying this would help readers assess the scope of the AI layer and distinguish it from agentic platforms that perform computation on behalf of the user.

    1. eLife Assessment

      This important study presents the development of a model to quantify the dynamics of stress-induced volatile emissions in plants and identify biologically relevant differences in these responses. The model is based on a convincing methodology and represents a starting point for studies of similar responses in other plants or investigations of induced phenotypes in different biological systems.

    2. Reviewer #1 (Public review):

      Summary:

      The manuscript by Waterman et al. describes the development of a mathematical model that quantifies plant volatile emissions dynamics in response to mechanical/biotic stress. Model outputs were based on volatile emission measurements from maize plants using PTR-MS. Modeling revealed differences in emission patterns dependent on the intensity of wounding damage, application of herbivore oral secretions, age of leaf, circadian clock, and genotype. Differences were also observed between different types of volatiles, and the response curves somewhat correlated with expression patterns of biosynthetic genes. Moreover, the model showed priming effects from overlapping response curves upon multiple wounding events.

      Strengths:

      As a non-expert in modeling, this reviewer assesses the work from a broader point of view. Overall, I consider this model to be useful for other researchers to quantify volatile emission dynamics for their plant system. Generating the models does not seem to be overly complicated as long as emissions can be measured with a real-time system such as PTR-MS, which is costly and not available to every lab. The advantage of this approach is that it does not rely on parameters of underlying enzymatic pathways or transport processes. The authors claim that it can be easily applied to other biological responses.

      Weaknesses:

      The manuscript lacks a deeper discussion of how the model can help make predictions of volatile emission dynamics from plants in the greenhouse or field. Can the model be trained and validated with volatile measurements from plants under different environmental conditions? How realistic is this approach given the complexity of a field environment? It would be helpful to provide a better outlook of the application of the model for scientists in the field of plant volatile biology and beyond.

      The authors state that "emissions can be regulated independently of each other" (Line 359). I would assume that regulatory mechanisms in different genotypes are similar but show genotype-specific variation.

    3. Reviewer #2 (Public review):

      This is a study of the dynamics of plant volatile emissions, using a curve-fitting approach to describe salient properties of the dynamics of plant volatile chemicals. The study is interesting and unique in taking this approach. Some of the dynamics uncovered (e.g. lagged emission of many sesquiterpenes) are already well known using less sophisticated approaches, while other properties (diurnal cycles in emission dynamics) are newly uncovered. The approach in general is new for the topic of plant volatile emissions, but curve-fitting is widely used to describe the dynamics or function-valued responses of plants and other organisms. The study thus reads as rather methods-focused, giving tidbits of interesting properties of the dynamics of plant VOCs rather than being structured strongly around clear biological hypotheses. The method seems like a logical and robust way to analyze the dynamics of plant VOCs. I believe the impact of the work will largely depend on whether there are substantial and meaningful outcomes (for herbivores, downstream processes of induction, etc) due to the differences in VOC dynamics described via these methods that would be hard to observe in other ways. If so, there will be a need to adopt robust methods such as this to describe the salient features of those dynamics. At present, I do not believe there is evidence one way or another as to whether the subtle differences in VOC dynamics have large consequences.

      The paper sells itself as describing a new technique for describing response curves generally across biological systems, but it only uses this technique to look at the dynamics of induced plant volatiles. I believe to show general utility of this approach, a wider range of examples of plastic responses to stimuli across organismal groups would be needed. I am, however, convinced that this approach is both novel and useful within the scope in which the examples are shown (i.e. in describing the dynamics of induced plant responses). Some of the text purporting novelty in uncovering shared and divergent responses across the tree of life seems pretty overstated.

      Much of the introductory and discussion text is quite broad, and I wonder if the technique is really meant to be applicable to the specific case that is described (repeated measures of an induced volatile response). Likewise, there has been considerable work in such realms as behavioral science, function-valued traits (e.g. Stinchcombe et al 2012), performance curves (Kingsolver various papers), etc to describe dynamic or variable responses phenomenologically, and there are approaches including GAMs, parametric curve fitting, and other techniques that probably report the same salient features as the approach here. Indeed, there are already statistical techniques to assess the macroevolution of response curves (e.g. Goolsby 2015) and wide discussions as to how to compare function-based responses among organisms (The Functional Phylogenies Group 2012). So in the broad scheme of biology, I am not sure I'm convinced of the novelty of the approach. However, I believe it is novel within the context in which it is used here. The salient part of the methods is that it uses predefined attributes of dynamics (onset, duration, etc) based on a gamma distribution that the researchers (with good reason) believe to be biologically meaningful. This is in contrast to multivariate approaches (e.g. Izem et al 2005) that attempt to find salient dynamic features in a less constrained way.

      I would have liked to see a clear description of model fits (e.g., how much of the variation in the real data is described by the fitted model). This seems important because there are quite a number of constraints placed on model fitting - so presumably when a model blind to those constraints picks unrealistic parameters, that would suggest that the constrained model probably does not fit the data all that well.

      I am curious about the normalization process in the 'normalized emission' that is analyzed throughout the study. Normalization to leaf size makes sense, though I was less clear about L459: "Additionally, values were normalized to the maximum response observed in each experiment, yielding a range of positive values < 1." Why was this needed? Is the 'maximum response observed in each experiment' across all plants/compounds/treatments or within a single plant? In general, is there a way of reporting VOC emission rates in absolute values (e.g. umol / Liter air)? Normalization would presumably not impact most curve properties very much, but it could have effects on 'integral', and the need for within-experiment normalization would suggest a lack of transferability or comparability among datasets from different experiments (at least as regards 'integral'), which is suggested as a major advantage of this approach in the discussion.

    1. eLife Assessment

      This important study presents a cell-based screen for small-molecule activators of GCN2, an eIF2α kinase that regulates the Integrated Stress Response under diverse stress conditions. The authors identify a compound as a potent GCN2 activator with GCN1-independent activity under the tested conditions, providing a new pharmacological tool for probing GCN2 regulation. The revised manuscript provides compelling support for the central conclusions through a clearer description of the screening workflow, extended kinetic analyses, and demonstration that the identified compound causes a GCN2-dependent reduction in protein synthesis.

    2. Reviewer #1 (Public review):

      Summary:

      This manuscript describes a chemical screen for activators of the eIF2 kinase GCN2 (EIF2AK4) in the integrated stress response (ISR). Recently, reported inhibitors of GCN2 and other protein kinases have been shown at certain concentrations to paradoxically activate GCN2. The study uses CHO cells and ISR reporter screens to identify a number of GCN2 activator compounds, including a potent "compound 20." These activators have implications for the development of new therapies for ISR-related diseases. For example, although not directly pursued in this study, these GCN2 activators could be helpful for the treatment of PVOD, which is reported for patients with certain GCN2 loss-of-function mutations. The identified activators are also suggested to engage with the GCN2 directly and can function devoid of GCN1, a co-activator of GCN2.

      Strengths:

      The manuscript appears to be a largely rigorous study that flows in a logical manner. The topic is interesting and significant.

      Weaknesses:

      Portions of the manuscript are not fully clear. There are some experimental presentation and design concerns that should be addressed to support the stated conclusions.

    3. Reviewer #2 (Public review):

      Summary:

      In this manuscript Zhu, Emanuelli and colleagues describe a novel pharmacological activator of the Integrated Stress Response kinase GCN2. The work is conclusive and biochemically solid. This work significantly adds to the pharmacological arsenal targeting the ISR and in particular GCN2.

      Strengths:

      Strong biochemistry, novel molecular activator of GCN2 (GCN1 independent).

      Weaknesses:

      Rationale for the screen not exploited in the results (e.g. pathogenic GCN2 mutants), lots of cell-based read-outs not endogenous.

      Comments on revised version.

      The authors did a great job at addressing my initial critique on their manuscript and consequently I have no further comment.

    4. Reviewer #3 (Public review):

      Summary:

      In this manuscript, the authors describe the results of a high throughput screen for small molecule activators of GCN2. Ultimately, they find 3 promising compounds. One of these three, compound 20 (C20) is of the most interest both for its potency and specificity. The major new finding is that this molecule appears to activate GCN2 independent of GCN1, which suggests that it works by a potentially novel mechanism. Biochemical analysis suggests that each bind in the ATP binding pocket of GCN2, and that at least in vitro C20 is a potent agonist. Structural modeling provides insight into how the three compounds might dock in the pocket and generates testable hypotheses as to why C20 perhaps acts through a different mechanism than other molecules.

      Strengths:

      Of the 3 compounds identified by the authors, C20 is of the most interest, not just for its intriguing mechanistic distinction as being GCN1-independent (shown genetically in two distinct cell lines, CHO and 293T, and in contrast to other GCN2 activators) but also for its potency. Ultimately, C20 might be a tool for providing mechanistic insight into the details of GCN2 activation and regulation and could be exploited therapeutically.

      Weaknesses:

      The chief limitation of this work is that the experiments exploring the effects of C20 on ISR output in cells are limited, so how useful these compounds are both experimentally and therapeutically remains to be determined.

      Comments on revised version.

      The authors have satisfactorily addressed my comments. A more extensive analysis of UPR signaling in cells (transcription and cell death in particular) would have further strengthened the paper, but that can be left to future work.

    5. Author response:

      The following is the authors’ response to the original reviews

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This manuscript describes a chemical screen for activators of the eIF2 kinase GCN2 (EIF2AK4) in the integrated stress response (ISR). Recently, reported inhibitors of GCN2 and other protein kinases have been shown at certain concentrations to paradoxically activate GCN2. The study uses CHO cells and ISR reporter screens to identify a number of GCN2 activator compounds, including a potent "compound 20." These activators have implications for the development of new therapies for ISR-related diseases. For example, although not directly pursued in this study, these GCN2 activators could be helpful for the treatment of PVOD, which is reported for patients with certain GCN2 loss-of-function mutations. The identified activators are also suggested to engage with the GCN2 directly and can function while devoid of GCN1, a co-activator of GCN2.

      Strengths:

      The manuscript appears to be a largely rigorous study that flows in a logical manner. The topic is interesting and significant.

      Weaknesses:

      Portions of the manuscript are not fully clear. Some experimental presentation and design concerns should be addressed to support the stated conclusions.

      We thank the reviewer for their supportive comments. We agree that portions of the manuscript were not fully clear and that some aspects of the experimental presentation and design required clarification.

      To address this, we have revised the manuscript to make the experimental logic more transparent. In particular, we now explain more clearly the rationale for the screening strategy, including the use of histidinol as a canonical GCN2 activator, latrunculin A as a modulator of PPP1R15A-mediated eIF2α dephosphorylation, and tunicamycin as a PERK-dependent ER-stress control. We also clarify why a submaximal concentration of histidinol was used: this was intended to reveal compounds that enhance ISR signalling when GCN2 is partially activated.

      We have clarified the use of the two ATF4 reporter systems. The ATF4–NanoLuc reporter was used for sensitive primary screening, whereas the ATF4–luc2 reporter was used as a more stringent orthogonal assay to prioritise robust ISR activators. We now state explicitly why some initial hits were not retained after testing in the second reporter line, and why the NanoLuc system was subsequently used again for mechanistic experiments.

      Finally, we have revised the presentation of the orthogonal validation steps to make clearer how they support the stated conclusions. These include assays designed to distinguish GCN2-dependent ISR activation from indirect activation through ER stress, additional analysis of GCN2 dependence, and clearer interpretation of biochemical and docking data.

      We hope that these revisions address the reviewer’s concern that the experimental design and data presentation needed to be made clearer in order to support the conclusions.

      Reviewer #2 (Public review):

      Summary:

      In this manuscript, Zhu, Emanuelli, and colleagues describe a novel pharmacological activator of the Integrated Stress Response kinase GCN2. The work is conclusive and biochemically solid. This work significantly adds to the pharmacological arsenal targeting the ISR and, in particular, GCN2.

      Strengths:

      Strong biochemistry, novel molecular activator of GCN2 (GCN1 independent).

      Weaknesses:

      The rationale for the screen is not exploited in the results (e.g., pathogenic GCN2 mutants), and lots of cell-based read-outs are not endogenous.

      We thank this reviewer for their positive assessment of the work. We address the three major concerns in turn below.

      Major points

      (1) Regarding the justification of the work. Since the authors justify the screen for GCN2 activators with loss-of-function mutants associated with diseases, it would be of interest to evaluate whether the best compounds identified in the study are indeed able to prompt activation of those mutants (or at least of the most prevalent). This approach could actually go in parallel with the docking experiments carried out in the last figure of the manuscript, where mutants could be modelized as well.

      To address this point, we tested whether the lead compounds could activate disease-associated GCN2 variants linked to pulmonary veno-occlusive disease. In contrast to GCN2iB, the new compounds did not activate these variants. We now state this explicitly in the manuscript, thereby clarifying that although the compounds identify a new mode of GCN2 activation, they do not rescue the pathogenic GCN2 variants tested here.

      Results

      “Moreover, in contrast to GCN2iB (17), the current compounds did not activate disease-associated GCN2 variants linked to PVOD [data not shown].”

      (2) The compounds are only tested using « artificial » proximal signaling outputs. It would be interesting to evaluate whether the best identified compounds are capable of prompting endogenous eIF2alpha phosphorylation in cellular models.

      We thank the reviewer for this suggestion. Detecting eIF2α phosphorylation following activation of GCN2 is technically challenging and typically produces weaker signals compared to activation of other ISR kinases, such as PERK (e.g. by thapsigargin). For this reason, many studies rely on downstream reporter assays to monitor GCN2 activity. To address the reviewer’s concern, we have now included an orthogonal readout of ISR activation by assessing global translation using a puromycin incorporation assay. Using this approach, we show that compound 20 significantly reduces translation, and importantly, this effect is attenuated in GCN2-deficient cells, supporting a GCN2-dependent mechanism.

      Results

      “Studies with compound 18 were limited by poor aqueous solubility; therefore, time‑course analyses focused on compounds 20 and 21. To assess ISR activation over an extended period, live‑cell luciferase measurements were performed using CHO cells stably expressing an ATF4::Nanoluc-PEST reporter. Both compounds elicited maximal reporter activation between 6 and 8 h (Figure S1A&B). Compound 20, but not 21, induced a significant GCN2‑dependent reduction in mRNA translation, as measured by puromycin incorporation, with a progressive effect observed up to 7 h (Figure S1C-F).”

      (3) Other GCN2 activators (other than GCN2iB, e.g., HC-7366) were recently identified. In this context, it would be of interest to carry out a small benchmarking study to evaluate how the compounds identified in the current study perform against the previously identified molecules.

      We thank the reviewer for this suggestion. In response, we obtained HC-7366 and assessed its activity alongside our compounds in the CHO ATF4::NanoLuc reporter assay. In this system, compound 20 demonstrated greater potency than HC-7366 (see reviewer figure below). However, we note that HC-7366 showed relatively limited activity in CHO cells in our hands, despite previously reported strong effects in other cellular systems and in vivo models. This context-dependent activity makes direct benchmarking difficult. Accordingly, we have included this comparison in the revised manuscript and discuss this limitation in the Discussion.

      Discussion

      “We also evaluated the reported GCN2 activator HC-7366 in our CHO ATF4::NanoLuc reporter system. In this context, HC-7366 showed limited activity relative to compound 20, despite its reported efficacy in other cellular systems and in vivo (data not shown). This highlights potential context dependence in small‑molecule activation of GCN2 and limits direct cross-study comparison.”

      Author response image 1.

      ISR activation by compound 20 and GC-7366 in CHO cells

      Normalised fold-change in ATF4 signal in CHO ATF4::NanoLuc reporter cells treated for 19 hours with Compound 20 or HC-7366. DMSO was used as vehicle control. (representative experiment, mean ± SEM, n=3 technical replicates).

      Reviewer #3 (Public review):

      Summary:

      In this manuscript, the authors describe the results of a high-throughput screen for small-molecule activators of GCN2. Ultimately, they find 3 promising compounds. One of these three, compound 20 (C20), is of the most interest both for its potency and specificity. The major new finding is that this molecule appears to activate GCN2 independent of GCN1, which suggests that it works by a potentially novel mechanism. Biochemical analysis suggests that each binds in the ATP-binding pocket of GCN2, and that at least in vitro, C20 is a potent agonist. Structural modeling provides insight into how the three compounds might dock in the pocket and generates testable hypotheses as to why C20 perhaps acts through a different mechanism than other molecules.

      We agree that GCN1-independent activation suggests a potentially distinct mechanism of action. While we are currently unable to define the mechanistic basis underlying the GCN1-independence of compound 20, prior work provides some relevant context. Recent studies have shown that the ATP-competitive modulator GCN2iB can activate GCN2 independently of GCN1 under specific conditions, notably in the context of the GCN2 E26A mutant [Carlson, 2023]. This observation raises the possibility that, under certain conditions, engagement of the kinase domain, potentially via the ATP-binding pocket, may bypass the requirement for GCN1. However, in our system, we did not observe GCN1-independent activation with GCN2iB at the concentrations tested. This discrepancy may reflect a narrow or context-dependent window for such activity, or differences between wild-type and mutant GCN2. These findings suggest that GCN1-independent activation of GCN2 may occur under specific conditions or with distinct classes of compounds, although further work will be required to define the underlying mechanism for compound 20. We have added the following to the main text:

      Discussion

      “Recent work suggests that the ATP-competitive modulator GCN2iB can activate GCN2 independently of GCN1 under specific conditions using a GCN2 E26A mutant (9). In our hands, we did not observe GCN1‑independent activation with GCN2iB at the concentrations tested. This discrepancy may reflect a narrow concentration window for GCN1‑independent activation or context‑dependent effects of the E26A mutation. These findings raise the possibility that GCN1‑independent activation of GCN2 may occur under specific conditions or with distinct classes of compounds.”

      Strengths:

      Of the 3 compounds identified by the authors, C20 is the most interesting, not just for its intriguing mechanistic distinction as being GCN1-independent (shown genetically in two distinct cell lines, CHO and 293T in Figure 4, and in contrast to other GCN2 activators) but also for its potency. In in-cellulo assays, compound 21 appears as more of an ISR enhancer than an activator per se, and although compound 18 and compound 21 lead to upregulation of the ISR targets (Figure 2), that degree of upregulation is probably not significantly different from that induced by those compounds in Gcn2-/- cells. For C20, the effect appears stronger (although it is unclear whether the authors performed statistical analysis comparing the two genotypes in Figure 2D). In Figure 3, only C20 activates the ISR robustly in both CHO and 293T. Ultimately, C20 might be a tool for providing mechanistic insight into the details of GCN2 activation and regulation, and could be exploited therapeutically.

      Prompted by this suggestion, we assessed C20 in two additional commonly used cell lines: human colon carcinoma HCT116 cells and African green monkey COS7 cells. C20 showed no activity in these models. In contrast, primary mesothelioma cells (Mesobank T12) exhibited robust PPP1R15A induction in response to the compound. The following text has been added to the manuscript.

      Results

      “We went on to examine downstream cellular consequences of GCN2 activation in multiple models. While compounds did not induce detectable ISR signalling in HCT116 or COS‑7 cells under the conditions tested, induction of PPP1R15A was observed in Mesobank T12 primary mesothelioma cells, indicating context-dependent biological responses… [data not shown].”

      Weaknesses:

      There are some limitations to the existing work. As the authors acknowledge, they do not use any of the compounds in animals; their in vivo efficacy, toxicity, and pharmacokinetics are unknown. But even in the context of the in cellulo experiments, it is puzzling that none of the three compounds, including C20, has any effects in HeLa cells when Neratinib does. It's beyond the scope of this paper to address definitively why that is, but it would at least be reassuring to know that C20 activates the ISR in a wider range of cells, including ideally some primary, non-immortalized cells. In addition, the ISR is a complex, feedback-regulated response whose output varies depending on the time point examined. The in cellulo analysis in this paper is limited to reporter assays at 18 hours and qRT-PCR assays at 4 and 8 hours. A more extensive examination of the behaviour of the relevant ISR mRNAs and proteins (eIF2, ATF4, CHOP, cell viability, etc.) for C20 across a more extensive time course would give the reader a clearer sense of how this molecule affects ISR output.

      We thank the reviewer for this insightful suggestion. To address the need for a more comprehensive assessment of ISR signalling, we have extended our analysis across a broader time course and incorporated additional functional readouts. Using the ATF4-NanoLuc reporter, compounds 20 and 21 exhibit peak ISR activation at approximately 6-8 h in wild-type cells. In parallel, we assessed global mRNA translation using puromycin incorporation and found that compound 20, but not compound 21, induces a progressive reduction in translation over this period, which is dependent on GCN2. While we agree that direct measurement of upstream ISR markers such as eIF2α phosphorylation can be informative, detection of GCN2-mediated eIF2α phosphorylation is technically challenging and often less robust than activation of other ISR kinases (e.g. PERK). For this reason, we have prioritised orthogonal downstream functional readouts, including reporter activity and translational output, to capture ISR pathway engagement. These additional data provide a clearer picture of the kinetics and functional consequences of compound-induced ISR activation and have been incorporated into the revised manuscript.

      Results

      “Studies with compound 18 were limited by poor aqueous solubility; therefore, time‑course analyses focused on compounds 20 and 21. To assess ISR activation over an extended period, live‑cell luciferase measurements were performed using CHO cells stably expressing an ATF4::Nanoluc-PEST reporter. Both compounds elicited maximal reporter activation between 6 and 8 h (Figure S1A&B). Compound 20, but not 21, induced a significant GCN2‑dependent reduction in mRNA translation, as measured by puromycin incorporation, with a progressive effect observed up to 7 h (Figure S1C-F).”

      I also find it a bit strange that the authors describe C20 as "demonstrat(ing) weak inhibition of ... PKR" - the measured IC50 is ~4 μM, which is right around its EC50 for GCN2 activation. This raises the confounding possibility that C20 would simultaneously activate GCN2 while inhibiting PKR. While perhaps inhibition of PKR is not relevant under the conditions when GCN2 would be activated either experimentally or therapeutically, examining in cells the effects of C20 on GCN2 and PKR across a dose range would shed light on whether this cross-reactivity is likely to be of concern.

      We thank the reviewer for highlighting compound 20 as the most interesting lead compound and for recognising its apparent ability to activate GCN2 independently of GCN1. The reviewer identified several limitations relating to cell-type specificity, the temporal behaviour of ISR activation, and possible PKR cross-reactivity.

      In response to the concern about cell-type specificity, we tested compound 20 in additional cellular models. Compound 20 did not induce detectable ISR signalling in HCT116 or COS-7 cells under the conditions tested, consistent with the reviewer’s observation that activity is not universal across cell types. However, we observed induction of PPP1R15A in primary Mesobank T12 mesothelioma cells. We have therefore revised the manuscript to present compound 20 activity as cell-context dependent rather than broadly generalisable across all cell types.

      To address the reviewer’s concern that the ISR output was examined only at limited time points, we extended the time-course analysis for compounds 20 and 21. Using live-cell ATF4::NanoLuc reporter measurements, both compounds showed maximal reporter activation at approximately 6-8 hours. We also measured translational output by puromycin incorporation and found that compound 20, but not compound 21, caused a progressive GCN2-dependent reduction in translation. These data provide a clearer view of the kinetics and functional consequences of compound 20-mediated ISR activation.

      The reviewer also noted that compound 20 inhibits PKR in vitro at concentrations close to those required for GCN2 activation in cells. We agree that this is an important potential liability. Because PKR signalling was not robustly or reproducibly inducible in our CHO-based reporter system, we were unable to perform a reliable cellular dose-response analysis of PKR engagement in the present study. We have therefore revised the Discussion to acknowledge kinase cross-reactivity, including possible PKR inhibition, as an important limitation and an issue for future development of this chemical series.

      Finally, we have moderated our mechanistic interpretation of compound 20. Although the data support direct engagement of GCN2 and suggest a mechanism distinct from canonical GCN1-dependent activation, we now discuss GCN1-independent activation more cautiously and in the context of prior reports that GCN2iB can display GCN1-independent activity under specific experimental conditions.

      Discussion

      “While the functional relevance of PKR inhibition in our cellular systems is uncertain, these observations highlight the potential for kinase cross-reactivity, which will be important to address in future studies.”

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):<br /> (1) The description of the chemical screen for Gcn2 activators is not sufficiently clear and detailed. a) Briefly and early on provide the rationales (modes of action) for using histindinol and latruculin A. Explain further the rationale in Figure 2, outlining the purpose for the combined compound + submaximal dose of histindinol.

      The text has been amended.

      Results

      “Histidinol activates GCN2 by inhibiting histidyl‑tRNA synthetase, leading to the accumulation of uncharged tRNAHis. This uncharged tRNA binds to GCN2, relieving its autoinhibition and activating the kinase (35). Latrunculin A sequesters G‑actin, thereby inhibiting PPP1R15A activity (36, 37). Tunicamycin inhibits protein glycosylation in the endoplasmic reticulum (ER), resulting in activation of PERK (38).”

      “This submaximal concentration was used to allow detection of compounds that enhance ISR signalling when GCN2 is partially activated.”

      b) What is unique about the second CHO: ATF4-luc2 reporter line? Why do only 89 out of the original 130 compounds induce the ISR in this line versus the original CHO: ATF4-Nanoluc cell line? This is confusing for the reader about how compounds were triaged for characterization.

      The ATF4‑Nanoluc and ATF4‑luc2 reporter lines differ only in the luciferase used, but this has important practical consequences. The Nanoluc reporter is substantially more sensitive, so it was used for the primary screen to detect even weak ISR activation. The luc2 reporter has lower sensitivity and a narrower dynamic range, making it a more stringent orthogonal assay. As a result, not all hits from the Nanoluc screen (130 compounds) reproduced in the luc2 line; the 89 compounds retained are those that robustly activate the ISR under these more stringent conditions. This step was therefore used to prioritise stronger, more reproducible activators for downstream characterisation.

      Results

      “While primary screening was performed in ATF4‑Nanoluc lines for maximal sensitivity, hits were subsequently re-tested in a second CHO ATF4::luc2 reporter line as a more stringent orthogonal assay to prioritise robust ISR activators. Of the 130 hits identified in the sensitive Nanoluc screen and passing early toxicity assessment, 89 were confirmed in the luc2 assay, consistent with enrichment for higher-amplitude ISR activators under more stringent detection conditions.”

      c) The study uses a second CHO reporter line in the flow scheme (CHO:ATF4-luc2) and then switches back to an ATF4-Nanoluc line to establish GCN2 dependence. What is the rationale for switching back to the original reporter line?

      The luc2 reporter line was used as a more stringent, orthogonal validation step to prioritise robust ISR activators. For subsequent mechanistic studies, including assessment of GCN2 dependence, we returned to the ATF4‑Nanoluc line because its higher sensitivity and simpler single‑reagent assay format are better suited to multi‑point measurements and comparative analyses. In effect, the luc2 reporter was used for triage, whereas the Nanoluc system was retained for mechanistic characterisation and downstream screening.

      Results

      “In subsequent mechanistic studies, the ATF4‑Nanoluc reporter was again used to take advantage of its higher sensitivity and simpler assay format for multi‑condition comparisons.”

      d) The rationale for the first orthogonal screen described in the results section to identify inducers of ER stress is not clearly explained. The compounds were already determined to be dependent on GCN2 prior to this test, and one would have thought that this criterion would have covered ER stress and alternative eIF2 kinase activators.

      We agree with the reviewer that, in principle, establishing GCN2 dependence should reduce the likelihood of capturing compounds acting through alternative eIF2α kinases. However, we performed this orthogonal ER stress screen to address two practical considerations. First, high‑throughput screening is inherently prone to false positives, as it is typically conducted at a single concentration and time point, and compound libraries may contain degraded or chemically inconsistent material. We therefore used a lower‑throughput, more controlled ER stress assay with freshly sourced compounds and additional readouts (e.g. CHOP and XBP1) to improve confidence in the hits. Second, despite prior evidence of GCN2 dependence, ER stress signalling via PERK converges on the same downstream endpoints: eIF2α phosphorylation and ATF4 induction. We therefore wished to explicitly exclude compounds that activate the ISR indirectly via ER stress. In practice, this proved important, as the orthogonal assay did identify compounds that induced ER stress, which we subsequently excluded from the lead set.

      Results

      “Although hits were prioritised for GCN2 dependence, we performed an additional orthogonal screen to exclude compounds that activate the ISR indirectly via ER stress, which converges on the same downstream outputs.”

      (2) A major point of the manuscript is that there is GCN1 independence for the small molecule activation of GCN2, and this has not yet been reported. One report for this GCN1 independence is reference 9 [Carlson … Wek 2023] (Figure 5). In this report, low doses of GCN2iB that can activate GCN2 (although by the present manuscript at much lower levels than the identified new compounds) induce ATF4 expression in cells expressing an E26A mutant of GCN2 that is suggested to negate GCN1 binding and enhancement of GCN2 activity. Halofuginone induction of ATF4 expression was thwarted by the GCN2 E26A mutant.

      We thank the reviewer for highlighting this important point. We agree that Carlson et al. (2023) suggest that, under certain conditions, GCN2iB can activate GCN2 independently of GCN1 using the E26A mutant. In our experiments, however, we did not observe GCN1‑independent activation with GCN2iB under the conditions tested, i.e. similar low doses. One possible explanation is that the GCN1‑independent activity reported by Carlson et al. occurs only within a narrow concentration range; their observations were made at very low compound concentrations, whereas higher concentrations may engage additional regulatory mechanisms. In our study, we used concentrations optimised for robust ISR activation, which may mask such effects. We also note that the E26A mutation (E18A in yeast), originally identified by two‑hybrid analysis, disrupts the GCN2-GCN1 interaction but may not completely eliminate all modes of functional coupling under all conditions. Taken together, these observations raise the possibility that GCN1‑independent activation represents a context‑dependent mechanism that may be unmasked only under specific experimental conditions or by particular classes of compounds.

      We have revised the Discussion to acknowledge this prior report explicitly and to clarify how our findings relate to it.

      Discussion

      “Recent work suggests that the ATP-competitive modulator GCN2iB can activate GCN2 independently of GCN1 under specific conditions using a GCN2 E26A mutant (9). In our hands, we did not observe GCN1‑independent activation with GCN2iB at the concentrations tested. This discrepancy may reflect a narrow concentration window for GCN1‑independent activation or context‑dependent effects of the E26A mutation. These findings raise the possibility that GCN1‑independent activation of GCN2 may occur under specific conditions or with distinct classes of compounds.”

      (3) The authors state that the ISR was exaggerated in Ppp1r15a KO cells. It would be helpful to include statistical analyses to support this statement.

      Thank you for enabling us to be more precise. New text added:

      Results

      “Activation of the ISR by tunicamycin was exaggerated in the Ppp1r15a<sup>-/-</sup> cells owing to their defective dephosphorylation of eIF2a (wild type vs Ppp1r15a<sup>-/-</sup>, p<0.05).”

      (4) The results state that Chop and Ppp1r15a mRNAs were measured following 4 hours of treatment with compound 18, 20, or 21, but Figure 2D shows treatment from 0 to 8 hours? It appears that the compound still induces these mRNAs in GCN2 KO cells, possibly with delayed kinetics. A lengthened time course study would help determine if this is indeed the case.

      We thank the reviewer for this careful observation. To address the reviewer’s point regarding delayed or GCN2‑independent signalling, we have extended our analysis using compounds 20 and 21, which are more tractable experimentally (solubility). Using the ATF4‑Nanoluc reporter, both compounds show peak ISR activation at ~6–8 h in wild-type cells over an extended time course. In parallel, functional readouts of mRNA translation (puromycin incorporation) demonstrate that compound 20, but not 21, induces a progressive, GCN2‑dependent reduction in translation over this period. These clarify the temporal aspects of signalling by these two compounds.

      Results

      “Studies with compound 18 were limited by poor aqueous solubility; therefore, time‑course analyses focused on compounds 20 and 21. To assess ISR activation over an extended period, live‑cell luciferase measurements were performed using CHO cells stably expressing an ATF4::Nanoluc-PEST reporter. Both compounds elicited maximal reporter activation between 6 and 8 h (Figure S1A&B). Compound 20, but not 21, induced a significant GCN2‑dependent reduction in mRNA translation, as measured by puromycin incorporation, with a progressive effect observed up to 7 h (Figure S1C-F).”

      Legend

      “Supplementary Figure S1. Kinetics of responses to compounds 20 and 21

      (A-B) Wild-type CHO cells stably expressing the ATF4::nanoLuc-PEST reporter were treated with Nano-Glo and either (A) 13mM compound 20 or (B) 13mM compound 21. Median bioluminescence (fold change normalised to DMSO control) ± 95% confidence. Representative experiment (n=4 technical repeats). (C-F) Representative immunoblot of lysates from wild-type or Eif2ak4<sup>-/-</sup> CHO cells treated with 10μM compound 20 or 7.5μM 21 for the indicated times. Immediately before harvesting, cells were treated with 10μg/mL puromycin to label newly synthesised polypeptides. “-“ indicates cells not incubated with puromycin. “U” cells were treated with puromycin but without test compound. “CHX” represents the cycloheximide control (100μg/mL). Molecular size in kDa. (E-F) Quantification of puromycinylated proteins normalised to GAPDH. Mean ± SEM. CHO WT (black) and Eif2ak4<sup>-/-</sup> cells (turquoise. N = 4 independent experiments. Two-way ANOVA with Šídák's multiple comparisons test; ***: p ≤ 0.001.”

      (5) In the section describing the differences between cell lines in the ability of compounds to induce the ISR, this is difficult for the reader to interpret, as no controls are included. How does histidinol (or other canonical inducers of the ISR) behave in the three reporter assays (CHO, 293T, and HeLa)?

      As requested, we now provide ATF4::Nanoluc reporter activation (3mM, 20 hours because of this drug’s slow kinetics)

      Results

      “To benchmark ISR activation in these models, each cell type was treated with 3mM histidinol (Figure S2). Reporter activation was most robust in CHO cells, followed by 293T cells, then HeLa cells.”

      Discussion

      “Moreover, histidinol-induced ISR activation showed a clear hierarchy across cell lines, with CHO cells being the most responsive and HeLa cells the least.”

      Legend

      “Supplementary Figure S2. Cell-type differences in response to histidinol

      Fold-change of ATF4::NanoLuc reporter signal in HEK293T, HeLa and CHO cells transiently transfected with reporter and treated for 20 hours with 3mM histidinol. Fold-change calculated relative to vehicle control. Mean ± SEM).”

      We thank the reviewer for this important question. However, we respectfully disagree that a direct correspondence between the concentrations required for target engagement in the BRET assay and for ISR activation in functional assays should necessarily be expected. BRET (including NanoBRET) is a target engagement assay that measures compound binding to the protein in intact cells, typically by competition with a labelled tracer, and thus reports on apparent intracellular affinity and occupancy rather than downstream biological effect (Robers 2019, PMID 30519940). By contrast, ISR activation is a functional readout that reflects amplification through signalling networks, and can be influenced by multiple additional variables including pathway non-linearity, feedback, and kinase regulation. Consequently, it is well established that potencies derived from target engagement assays do not always align with those measured in functional assays. For example, intracellular kinase profiling studies using NanoBRET have demonstrated systematic potency offsets between binding/engagement measurements and downstream cellular activity, arising from factors such as intracellular ATP competition and pathway context (PMID Capener 2026, PMID 41495225). More generally, target engagement assays provide a quantitative measure of binding, whereas functional assays measure biological outcome, and these readouts need not coincide because they capture distinct aspects of a compound’s mechanism of action. Accordingly, we interpret our BRET data as evidence of direct interaction with GCN2 in cells, rather than as a predictor of the concentration required to activate the ISR. The observation that higher concentrations are required in the BRET assay is therefore not unexpected and does not argue against a requirement for kinase-domain engagement in ISR activation. Instead, it reflects the different mechanistic endpoints captured by the two assay formats.

      We will clarify this point explicitly in the revised manuscript.

      Results

      “The concentrations required to detect target engagement in NanoBRET assays did not directly mirror those required for ISR activation, reflecting the distinction between ligand binding and downstream pathway output.”

      (7) In Figure 5E, the authors suggest that compounds 18 and 20 are non-competitive inhibitors of GCN2 since the Vmax increases with increasing ATP concentration. What is the Km for ATP in the absence or presence of compound 18 or 20? It would be helpful to include progress curves as supplementary data to support the Vmax plots in Fig. 5E. Consider providing more specific units (currently arbitrary units) for the y-axis.

      We thank the reviewer for this insightful comment and agree that our original wording overstated the mechanistic interpretation of these data. In particular, the use of the term “non‑competitive” is not well supported by the current analysis and may be misleading, especially given that our data are consistent with binding within or proximal to the ATP-binding pocket. We have therefore revised the text to remove this designation and instead describe the data more conservatively in terms of changes in apparent Vmax, without assigning a specific inhibition mechanism. With respect to kinetic analysis, we agree that full determination of K<sup>m</sub> values and inclusion of progress curves would provide a more rigorous mechanistic interpretation. However, given the primary focus of this manuscript on identifying and characterising small‑molecule activators of GCN2 in cells, we believe that a detailed steady‑state kinetic analysis would be beyond the scope of the current study. We have therefore moderated our conclusions accordingly and now present these data as preliminary kinetic observations rather than definitive evidence of inhibition modality.

      Results

      “Compounds 18 and 20 altered the apparent kinetic parameters of GCN2, including an increase in the observed V<sub>max</sub>; however, these data do not allow assignment of a specific inhibition modality.”

      (8) The full-length GCN2 assay presented in Figure 5G appears to be unresponsive to uncharged tRNA, a known regulator of GCN2. The statement that compound 20 induces eIF2 phosphorylation to a greater extent than tRNAs is true for the in vitro assay, but arguably is because the in vitro assays do not recapitulate the in vivo arrangement.

      We thank the reviewer for this comment. We respectfully disagree that the assay is unresponsive to uncharged tRNA. In our hands, GCN2 does exhibit activation in response to tRNA; however, the magnitude of this effect is modest (~4‑fold) compared to the substantially stronger activation observed with compound 20 (~40‑fold). As a result, the tRNA response can appear compressed when both are plotted on the same scale. We agree with the reviewer that the in vitro assay does not fully recapitulate the in vivo regulatory environment, where factors such as GCN1 and ribosome association are known to potentiate GCN2 activation. This limitation likely explains the relatively weaker response to tRNA under our assay conditions and was a key motivation for incorporating cellular assays in our study. Interestingly, the marked difference in activation magnitude between tRNA and compound 20 in vitro raises the possibility that compound-mediated activation may, at least in part, bypass regulatory features that normally constrain GCN2 activity in a GCN1‑dependent manner. While we have not directly tested this hypothesis, we will temper the wording and include this as a speculative point in the Discussion.

      Results

      “Of note, uncharged tRNA produced a modest (~4‑fold) activation of GCN2 under these conditions, whereas compound 20 induced substantially greater (~40‑fold) activation.”

      Discussion

      “The markedly greater activation observed with compound 20 compared with uncharged tRNA in vitro raises the possibility that such compounds may partially bypass regulatory constraints on GCN2 activation, including those normally mediated by GCN1.”

      (9) Using purified GCN2 kinase domain at low ATP concentrations (10 μM), compounds 18 and 20 were shown not to inhibit GCN2 up to concentrations of 3 μM. In previous assays, much higher concentrations of compounds 18 and 20 were used to inhibit GCN2. Why were different concentrations used? This makes this interpretation of this data difficult for the reader to draw conclusions.

      We thank the reviewer for this comment and agree that the use of different concentration ranges across assays may not have been sufficiently clear. The kinase‑domain assay performed at low ATP (10 μM) was specifically designed to assess whether compounds 18 and 20 have a propensity to inhibit GCN2 under conditions that sensitise detection of ATP‑competitive effects and facilitate comparison with related eIF2α kinases. This assay was therefore optimised for detecting inhibition, rather than activation. In contrast, the higher concentrations used in other experiments were selected to robustly measure ISR activation in cellular or full‑length protein contexts, where higher compound exposure is required to observe downstream signalling outputs. These two assay systems therefore address distinct mechanistic questions—targeting inhibition under controlled biochemical conditions versus activation in more complex functional settings—and are not directly comparable in terms of concentration–response relationships. We will revise the manuscript to clarify this distinction and to emphasise that the kinase‑domain assay was not intended to define the activation potency of the compounds.

      Results

      “This assay was performed at low ATP concentrations to sensitise detection of ATP-competitive inhibition and was not optimised to detect compound-mediated activation of GCN2.”

      (10) In silico docking studies support the binding of compound 20 in the ATP-binding pocket of the GCN2 kinase domain. How is this compatible with the stated ATP non-competitive mechanism?

      We agree with the reviewer that our previous description was misleading. The designation of compounds 18 and 20 as “ATP non‑competitive” is not supported by the available data and is inconsistent with the docking results suggesting binding within the ATP‑binding pocket. We have therefore revised the manuscript to remove this terminology and to describe the kinetic behaviour more cautiously, without assigning a specific mode of inhibition.

      (11) The study does not appear to feature biological assays demonstrating the effects of GCN2 activation. For example, does compound 20 reduce translation or growth of cells in a GCN1/GCN2-dependent manner, or do the compounds overcome PVOD mutations akin to the authors' Hum Mol Genet 2024 Aug 18;33(17):1495-1505 article?

      We thank the reviewer for this suggestion. We agree that defining downstream biological consequences of GCN2 activation is an important goal. We did explore this using several cellular systems; however, these effects were context-dependent and not consistently observed across models. Specifically, compounds did not induce a detectable ISR in HCT116 or COS‑7 cells under the conditions tested, despite responsiveness of these systems to canonical activators such as histidinol. In contrast, we did observe induction of PPP1R15A in primary mesothelioma cells, indicating that biological responses can be elicited in certain cellular contexts. We also tested whether these compounds could rescue disease-associated GCN2 variants linked to PVOD, as previously reported for GCN2iB, but did not observe activation of these mutants. These findings suggest that while the compounds robustly activate GCN2 signalling in reporter assays, downstream biological outputs are context-dependent and may require specific cellular conditions or co-factors. Given the variability across systems, we have limited our conclusions to ISR activation and have not generalised broader biological effects. We will clarify this point in the revised manuscript.

      Results

      “We went on to examine downstream cellular consequences of GCN2 activation in multiple models. While compounds did not induce detectable ISR signalling in HCT116 or COS‑7 cells under the conditions tested, induction of PPP1R15A was observed in Mesobank T12 primary mesothelioma cells, indicating context-dependent biological responses. Moreover, in contrast to GCN2iB (17), the current compounds did not activate disease-associated GCN2 variants linked to PVOD [data not shown].”

      (12) For the control of neratinib and other activators linked with ATP binding that are suggested to be dependent on GCN1 in Fig. 4, include reporter induction by drug treatment in Gcn2-/- cells. It would be helpful to be clear in the list of compounds between those suggested to be direct activators versus those that may create stress that leads to GCN2 activation.

      We thank the reviewer for this suggestion. In response, we have performed additional reporter assays in both WT and GCN2<sup>-/-</sup> cells to assess the dependence of drug-induced ISR activation on GCN2. These experiments reveal clear GCN2 dependence for reporter induction upon treatment with sunitinib, NXP800, WEE1-in-4, Debio0123, gefitinib, and erlotinib, as signal is markedly reduced in GCN2⁻/⁻ cells compared to WT.

      In contrast, dovitinib and AZD1775 show less clear dependence, with relatively low reporter signal even in WT cells (notably lower than observed in Fig. 4C), limiting interpretation. Interestingly, dabrafenib induces stronger reporter activity in GCN2<sup>-/-</sup> cells than in WT, indicating that its effects are independent of GCN2 and may reflect activation of alternative stress or signalling pathways.

      Results

      “To further assess the mechanism of compound-induced ISR activation, we evaluated reporter responses in GCN2-deleted cells (Supplementary Figure S3). Several compounds, including sunitinib, NXP800, WEE1-in-4, Debio0123, gefitinib, and erlotinib, showed reduced reporter activity in GCN2-deficient cells, consistent with GCN2-dependent activation. In contrast, dovitinib, AZD1775 and dabrafenib produced weaker or inconclusive responses even in the paired wild-type lines, limiting analysis. These data allow us to distinguish compounds consistent with direct or GCN2-dependent activation from those more likely to induce ISR indirectly through cellular stress upstream of GCN2. These findings support a distinction between compounds that activate the ISR through GCN2-dependent mechanisms and those that likely act indirectly via alternative stress pathways.”

      Legend

      “Supplementary Figure S3. GCN2-dependence of ISR activation by putative GCN2 agonists

      Normalised fold-change in ATF4 signal in CHO WT (purple) and Gcn2-/- (blue) ATF4::NanoLuc reporter cells treated for 19 hours with a panel of ATP-competitive kinase inhibitors reported to activate GCN2 (at 1 and 3µM, sunitinib used at 3 and 10µM, AZD1175 used at 0.3 and 1µM, gefitinib and erlotinib used at 3 and 10µM). DMSO was used as vehicle control. (n=3; mean ± SEM).”

      Minor Comments:

      (1) In the introduction, the authors state that "Several type 1 and 1.5 kinase inhibitors can activate GCN2 at low concentrations while inhibiting at higher concentrations". What is the evidence that type I inhibitors can activate GCN2?

      Thank you for this query. We believe our initial phrasing was open to misinterpretation. The new wording is

      Introduction

      “Several kinase inhibitors classified as type 1 or type 1.5 with respect to their canonical targets can activate GCN2 at low concentrations while inhibiting it at higher concentrations; however, their binding mode to GCN2 remains undefined.”

      (2) A CHO:ATF4-Nanoluc translation reporter screen was used to screen 123K compounds and divide them into a pool of inhibitors and a pool of activators. The criteria used to make this distinction are not sufficiently described in the manuscript, and screen results are not supplied as supplementary data.

      We thank the reviewer for highlighting the need for greater clarity regarding the screening criteria and reproducibility. Compounds from the primary screen (BioAscent library of 123,222 drug-like compounds performed by the ALBORADA Drug Discovery Institute) were classified based on Z-score thresholds, with activators defined as those with Z-score > 3. To assess robustness, the primary screen data were re-analysed independently. In an initial analysis of 121,000 compounds (excluding plates failing quality control), 6,521 compounds met the activator threshold. Of these, 6,461 overlapped with the original hit list. The small number of discrepancies included: (i) compounds absent from the analysed dataset (e.g. originating from failed plates), and (ii) compounds with Z-scores close to the threshold (typically between −3 and −3.02), consistent with minor analytical variation (e.g. rounding). Overall, these analyses show a high degree of concordance in hit identification, with differences restricted to borderline cases near the selection threshold. We have clarified these criteria below.

      Results

      “Compounds were classified based on Z-score thresholds derived from the primary screen, with activators defined as those with Z-score > 3. An initial analysis of 121,000 compounds, excluding plates failing quality control, identified 6,521 activators, of which 6,461 overlapped with the subset selected for follow-up screening. Minor discrepancies were restricted to compounds absent from the analysed dataset (e.g. originating from failed plates) or those with Z-scores close to the threshold (3 to 3.02), consistent with limited analytical variation. Using this approach, we assembled a subset of 6,461 compounds enriched for potential ISR activators and screened these in 384-well format at 10 µM for 16 hours.”

      Reviewer #3 (Recommendations for the authors):

      (1) The legend to Figure 4 should read "Compound 20 displays GCN1 independence", not "dependence".

      Thank you for spotting this error. We have made the correction.

      (2) The text describes Figure 2D as examining 4h, but it also examines 8h.

      We have amended the text to:

      “After validating their effects using the assays described above, we next confirmed activation of the ISR by these compounds at the transcriptional level at 4 and 8 hours.”

      (3) Unless I'm misreading Table S1, compound Z134826202 is listed as activating the ISR, but the authors describe it in the text as inactive.

      Thank you. We have corrected this error.

    1. eLife Assessment

      This important paper describes the role of the Pre-rRNA in meiotic sex chromosome inactivation in mouse spermatocytes. The cytological analyses of nucleolar components and the chemical inhibition of RNA polymerase I for rDNA transcription provided solid evidence, supporting the authors' conclusions. However, the results were not well described or explained in the text, making the logic difficult to follow. This paper will be of interest to researchers in meiotic chromosome structure and the nucleolus.

    2. Reviewer #1 (Public review):

      The authors show that during prophase I of male meiosis, nucleoli disassemble and nucleolar components relocalize to the sex chromosome (XY) body. They further demonstrate that this process is regulated by the ATR-dependent signaling pathway that mediates meiotic sex chromosome inactivation (MSCI). Pharmacological disruption of pre-rRNA synthesis using the RNA polymerase I inhibitor BMH-21 leads to the recruitment of RNA polymerase II to the sex chromosomes and ectopic expression of sex chromosome-linked genes. These findings uncover a previously unrecognized role for pre-rRNAs in maintaining transcriptional silencing during meiosis. The study employs a combination of cell biology, genetics, and genomics approaches, and the conclusions are supported by compelling, well-organized data.

      Comments:

      (1) The current study focuses on transcriptional regulation of the sex chromosomes. It would be interesting to know whether perturbation of pre-rRNA synthesis also affects transcription of autosomal genes.

      (2) Is ribosome biogenesis still active during prophase I of male meiosis? Additional discussion of the timing and extent of rRNA synthesis at this stage would help place the findings in a broader biological context.

      (3) A recent preprint reports active RNA polymerase II-mediated transcription of Y chromosome genes within nucleolus-like bodies (NLBs) during prophase I of meiosis in Drosophila male germ cells (https://doi.org/10.64898/2026.05.20.726666). These findings suggest that the meiotic nucleolus may have species-specific roles in regulating sex chromosome gene expression. It would be valuable for the authors to discuss how their findings compare with these observations and the potential evolutionary implications.

    3. Reviewer #2 (Public review):

      Summary:

      The authors showed the localization pattern of nucleolus components, including Pre-rRNA, a precursor of rRNAs, changes during meiotic prophase I, particularly with the localization of these nucleolar components to the X-Y body, which shows inactivation of RNA polymerase II transcription, during pachynema. The localization of Pre-rRNA depends on ATR kinase and gammaH2AX. The chemical inhibition of rRNA transcription disrupts the binding of pre-rRNA to the X-Y body and suppresses the inhibition of the RNA polymerase II-mediated transcription on the sex chromosomes.

      Strengths:

      The cytological analysis, combined with the chemical inhibition, provided solid evidence to support the idea that, together with the remodeling of the nucleolus structure, pre-rRNA is an essential component of sex chromosome inactivation in male mouse meiosis. The role of pre-rRNA in sex chromosome inactivation in male meiosis helps our understanding of how the X-Y body, which would be a biological condensate, would be formed; e.g. for example, this Pre-rRNA may promote phase separation.

      Weaknesses:

      However, there is limited information on how Pre-rRNA is recruited to only sex chromosomes and how the RNA promotes the inactivation of sex chromosomes. Of course, these will be a target of future study. One major weakness of this paper is a poor description of the results, with fair presentation and interpretation of the data.

    1. eLife Assessment

      The authors addressed a significant biological question, namely the role of glutamine metabolism in humoral responses, and they obtained solid conclusions. The strength of this study is that the authors used state-of-the-art transgenic mouse models together with in vitro analysis, thereby providing important insights into the question posed. The manuscript has been further substantiated by adding more appropriate experimental controls and describing more in-depth functionality/physiological relevance.

    2. Reviewer #2 (Public review):

      Summary:

      In this manuscript, the authors investigate the functional requirements for glutamine and glutaminolysis in antibody responses. The authors first demonstrate that the concentrations of glutamine in lymph nodes are substantially lower than in plasma, and that at these levels, glutamine is limiting for plasma cell differentiation in vitro. The authors go on to use genetic mouse models in which B cells are deficient in glutaminase 1 (Gls), the glucose transporter Slc2a1, and/or mitochondrial pyruvate carrier 2 (Mpc2) to test the importance of these pathways in vivo. Interestingly, deficiency of Gls alone showed clear antibody defects when ovalbumin was used as the immunogen, but not the hapten NP. For the latter response, defects in antibody titers and affinity were observed only when both Gls and either Mpc2 or Slc2a1 were deleted. These latter findings form the basis of the synthetic auxotrophy conclusion. The authors go on to test these conclusions further using in vitro differentiations, Seahorse assays, pharmacological inhibitors, and targeted quantification of specific metabolites and amino acids. Finally, the authors document reduced STAT3 and STAT1 phosphorylation in response to IL-21 and interferon (both type 1 and 2), respectively, when both glutaminolysis and mitochondrial pyruvate metabolism are prevented.

      Strengths:

      (1) The main strength of the manuscript is the overall breadth of experiments performed. Orthogonal experiments are performed using genetic models, pharmacological inhibitors, in vitro assays, and in vivo experiments to support the claims. Multiple antigens are used as test immunogens--this is particularly important given the differing results.

      (2) B cell metabolism is an area of interest but understudied relative to other cell types in the immune system.

      (3) The importance of metabolic flexibility and caution when interpreting negative results is made clear from this study.

      Weaknesses:

      (1) All of the in vivo studies were done in the context of boosters at 3 weeks and recall responses 1 week later. Primary responses, including germinal centers, may still be ongoing at 3 weeks after the initial immunization and defects in GCs may contribute to the findings. Nonetheless, the authors do check antibody levels prior to the boost, and it is likely that most of the observed defects in Gls/Mpc2-deficiency are driven by faulty recall responses.

    3. Reviewer #3 (Public review):

      Summary:

      In their manuscript, the authors investigate how glutaminolysis (GLS) and mitochondrial pyruvate import (MPC2) jointly shape B cell fate and the humoral immune response. Using inducible knockout systems and metabolic inhibitors, they uncover a "synthetic auxotrophy": When GLS activity/glutaminolysis is lost together with either GLUT1-mediated glucose uptake or MPC2, B cells fail to upregulate mitochondrial respiration, IL 21/STAT3 and IFN/STAT1 signaling is impaired, and the plasma cell output and antigen-specific antibody titers drop significantly. This work thus demonstrates the promotion of plasma cell differentiation and cytokine signaling through parallel activation of two metabolic pathways. The dataset is technically comprehensive and conceptually novel, but some aspects leave the in vivo and translational significance uncertain.

      Strengths:

      (1) Conceptual novelty: the study goes beyond single-enzyme deletions to reveal conditional metabolic vulnerabilities and fate-deciding mechanisms in B cells.

      (2) Mechanistic depth: the study uncovers a novel "metabolic bottleneck" that impairs mitochondrial respiration and elevates ROS and directly ties these changes to cytokine-receptor signaling. This is both mechanistically compelling and potentially clinically relevant.

      (3) Breadth of models and methods: inducible genetics, pharmacology, metabolomics, seahorse assay, ELISpot/ELISA, RNA-seq, two immunization models.

      (4) Potential clinical angle: the synergy of CB839 with UK5099 and/or hydroxychloroquine hints at a druggable pathway targeting autoantibody-driven diseases.

      Comments on revised version.

      Authors extensively modified the text with great care and provided new data e.g. Fig. 5. Collectively, this is convincing and hence, I have no further comments.

    4. Author response:

      The following is the authors’ response to the current reviews.

      We thank the referees for noting the substantive revisions and for the praise of the work. While we each have somewhat different weightings of likelihood, we feel the appraisals are fair and reasonable.


      The following is the authors’ response to the original reviews.

      eLife Assessment

      The authors addressed an important biological question, namely the role of glutamine metabolism in humoral responses, and they obtained solid conclusions. The strength of this study is that the authors used state-of-the-art transgenic mouse models together with in vitro analysis, thereby providing significant insights into the question posed. The following would strengthen the manuscript: i) adding more in-depth functionality/physiological relevance in the discussion part, and ii) regarding the experiments, the inclusion of more appropriate controls and a clearer and more accurate description of the methods.

      We are grateful for the decision of the Editors to select this submission for in-depth peer review and to the Reviewing Editor and referees for the thoughtful and constructive comments.

      We mostly agree with the specific comments and evaluation of strengths of what the work adds as well as with indications of limitations and caveats that apply to the breadth of conclusions. We have edited the text to be more clear and provide more details about certain aspects of the Methods and Legends. In addition, although we try to avoid Discussion sections that are unduly long or have flights of fancy, we will add to the Discussion as well as edit it for directness about potential relevance, basic explorations of mechanisms, and functionality.

      The revised manuscript also contains new data, some of it dealing with comments of the referees, other additions representing work done while the manuscript was under review. While we would be inclined to do more, the sad practical problem is one of limits placed by both the absence of any grant funds and the institution's terminations (RIFs) of the two experimenters in the lab.

      While we believe the original data interpretable as presented originally, up to a point it nonetheless is good to enhance scope or have even better data and add refinements about some of the technical issues. Ultimately, the question becomes "when is enough enough?"

      In the detailed point-by-point response below, we outline changes prompted by the reviewers. We also comment on a few points more expansively that would be suitable for the paper itself, and offer some skepticism or disagreement, (longer and more detailed explanations.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      In this manuscript, Cho et al. present a comprehensive and multidimensional analysis of glutamine metabolism in the regulation of B cell differentiation and function during immune responses. They further demonstrate how glutamine metabolism interacts with glucose uptake and utilization to modulate key intracellular processes. The manuscript is clearly written, and the experimental approaches are informative and well-executed. The authors provide a detailed mechanistic understanding through the use of both in vivo and in vitro models. The conclusions are well supported by the data, and the findings are novel and impactful. I have only a few, mostly minor, concerns related to data presentation and the rationale for certain experimental choices.

      Detailed Comments:

      (1) In Figure 1b, it is unclear whether total B cells or follicular B cells were used in the assay. Additionally, the in vitro class-switch recombination and plasma cell differentiation experiments were conducted without BCR stimulation, which makes the system appear overly artificial and limits physiological relevance. Although the effects of glutamine concentration on the measured parameters are evident, the results cannot be confidently interpreted as true plasma cell generation or IgG1 class switching under these conditions. The authors should moderate these claims or provide stronger justification for the chosen differentiation strategy. Incorporating a parallel assay with anti-BCR stimulation would improve the rigor and interpretability of these findings.

      We edited the manuscript to be clear that total splenic B cells were used in this set-up figure and the rest of the paper. In addition, we performed new experiments to improve this "set-up figure (Fig. 1)" and moved the older data using alternative experimental conditions to a supplemental figure, Figure 1 - supplement 1. We also used new conditions that included styles of stimulating proliferation and differentiation - to foster an increased sense of generality. The findings in no way change the supported conclusions of the work. Specifically, we used mitogenic stimulation with anti-IgM <sup>+</sup> anti-CD40, all with BAFF, IL-4, and IL-5 in addition to the anti-CD40 stimulation of the original manuscript, bearing in mind excellent work from Aiba et al, Immunity 2006; 24: 259-268, and similar papers. In addition, we added a panel with representative flow cytometric profiles. These new data are presented in Figure 4 - supplement 1 (panels ae).

      To be transparent and add to a more open public discussion (using the virtues of this forum), the senior author and colleagues would caution about whether any in vitro conditions exist that warrant complete confidence. That is the reason for proceeding to immunization experiments in vivo. That is not said to cast doubt on our own in vitro data - there are some experiments (such as those of Fig. 1a-c and associated Fig 1 - supplement 1) that only can be done in vitro or are better done that way (e.g., because of rapid uptake of early apoptotic B cells in vivo).

      For instance: Well-respected papers use the CD40LB and NB21.2D9 systems to activate B cells and generate plasma cells. Those appear to be BCR-independent and yet continue in common use. [We found that these cellular systems (CD40LB; NB21.2D9) cannot be used in experiments with a.a. deprivation or the inhibitors due to effects on the engineered stroma-like cells.] In considering BCR engagement, Reth has published salient points about signaling and concentrations of the Ab, the upshot being that this means of activating mitogenesis and plasma cell differentiation (when the B cells are costimulated via CD40 or TLR (4 or 7/8) is also artificial. Moreover, although Aiba et al, Immunity 2006; 24: 259-268 is a laudable exception, one rarely finds papers using BAFF despite the strong evidence it is an essential part of the equation of B cell regulation in vivo and a cytokine that modulates BCR signaling - in the cultures.

      (2) In Figure 1c, the DMK alone condition is not presented. This hinders readers' ability to properly asses the glutaminolysis dependency of the cells for the measured readouts. Also, CD138<sup>+</sup> in developing PCs goes hand in hand with decreased B220 expression. A representative FACS plot showing the gating strategy for the in vitro PCs should be added as a supplementary figure. Similarly, division number (going all the way to #7) may be tricky to gate and interpret. A representative FACS plot showing the separation of B cells according to their division numbers and a subsequent gating of CD138 or IgG1 in these gates would be ideal for demonstrating the authors' ability to distinguish these populations effectively.

      In the revised manuscript, we have added new experimental data (Figure 1).

      We agree that exact placement of divisions and deconvolution by FlowJow is more fraught than might be thought from presentations in many or most papers. We include the data shown to the right as representative FACS plot(s) with old and new data that illustrate the gating on CTV fluorescence. With the representative examples pasted in here and presented in Fig 1 - supplement 1f, g of the revised manuscript, we will aver that using divisions 0-6, and ≥7 was and is entirely reasonable.

      Ditto for DMK with normal glutamine. However, in the spirit of eLife transparency lacking in many other journals, this comparison is more fraught than the referee comment would make things seem. The concentration tolerated by cells is highly dependent on the medium and glutamine concentration, and perhaps on rates of glutaminolysis (due to its generation of ammonia). In practice, DMK becomes more toxic to B cells unless glutamine is low or glutaminolysis is restricted. Thus, the concentration of DMK that is tolerated and used in Fig. 1b, c can become toxic to the B cells when using the higher levels of glutamine in typical culture media (2 mM or more) - at which point the "normal conditions <sup>+</sup> DMK" "control" involves the surviving cells in conditions with far greater cell death and less population expansion than the "low glutamine <sup>+</sup> DMK". condition.

      (3) A brief explanation should be provided for the exclusive use of IgG1 as the readout in classswitching assays, given that naïve B cells are capable of switching to multiple isotypes. Clarifying why IgG1 was preferentially selected would aid in the interpretation of the results.

      On lines ~112-3 and ~182-5, we edited the text in light of the referee's suggestion that we focus the presentation of serologic data on IgG1 in the immunization experiments. We also rearranged figures and panels to be more explicit and harmonize. That said, and [Brief explanation - IgG1 provides the strongest signal and hence better signal/noise both in vitro and with the alum-based immunizations that are avatars for the adjuvant used in the majority of protein-based vaccines for humans. Perhaps for this reason, the majority of papers on molecular mechanisms seem only to analyze IgG1. Nonetheless, since molecular regulation can differ according to isotype, and the more pro-inflammatory mouse IgG2c is more pertinent to some forms of anti-pathogen immunity and some auto-immune disease models, we believe it valuable to retain these data in supplements to the related Figures.]

      (4) The immunization experiments presented in Figures 1 and 2 are well designed, and the data are comprehensively presented. However, to prevent potential misinterpretation, it should be clarified that the observed differences between NP and OVA immunizations cannot be attributed solely to the chemical nature of the antigens - hapten versus protein. A more significant distinction lies in the route of administration (intraperitoneal vs. intranasal) and the resulting anatomical compartment of the immune response (systemic vs. lung-restricted). This context should be explicitly stated to avoid overinterpretation of the comparative findings.

      We appreciate the positive assessment, and agree with the referee that it is possible the conditions of immune challenge or re-exposure may contribute to the observed differences. We edited the text of the revised manuscript accordingly [lines ~152-153; ~159-160]. Certainly, the difference in how the anti-ova response is elicited compared to the anti-NP response in the same mice or with a bit different an immunization regimen might be another factor - or the major factor - explaining why glutaminolysis was important after ovalbumin inhalations (used because emergence of anti-ova Ab / ASCs is suppressed by the NP hapten after NP-ova immunization) but not needed for the anti-NP response unless Slc2a1 or Mpc2 also was inactivated. Thank you prompting addition of this important caveat!

      Nevertheless, it seems fair to note that in Figures 1 and 2, the ASCs and Ab are being analyzed for NP and ova in the same mice, albeit with the NP-specific components not being driven by the inhalations of ovalbumin. With that in mind, when one compares the IgG1 anti-NP ASC and Ab to those for IgG1 anti-ovalbumin (ASC in bone marrow; Ab), the ovalbumin-specific response was reduced whereas the anti-NP response was not. [lines ~171-172]

      (5) NP immunization is known to be an inducer of an IgG1-dominant Th2-type immune response in mice. IgG2c is not a major player unless a nanoparticle delivery system is used. However, the authors arbitrarily included IgG2c in their assays in Figures 2 and 3. This may be confusing for the readers. The authors should either justify the IgG2c-mediated analyses or remove them from the main figures. (It can be added as supplemental information with proper justification).

      We rearranged the Figure panels to move IgM and IgG2c data to Supplemental Figures (Figure 3 - supplements 1, 2, 4, 5 in the eLife system).

      For purposes of public discourse, we note first that in contrast to the premise about weak IgG2c responses, the data [previously, Figure 3(c, g); now in the supplements] show substantial levels of NP-specific IgG2c. The referee is quite right that the class switching and in vitro ASC generation were done with IL-4 / IgG1-promoting conditions.

      To assist readers, the revised manuscript takes note of the important role of IgG2c (mouse - IgG1 in humans) in controlling or clearing various pathogens as well as in autoimmunity [lines ~182-5]. Moreover, we continue to think that these measurements add substantial value both from the standpoint of providing a better sense of generality to the loss-of-function effects, and in considering potential ways of translating the findings to B cell-dependent autoimmune conditions such as systemic lupus erythematosus.

      [As a scientific aside, we speculate that a greater or lesser IgG2c anti-NP response may arise due to different preparations of NP-carrier obtained from the vendor (Biosearch) having different amounts of TLR (e.g., TLR4) ligand. In any case, the points of presenting the IgG2c (and IgM) data were to push against the limiting boundaries of convention (which risks perpetuating a narrow view of potential outcomes) and make the breadth of results more apparent to readers.

      (6) Similarly, in affinity maturation analyses, including IgM is somewhat uncommon. I do not see any point in showing high affinity (NP2/NP20) IgMs (Figure 3d), since that data probably does not mean much.

      As noted in the reply immediately preceding this one, we appreciate this suggestion from the reviewer and moved the IgM and IgG2c to supplemental status.

      Nonetheless, in collegial discourse we disagree a bit with the referee in light of our data as well as of work that (to our minds) leads one to question why inclusion of affinity maturation of IgM is so uncommon - as the referee accurately notes. Of course a defect in the capacity to class-switch is highly deleterious in patients but that is not the same as concluding that recall IgM or its affinity is of little consequence.

      In some of the pioneering work back in the 1980's, Bothwell showed that NP- carrier immunization generated hybridomas producing IgM Ab with extensive SHM (~11% of the 18 lineages; ~ 1/3 of the IgM hybridomas) [PMID: 8487778], IgM B cells appear to move into GC, and there is at least a reasonable published basis for the view that there are GC-derived IgM (unswitched) memory B cells (MBC) that would be more likely, upon recall activation, to differentiate into ASCs. [As an example, albeit with the Jenkins lab anti-rPE response, Taylor, Pape, and Jenkins generated quantitative estimates of the numbers of Ag-specific IgM<sup>+</sup> vs switched MBC that were GC-derived (or not). [PMID: 22370719]. While they emphasized that ~90% of IgM<sup>+</sup> MBC appeared to be GC-independent, their data also indicated that ~1/2 of all GC-derived MBC were IgM<sup>+</sup> rather than switched (their Fig. 8, B vs C; also 8E, which includes alum-PE). And while we immensely respect the referee, we are perhaps less confident that IgM or high-affinity Ag-specific IgM doesn't mean that much, if only because of evidence that localized Ab compete for Ag and may thus influence selective processes [PMCID: PMC2747358; PMID: 15953185; PMID: 23420879; PMID: 27270306].

      (7) Following on my comment for the PC generation in Figure 1 (see above), in Figure 4, a strategy that relies solely on CD40L stimulation is performed. This is highly artificial for the PC generation and needs to be justified, or more physiologically relevant PC generation strategies involving anti-BCR, CD40L, and various cytokines should be shown.

      In line with our response to point (1), we tested BCR-stimulated B cells (anti-CD40 plus anti-IgM with BAFF, IL-4, and IL-5, parallel to the analyses with anti-CD40 but no BCR engagement). These results align with and reinforce the utility of the data with anti-CD40 as the sole mitogen.

      (8) The effects of CB839 and UK5099 on cell viability are not shown. Including viability data under these treatment conditions would be a valuable addition to the supplementary materials, as it would help readers more accurately interpret the functional outcomes observed in the study.

      We added presentation of data that provide cues as to relative viability / cxmsurvival under the experimental conditions used.

      [FSC X SSC as well as 7AAD or Ghost dye panels; we also generated new data that in[ further experiments scoring annexin V staining (see Fig 4 - supplement 1d, e, and Fig 5 - supplement 1e, f)].

      (9) It is not clear how the RNA seq analysis in Figure 4h was generated. The experimental strategy and the setup need to be better explained.

      Including text added at lines ~291-293 and ~582-585, the revised manuscript provides more information in the Results, Methods and Legend for Fig 4j-l. We agree entirely with the concern and apologize that in this and a few other instances we inadvertently sacrificed sufficiency of detail on the altar of attempting brevity.

      [As a synopsis: In three temporally and biologically independent experiments, cultures were harvested 3.5 days after splenic B cells were purified and cultured as in the experiments of Fig. 4a-e. Total cellular RNA was prepared from the twelve samples (three replicates for each of four conditions - DMSO vehicle control, CB839, UK5099, and CB839 <sup>+</sup> UK5099), then analyzed by RNA-seq. RNA-seq data were initially processed using the pipeline described in the Methods. For panels g & h of Fig 4, DESeq2 was used to quantify and compare read counts in the three CB839 <sup>+</sup> UK5099 samples relative to the three independent vehicle controls and identify all genes for which variances yielded P<0.05. In Fig 4g, all such genes for which the difference was 'statistically significant' (i.e., P<0.05) were entered into the indicated Immgen tool and thereby mapped to the B lineage subsets shown in the figure panels (i.e., g, h). In (g), these are displayed using one format, whereas (h) uses the 'heatmap' tool in MyGeneSet.

      Reviewer #2 (Public review):

      Summary:

      In this manuscript, the authors investigate the functional requirements for glutamine and glutaminolysis in antibody responses. The authors first demonstrate that the concentrations of glutamine in lymph nodes are substantially lower than in plasma, and that at these levels, glutamine is limiting for plasma cell differentiation in vitro. The authors go on to use genetic mouse models in which B cells are deficient in glutaminase 1 (Gls), the glucose transporter Slc2a1, and/or mitochondrial pyruvate carrier 2 (Mpc2) to test the importance of these pathways in vivo.

      Interestingly, deficiency of Gls alone showed clear antibody defects when ovalbumin was used as the immunogen, but not the hapten NP. For the latter response, defects in antibody titers and affinity were observed only when both Gls and either Mpc2 or Slc2a1 were deleted. These latter findings form the basis of the synthetic auxotrophy conclusion. The authors go on to test these conclusions further using in vitro differentiations, Seahorse assays, pharmacological inhibitors, and targeted quantification of specific metabolites and amino acids. Finally, the authors document reduced STAT3 and STAT1 phosphorylation in response to IL-21 and interferon (both type 1 and 2), respectively, when both glutaminolysis and mitochondrial pyruvate metabolism are prevented.

      Strengths:

      (1) The main strength of the manuscript is the overall breadth of experiments performed. Orthogonal experiments are performed using genetic models, pharmacological inhibitors, in vitro assays, and in vivo experiments to support the claims. Multiple antigens are used as test immunogens--this is particularly important given the differing results.

      (2) B cell metabolism is an area of interest but understudied relative to other cell types in the immune system.

      (3) The importance of metabolic flexibility and caution when interpreting negative results is made clear from this study.

      Weaknesses:

      (1) All of the in vivo studies were done in the context of boosters at 3 weeks and recall responses 1 week later. This makes specific results difficult to interpret. Primary responses, including germinal centers, are still ongoing at 3 weeks after the initial immunization. Thus, untangling what proportion of the defects are due to problems in the primary vs. memory response is difficult.

      We performed new experiments and added the data on differences prior to a boost [see below; new Fig 3d, e; etc].

      (2) Along these lines, the defects shown in Figure 3h-i may not be due to the authors' interpretation that Gls and Mpc2 are required for efficient plasma cell differentiation from memory B cells. This interpretation would only be correct if the absence of Gls/Mpc2 leads to preferential recruitment of low-affinity memory B cells into secondary plasma cells. The more likely interpretation is that ongoing primary germinal centers are negatively impacted by Gls and Mpc2 deficiency, and this, in turn, leads to reduced affinities of serum antibodies.

      We have edited the wording of the conclusion to add a possibility we consider unlikely and downplay a conclusion that MBCs bearing switched BCRs are affected once reactivated. [see lines ~221-230] We also have added citations pertaining to the topic, including work from the Victora lab which seems to put the point succinctly: "Recall GCs in mice consist almost entirely of naïve B cells, whereas recall antibodies derive overwhelmingly from memory B cells." [emphasis added] [PMID: 38838672; new ref #83]. While unclear as to the reasoning - as one looks at the data - and skeptical as to the accuracy of the referee's point (2), it suggests that the matter is open to reasonable doubt. In line with the point and the edits, we also have added citation of a bioRxiv preprint from the Victora lab, which touches on the concept of what one could call boost-induced reinvigoration of a pre-existing GC [new ref #82].

      Beyond the textual changes, we performed a new series of experiments to investigate partially, and present the results in Fig 3d, e as well as Fig 3 - supplement 1d, e. Unfortunately, time before lab closure was an enemy both for the period between primary and recall immunizations in performance and multiple replication of work to extend that presented in Figure 3, panels g & h, and the related Supplemental Data (Fig 3 - supplements 4d, 5a-g). Unfortunately, it was not possible to do a longer-term memory experiment with recall immunization out at 8 weeks.

      The intriguing concerns and questions of points 1 & 2 provide a springboard for consideration of generalizations and simplifications. Germinal center durability is not at all monolithic, and instead is quite variable**. It is true that in the literature (especially with the substantially different approach of transferring BCR-transgenic / knock-in versions of an NP-biased BCR) there may be meaningful pools of IgG1 and IgG2c GC B cells. The premise (cognitive bias, perhaps?) in our interpretation is that in our previous work we measured few if any GC B cells - NP-APC-binding or otherwise - above the background (non-immunized controls) three weeks after immunization with NP-ovalbumin in alum. While recognizing that the immunogen can matter, we note for the readers and referee that Fig. 1 of the Taylor, Pape, & Jenkins paper considered above [PMID: 22370719] reported 10-fold more Ag-specific MBCs than GC B cells at day 29 post-immunization (the point at which the boost/recall challenge was performed in our Figure 3g, h. [That work did not use NP-carrier in alum to immunize, or measure the anti-NP response.]

      Viewing Fig. 3i from that perspective, the surmise of the comment is that a major contribution to the differences in both all-affinity and high-affinity anti-NP IgG1 (whose production requires differentiation into plasma cells) derived from the immunization at 4 wk stimulating persistent GC B cells as opposed to memory B cells.

      The issue and question also relate to rates of output of plasma cells or rises in the serum concentrations of class-switched Ab. To this point, our prior experiences agree with the long-published data of the Kurosaki lab in Figure 3c of the Aiba et al paper noted above (Immunity, 2006) (and other such time courses). Readers can note that the IgG1 anti-NP response (alum adjuvant, as in our work) hits its plateau at 2 wk, and did not increase further from 2 to 3 wk. The most likely interpretation is that GC are on the decline and Ab production has reached its plateau by the time of the 2nd immunization in Fig. 3h.

      Assuming we understand the comment and line of reasoning correctly, we also lean towards disagreeing with the statement " This interpretation would only be correct if the absence of Gls/Mpc2 leads to preferential recruitment of low-affinity memory B cells into secondary plasma cells. Our evidence shows that both low-affinity as well as high-affinity anti-NP Ab (IgG1) were reduced due to combined gene-inactivation after the peak primary response (Fig. 3h; also, see the new data in Fig 3 and Fig 3 - supplement 1). Recent papers show that affinity maturation is attributable to greater proliferation of plasmablasts with high-affinity BCR. Accordingly, the findings with loss of GLS and MPC function are quite consistent with the interpretation that much of the response after the second immunization draws on MBC differentiation into plasmablasts and then plasma cells, where the proliferative advantage of high-affinity cells is blunted by the impaired metabolism. Notwithstanding these issues, the revised manuscript includes the alternative, if less likely, interpretation proposed by the review [lines ~221-230].

      **In some contexts, of course, especially certain viral infections or vaccination with lipid nanoparticles carrying modified mRNA, germinal centres are far more persistent; also, in humans even the seasonal flu vaccine

      (3) The gating strategies for germinal centers and memory B cells in Supplemental Figure 2 are problematic, especially given that these data are used to claim only modest and/or statistically insignificant differences in these populations when Gls and Mpc2 are ablated. Neither strategy shows distinct flow cytometric populations, and it does not seem that the quantification focuses on antigen-specific cells.

      The revised manuscript improves these aspects of the presentation, using old and new data. See Fig 3 - supplement 3a, c; Fig 3 - supplement 4a. We note for readers that many other papers in the best journals show plots in which the separation of, say, GC-Tfh from overall Tfh is based on cut-off within what essentially is a continuous spectrum of emission as adjusted or compensated by the cytometer (spectral or conventional).

      The revised manuscript presents results from new experiments that deal with the subset of GC B cells whose BCRs bind NP-APC with enough affinity to retain a positive signal after washing. These new data are presented in Fig 3 - supplement 3c & 3e. In practice, the new findings suggest that the metabolic requirement applied more to the NP-binding B cells than the overall GC B cell population.

      (4) Along these lines, the conclusions in Figure 6a-d may need to be tempered if the analysis was done on polyclonal, rather than antigen-specific cells. Alum induces a heavily type 2-biased response and is not known to induce much of an interferon signature. The authors' observations might be explained by the inclusion of other ongoing GCs unrelated to the immunization.

      We apologize for ambiguity or insufficient clarity and, as noted above, have edited the text to be more clear that the in vitro experiments do not represent GC B cells and that the RNA-seq data were from experiments that did not involve alum and were not an Ag (SRBC)-specific subset.

      New text in the Results, an expanded Legend, and tweaking the Methods make it more readily clear that the RNA-seq data (and hence the GSEA) involved immunizations with SRBC (not the alum / NP system. That said, we note that the hapten-carrier experiments in which the immunogen was adjuvantized with alum actually generated a robust IgG2c (type 1-driven) response along with the type 2-enhanced IgG1 response, in line with what has been reported by others with alum-adjuvanted vaccination.

      Reviewer #3 (Public review):

      Summary:

      In their manuscript, the authors investigate how glutaminolysis (GLS) and mitochondrial pyruvate import (MPC2) jointly shape B cell fate and the humoral immune response. Using inducible knockout systems and metabolic inhibitors, they uncover a "synthetic auxotrophy": When GLS activity/glutaminolysis is lost together with either GLUT1-mediated glucose uptake or MPC2, B cells fail to upregulate mitochondrial respiration, IL 21/STAT3 and IFN/STAT1 signaling is impaired, and the plasma cell output and antigen-specific antibody titers drop significantly. This work thus demonstrates the promotion of plasma cell differentiation and cytokine signaling through parallel activation of two metabolic pathways. The dataset is technically comprehensive and conceptually novel, but some aspects leave the in vivo and translational significance uncertain.

      Strengths:

      (1) Conceptual novelty: the study goes beyond single-enzyme deletions to reveal conditional metabolic vulnerabilities and fate-deciding mechanisms in B cells.

      (2) Mechanistic depth: the study uncovers a novel "metabolic bottleneck" that impairs mitochondrial respiration and elevates ROS, and directly ties these changes to cytokinereceptor signaling. This is both mechanistically compelling and potentially clinically relevant.

      (3) Breadth of models and methods: inducible genetics, pharmacology, metabolomics, seahorse assay, ELISpot/ELISA, RNA-seq, two immunization models.

      (4) Potential clinical angle: the synergy of CB839 with UK5099 and/or hydroxychloroquine hints at a druggable pathway targeting autoantibody-driven diseases.

      We agree and thank the referee for the positive comments and this succinct summary of what we view as contributions of the paper.

      Weaknesses:

      (1) Physiological relevance of "synthetic auxotrophy"

      The manuscript demonstrates that GLS loss is only crippling when glucose influx or mitochondrial pyruvate import is concurrently reduced, which the authors name "synthetic auxotrophy". I think it would help readers to clarify the terminology more and add a concise definition of "synthetic auxotrophy" versus "synthetic lethality" early in the manuscript and justify its relevance for B cells.

      We edited the Abstract, Introduction, and Discussion to try to do better on this score. Conscious of how expansive the prose and data are even in the original submission, we appear to have taken some shortcuts that we will try to rectify or at least mitigate. Thank you for highlighting this need to improve on key concepts !!

      Specifically, the revised text expands a bit on the notion that synthetic auxotrophy represents effects on differentiation that go beyond additional mechanisms of reducing division efficiency and a modest impact on selective death. [see the 10th - 11th lines in Abstract and lines ~84-85, Introduction] Even though decreased population expansion is observed and new evidence supports a model in which the altered metabolism contributes to enhanced death in vivo, at equal division numbers the frequency of CD138<sup>+</sup> progeny is lower once glutaminolysis and mitochondrial pyruvate are reduced by either genetic or pharmacological means.

      This comment of the review raises interesting semantic questions about what represents "physiological relevance". The fundamental point is to explore a basic science question - what, if any, are limits to metabolic flexibility? In principle, shouldn't B cells be able to use fatty acid metabolism to generate enough ATP and provide the backbones for biosynthesis during growth? Put a different way, the point is that a basic curiosity to understand why decreasing glucose influx did not have an even more profound effect than what was observed, combined with curiosity as to why glutaminolysis was dispensable in relatively standard vaccine-like models of immunize/boost, provided a springboard to identification of new vulnerabilities. The manuscript shows one physiological limitation (and hence vulnerability). Be that as it may, the revised text of the Discussion section more clearly addresses this issue (lines ~531-549 at the end of the Discussion).

      While the overall findings, especially the subset specificity and the clinical implications, are generally interesting, the "synthetic auxotrophy" condition feels a little engineered.

      CAR-T cells are 'a little engineered' (or more than a little) and yet they do seem to have had an impact on understanding the centrality of B cells in various autoimmune conditions as well as in the direction of cancer therapy research. So it is a matter of balancing this perspective of the referee against the strengths they highlight in points 1, 2, and 4. In editing the revision, we try to expand and be more explicit about this in the Discussion of the revised manuscript.

      In brief, even were the money not all gone, we would not believe that expanding the heft of this already rather large manuscript and set of data would be appropriate. As matters stand, a basic new insight about metabolic flexibility and its limits leads to evidence of a way to reduce generation of Ab and a novel impairment of STAT transcription factor induction by several cytokine receptors. The vulnerability that could be tested in later work on B cell-dependent autoimmunity includes the capacity to test a compound that already has been to or through FDA phase II in patients together with an FDA-approved standard-of-care agent.

      Therefore, the findings strongly raise the question of the likelihood of such a "double hit" in vivo and whether there are conditions, disease states, or drug regimens that would realistically generate such a "bottleneck".

      Hence, the authors should document or at least discuss whether GC or inflamed niches naturally show simultaneous downregulation/lack of glutamine and/or pyruvate. The authors should also aim to provide evidence that infections (e.g., influenza), hypoxia, treatments (e.g., rapamycin), or inflammatory diseases like lupus co-limit these pathways.

      Again, we appreciate some 'licensing' to be more expansive and explicit, and will try to balance editing in such points against undue tedium or tendentiously speculative length in the Discussion. In particular, we will note that a clear, simple implication of the work is to highlight an imperative to test CB839 in lupus patients already on hydroxychloroquine as standard-of-care, and to suggest development of UK5099 (already tested many times in mouse models of cancer) to complement glutaminase inhibition.

      As backdrop, we note that the failure to advance imaging mass spectrometry to the capacity to quantify relative or absolute (via nano-DESI) concentrations of nutrients in localized interstitia is a critical gap in the entire field. Techniques that sample the interstitial fluid of tumour masses or in our case LN as a work-around have yielded evidence that there can be meaningful limitations of glucose and glutamine, but it needs to be acknowledged that such findings may be very model-specific and, as can be the case with cutting-edge science, are not without controversy. That said, yes, we had found that hypoxia reduced glutamine uptake but given the norms of focused, tidy packages only reported on leucine in an earlier paper [PMID27501247; PMCID5161594].

      Beyond all that, another impetus to and inspiration for these experiments stems from quite data that we generated in a model of short-term protein-restricted diet (loosely akin to kwashiorkor in humans), based on an excellent publication showing that such a regimen quickly led to lower circulating glutamine and mTORC1 activity (**). In brief, we found that a low-protein diet did, in our experiments, preferentially lower glutamine but - importantly - led to reduced Ab responses (which would match what we have modeled here). The findings were not a well-enough connected evidentiary component to include in the "story" but I'll append slides with the relevant data to this Response to Reviews for the referee's perusal (and anyone else who reads this online discourse).

      It would hence also be beneficial to test the CB839 + UK5099/HCQ combinations in a short, proof-of-concept treatment in vivo, e.g., shortly before and after the booster immunization or in an autoimmune model. Likewise, it may also be insightful to discuss potential effects of existing treatments (especially CB839, HCQ) on human memory B cell or PC pools.

      We certainly agree that the suggestions offered in this comment are important next steps and the right approach to test if the findings reported here translate toward the treatment of autoimmune diseases that involve B cells, interferons, and pathophysiology mediated by auto-Ab. As practical points, performance and replication of such studies would take more time than the year allotted for return of a revised manuscript to eLife and in any case neither funds nor a lab remain to do these important studies.

      Concrete evidence for our concurrence was embodied in a grant application to NIH that was essential for keeping a lab and doing any such studies. [We note, as a suggestion to others, that an essential component of such studies would be to test the effects of these compounds on B cells from patients and mice with autoimmunity]. Perhaps unfortunately for SLE patients, the review panelists did not agree about the importance of such studies. However, it can be hoped that the patent-holder of CB839 (and perhaps other companies developing glutaminase inhibitors) will see this peer-reviewed preprint and the public dialogue, and recognize how positive results might open a valuable contribution to mitigation of diseases such as SLE.

      (2) Cell survival versus differentiation phenotype

      Claims that the phenotypes (e.g., reduced PC numbers) are "independent of death" and are not merely the result of artificial cell stress would benefit from Annexin-V/active-caspase 3 analyses of GC B cells and plasmablasts. Please also show viability curves for inhibitor-treated cells.

      This comment leads us to see that the wording on this point may have been overly terse in the interests of brevity, and thereby open to some odd misunderstanding. The CD138<sup>+</sup> events are scored among VIABLE CELLS, so a decrease in the %CD138<sup>+</sup> at similar division number represents an effect independent from (or beyond) survival and division-counting. Accordingly, we expanded the text of the Abstract and elsewhere in the manuscript, to be more clear. In addition, we added data from new experiments addressing death in vitro and among GC-phenotype B cells in vivo. To clarify in this public context, it is not that an increase in death (along with the reported decrease in cell cycling) can be or is excluded. The point is that beyond any such increase, and taking into account division number (since there is evidence that PC differentiation and output numbers involve a 'division-counting' mechanism), the frequencies of CD138<sup>+</sup> cells and of ASCs among the viable cells are lower, as is the level of Prdm1-encoded mRNA even before the big increase in CD138<sup>+</sup> cells in the population.

      (3) Subset specificity of the metabolic phenotype

      Could the metabolic differences, mitochondrial ROS, and membrane-potential changes shown for activated pan-B cells (Figure 5) also be demonstrated ex vivo for KO mouse-derived GC B cells and plasma cells? This would also be insightful to investigate following NP-immunization (e.g., NP+ GC B cells 10 days after NP-OVA immunization).

      We performed a series of new experiments to have enough biologically independent replications for meaningful and statistical analyses. The new results, added in as Fig 5 - supplement 1, showed that the combined pathway interruption by loss-of-function increased ROS, mtROS, and death (annexin V / 7AAD) upon analyzing GCphenotype B cells immediately upon harvest. The findings align well with the data in Fig 5 (cultured B cells).

      (4) Memory B cell gating strategy

      I am not fully convinced that the memory-B-cell gate in Supplementary Figure 2d is appropriate. The legend implies the population is defined simply as CD19+GL7-CD38+ (or CD19+CD38++?), with no further restriction to NP-binding cells. Such a gate could also capture naïve or recently activated B cells. From the descriptions in the figure and the figure legend, it is hard to verify that the events plotted truly represent memory B cells. Please clarify the full gating hierarchy and, ideally, restrict the MBC gate to NP+CD19+GL7-CD38+ B cells (or add additional markers such as CD80 and CD273). Generally, the manuscript would benefit from a more transparent presentation of gating strategies.

      In considering the referee's viewpoint, we further expanded the supplemental data displays to include more of the gating and analytic schemes, which we believe should mitigate one concern noted here. In addition, we now include flow data from the non-immunized control mice that had been analyzed concurrently in the experiments.

      Third and finally, we performed new experiments and analyses in which the focus was the frequencies of memory-phenotype (IgD<sup>neg</sup> GL7<sup>neg</sup> CD38<sup>+</sup> / CD38<sup>hi</sup> aka CD38<sup>+</sup><sup>+</sup>) NPbinding B cells after immunization. While this time, as opposed to previously, the NP-APC staining met our standard for interpretability, the gist of the findings was that the two independent repeat experiments yielded a split decision and a degree of variability. With time being up due to the funds running out, we have elected to delete the issue and the data panel in question.

      That said, it bears noting that in the previous figure panel, the labeling indicated that the gating included the important criterion that cells be IgD<sup>neg</sup>, which excludes the vast majority of naive B cells but measures memory-phenotype B cells independent from consideration of whether or not they were NP-binding.

      [In principle marginal zone (MZ) B cells might fall within this gate. However, the MZ B population is unlikely to explain the differences shown.

      (5) Deletion efficiency - [The] mRNA data show residual GLS/MPC2 transcripts (Supplementary Figure 8). Please quantify deletion efficiency in GC B cells and plasmablasts.

      Even were there resources to do this, the degree of reduction in target mRNA (Gls; Mpc2) renders this question superfluous. To the best of our understanding, the proteins (for which there might be some phenotypic lag) are translated from RNA. Might there be a small subpopulation of B cells (or their PC progeny) with only one, or even neither, allele converted from fl to D? Yes, but they would be a minor subset in light of the magnitude of mRNA reduction, in contrast to our published observations with Slc2a1. As to plasmablasts and plasma cells, the pre-existing populations make such an analysis misleading, while the scarcity of such cells recoverable with antigen capture techniques is so low as to make both RNA and genomic DNA analyses questionable. We also refer readers to the supplemental figure that presents the results of experiments testing the issue one might infer from the question about extents of deletion in PC (i.e., how much counter-selection might have occurred by the PC stage).

    1. eLife Assessment

      This interesting study adapts machine learning tools to analyze movements of a chromatin locus in living cells in response to serum starvation. The machine learning approach developed is useful, the experiments are well controlled, and the data are solid. The study will benefit from future work testing predictions by perturbation experiments. This work will be of interest to those studying chromosome biology and gene expression patterns.

    2. Reviewer #1 (Public review):

      Summary:

      Redchuk et al. explore the dynamic properties of chromatin upon serum starvation using machine learning approaches. They use CRISPR-tagging to visualize a region on chromosome 1 in human cells and show that in their system, chromosome 1, but not the previously reported chromosomes 10, 13, and X, undergo a change in radial position upon serum starvation. Live cell imaging showed a position change towards the periphery after serum starvation. They then apply a machine learning algorithm for the analysis of the imaging data, which reveals changes in nuclear area during serum starvation and longer displacements of the chromosome 1 locus near the nuclear periphery. Differential behavior of homologues is also reported.

      Strengths:

      (1) The study of chromatin dynamics is an interesting and important area of research.

      (2) The use of machine learning approaches to analyze live cell imaging data is timely.

      (3) With serum starvation, the authors use a simple, well-controllable model system.

      Weaknesses:

      (1) This study provides limited new insight into chromatin dynamics.

      (2) It was not immediately evident what the use of machine learning approaches added to this study. It appears that the main conclusions could have been reached by conventional analysis.

      Comments on revised version:

      The authors have added some technical information, but have not made any major efforts to clarify some of the major points or to strengthen the paper. The degree of advance remains limited and several conclusions are not convincingly supported by the presented data.

    3. Reviewer #2 (Public review):

      Summary:

      The study demonstrates that CRISPR-Sirius provides a powerful approach to investigating chromosome dynamics in living cells during environmental stress. By focusing on serum starvation, the authors show that this process induces global nuclear changes, including a reduction in nuclear area and increased morphological dynamism, while at the same time driving specific reorganization of chromosome 1. Chromosome 1 relocates toward the nuclear periphery and displays distinctive patterns of motion, maintaining overall motility but punctuated by occasional long-distance displacements, particularly near the nuclear envelope. Importantly, the analysis reveals that homologous copies of chromosome 1 do not behave uniformly: peripheral loci become more mobile and responsive to starvation, whereas central homologs remain comparatively stable, often associated with nucleolar subcompartments. By integrating live imaging with machine learning and explainable AI analysis, the study highlights the complexity of nuclear organization and provides valuable insights into how chromosome-specific and locus-specific responses to stress are orchestrated within the three-dimensional nuclear landscape.

      Strengths:

      The study uses live-cell imaging to investigate the dynamics of loci during starvation. Live-cell tracking and data interpretation are carried out using machine learning and AI models, which is a major strength.

      Weaknesses:

      The manuscript is at times difficult to follow, partly because the methodological descriptions are highly specialized, especially for non-expert biologists. In addition, the observations are not tested for a mechanistic basis. Experiments that could provide deeper insights are missing, for example, why chromosome 1 moves, why the peripheral homologue dislocates, or why a "long jump" is observed at the periphery even though the speed of the loci does not change. It is also unclear whether a displacement of 0.5 μm is functionally meaningful.

      Comments on revised version:

      The authors have added some technical information and provided a better discussion of the data, but beyond that, they have not strengthened the conclusions. The observations are not supported by any perturbation assays.

    4. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      This interesting study adapts machine learning tools to analyze movements of a chromatin locus in living cells in response to serum starvation. The machine learning approach developed is useful, the experiments are well controlled, and the data are solid. The study would be greatly strengthened by testing key predictions made using perturbation experiments. This work will be of interest to those studying chromosome biology and gene expression patterns.

      We thank eLife for this nice assessment. We indeed believe that the presented machine learning approach will be useful for many types of research questions, and this was the main aim of this manuscript.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      Redchuk et al. explore the dynamic properties of chromatin upon serum starvation using machine learning approaches. They use CRISPR-tagging to visualize a region on chromosome 1 in human cells and show that in their system, chromosome 1, but not the previously reported chromosomes 10, 13, and X, undergo a change in radial position upon serum starvation. Live cell imaging showed a position change towards the periphery after serum starvation. They then apply a machine learning algorithm for the analysis of the imaging data, which reveals changes in nuclear area during serum starvation and longer displacements of the chromosome 1 locus near the nuclear periphery. Differential behavior of homologues is also reported.

      Strengths:

      (1) The study of chromatin dynamics is an interesting and important area of research.

      (2) The use of machine learning approaches to analyze live cell imaging data is timely.

      (3) With serum starvation, the authors use a simple, well-controllable model system.

      Weaknesses:

      (1) This study only provides limited new insight into chromatin dynamics.

      We respectfully disagree with this conclusion. To the best of our knowledge, our study is the first to provide any insights into chromatin dynamics upon serum starvation. Previous studies are solely based on studies in fixed cells, and the dynamics have remained unexplored. Moreover, for example the notion that homologous chromosomes show differential dynamic behavior is novel and will likely have implications and relevance to many chromatin-based processes beyond the example studied here.

      (2) It was not immediately evident what the use of machine learning approaches added to this study. It appears that the main conclusions could have been reached by conventional analysis.

      First, we would like to point out that the other reviewer found our machine learning analysis pipeline a major strength of our manuscript. Indeed, analyzing single features and assessing their impact on the studied phenomenon could have been achieved relatively easily by conventional analysis. However, this analysis would have ignored the interactions (some of which were not intuitively obvious) between different features and thereby limited the knowledge gain from the experiment.

      Unbiased analysis of the interactions between the different features would have been already very difficult and time-consuming with conventional approaches. We believe that our analysis pipeline, especially with the Shapley values, addresses the key issue of combinatorial explosion prominent to multiparametric data, such as imaging data, and helps the researcher to navigate complex datasets.

      (3) There are several specific technical points:

      (a) It was not clear what the CRISRP-Sirius probes actually labelled. The chromosome 1 sgRNA sequence is provided, but I could not find information as to which region(s) of the chromosome are actually labelled (size, location, etc.).

      We have added a schematic as Supplementary Figure 1A to show the region of the chromosome that is labelled. In addition, the target sequence, together with the relevant references can be found in the Materials and methods (page 16). Please see also below Reviewer #1 (Recommendations for the authors) point 4a.

      (b) The authors visualize a relatively small region of chromosome 1 but make conclusions regarding the entire chromosome. Additional probes on the same chromosome should be used.

      Related to this point, the discussion of why the authors are unable to reproduce the prior findings of relocation of chromosomes 10, 13, and X is not satisfying. It would be worth comparing the FISH-based painting of entire chromosomes, which generated the results suggesting relocation of these chromosomes, with the point-labelling method used here.

      We agree that our approach to labeling chromosome 1 is very different than the FISH-based probes utilized before. However, we also feel that we discuss this aspect, and the difference between our and previous results, which may also stem from the used cell model, in quite a detail in the first paragraph of the results (page 4). Also, we are very careful throughout the manuscript to indicate that here we study the dynamics of a specific chromosome loci, not the entire chromosome, and have further amended the text to emphasize this. In the future, it would be very interesting to study the dynamics of also other loci of chromosome 1. As indicated also below in response to reviewer 2, we have failed to identify further gRNAs that would reliably and reproducibly label further chromosome 1 loci, suggesting that we would need to change the labeling system entirely. Unfortunately, this is not in the scope of this manuscript. Please see also below Reviewer #1 (Recommendations for the authors) point 1.

      (c) The study lacks controls. Since in their hands chromosomes 10, 13, and X do not change position, they should be used as a negative control in all experiments demonstrating a shift in the location of chromosome 1.

      We disagree that our study lacks controls, since we use telomeres as controls throughout the manuscript. Please see also below Reviewer #1 (Recommendations for the authors) point 2,3.

      (d) I did not find information about the spatial or temporal resolution of the imaging modality. This is important to assess whether the observed changes in position, relative to time, are meaningful.

      To estimate the spatial resolution, we have added new data using fixed cells (Supplementary figure 1E; corresponding text in results on page 5); temporal resolution is indicated in Materials and methods (page 17). Please see also below Reviewer #1 (Recommendations for the authors) point 4d.

      (e) The authors analyze surprisingly early timepoints (up to 40 minutes) of serum starvation. Would these results look different if longer serum starvation timepoints of several hours were analyzed?

      We chose to analyze early time points of serum starvation based on the previous literature reporting the chromosome relocation within the first 15 minutes of starvation. Indeed, the results might look very different later during serum starvation, since we already observe differences between 0-20 min vs 20-40 min into starvation (see for example Figure 5A-D). Analyzing further time points is not in the scope of this manuscript.

      (f) The authors can do a better job of explaining what the biological meaning of the various parameters (DistR, TDist, etc.) they measure is.

      We have amended Table 1 to describe the measured features more clearly. Please see also below Reviewer #1 (Recommendations for the authors) point 4e.

      (g) I did not understand the reasoning for the authors' conclusion of differential behavior of homologues. Please explain this better, or idealy use more direct labeling methods that identify the individual homologues.

      The differential behavior of homologues is best demonstrated in Figure 6H, which shows that in serum-containing media, the peripheral homolog has equal probability of being faster or slower compared to its homolog. However, the distribution changes upon starvation, with the peripheral loci being more frequently the faster homolog. We completely agree that further studies are needed to understand this phenomenon better, but changing the labeling method is not in the scope of this manuscript.

      (h) In many figures, statistical analysis of the data is missing, including, but not limited to, Figures 1B, C, G, Figures 4, 5, 6.

      We have added a Supplementary table to include inferential statistics. See also below Reviewer #1 (Recommendations for the authors) point 4b.

      (i) No information is provided throughout the manuscript as to how many cells were analyzed in each experiment. This should be indicated in every figure legend.

      The number of analyzed loci or nucleus is indicated in every figure. See also below Reviewer #1 (Recommendations for the authors) point 4c.

      Reviewer #2 (Public review):

      Summary:

      The study demonstrates that CRISPR-Sirius provides a powerful approach to investigating chromosome dynamics in living cells during environmental stress. By focusing on serum starvation, the authors show that this process induces global nuclear changes, including a reduction in nuclear area and increased morphological dynamism, while at the same time driving specific reorganization of chromosome 1. Chromosome 1 relocates toward the nuclear periphery and displays distinctive patterns of motion, maintaining overall motility but punctuated by occasional long-distance displacements, particularly near the nuclear envelope. Importantly, the analysis reveals that homologous copies of chromosome 1 do not behave uniformly: peripheral loci become more mobile and responsive to starvation, whereas central homologs remain comparatively stable, often associated with nucleolar subcompartments. By integrating live imaging with machine learning and explainable AI analysis, the study highlights the complexity of nuclear organization and provides valuable insights into how chromosome-specific and locus-specific responses to stress are orchestrated within the three-dimensional nuclear landscape.

      Strengths:

      The study uses live-cell imaging to investigate the dynamics of loci during starvation. Livecell tracking and data interpretation are carried out using machine learning and AI models, which is a major strength.

      Weaknesses:

      The manuscript is at times difficult to follow, partly because the methodological descriptions are highly specialized, especially for non-expert biologists. In addition, the observations are not tested for a mechanistic basis. Experiments that could provide deeper insights are missing, for example, why chromosome 1 moves, why the peripheral homologue dislocates, or why a "long jump" is observed at the periphery even though the speed of the loci does not change. It is also unclear whether a displacement of 0.5 μm is functionally meaningful.

      We appreciate the comment about the readability of our manuscript, and have seriously evaluated this point. We also completely agree that it would be interesting and important to understand the mechanistic and functional basis of the observed changes in chromatin dynamics take place upon serum starvation. However, we feel that it is not in the scope of the present manuscript. See also below Reviewer #2 (Recommendations for the authors) points 3,7-11.

      Recommendations for the authors:

      Reviewing Editor Comments:

      I would like to first offer my congratulations on a very interesting study; second, I would like to encourage you to test a few key predictions using a perturbation experiment. Two reviewers with deep expertise in this area were supportive of the work, and both noted that such an addition would greatly increase the impact and visibility of this work in the field. I welcome a revision that addresses this seminal point. Thank you for sending your work to eLife!

      We thank eLife for the positive assessment. We have aimed to address all of the reviewers comments and suggestions. However, we feel that some of the suggestions are not in the scope of this particular manuscript, since they would require setting up a different chromatin labeling system.

      Reviewer #1 (Recommendations for the authors):

      The following experiments would strengthen the study:

      (1) Please label additional regions on chromosome 1 so as not to rely on a single point to represent the behavior of the entire chromosome.

      This is an excellent suggestion, but unfortunately, despite our extensive efforts, we have failed to identify further gRNAs that would reliably label chromosome loci with the CRISPR-Sirius system. Changing the labeling system is not in the scope of the presented manuscript.

      (2) Please use chromosomes 10, 13, or X as a negative control since these chromosomes do not change position in the authors' hands.

      (3) Please compare the behavior of the homologues to that of either random loci or control loci on 10, 13, or X to assess whether the differential behavior observed for chromosome 10 is a specific effect.

      Related to points 2 and 3, we opted to use telomeres as controls in this study. Throughout the manuscript, the behavior of chromosome 1 loci is compared to telomeres, demonstrating the specific effect of serum starvation on chr 1. For example, Figure 5A and 5B show that when analyzing mean locus displacement, chr1 and telomeres show the opposite behavior.

      (4) In addition:

      (a) Please provide detailed information on the sequence and location of the probes used.

      We have added a schematic showing the location of the probes as Supplementary Figure S1A. In addition, the sequences are indicated in Materials and methods (page 16).

      (b) Please provide a statistical analysis in all graphs.

      To make statistical analysis more comprehensive, we have added supplementary table 1, showing the results of inferential statistics, namely, two-sided Mann-Whitney (MW) U-test. Descriptive statistics data are shown on figures as kernel density estimation, confidence intervals and bootstrapped changes distributions. See also below Reviewer #2 (Recommendations for the authors) point 5.

      (c) Please provide throughout the manuscript in each figure legend information as to how many cells were analyzed in each experiment.

      The number of analyzed loci (or nucleus) is indicated in each graph.

      (d) Please provide information on the spatial and temporal resolution of the imaging modality.

      The imaging settings are indicated in Materials and methods, including the temporal resolution of 0.25 frames per second (page 17). To estimate spatial resolution, and especially its relationship with the observed repositioning of the chromosome loci, we performed experiments in fixed cells, using an optically identical set-up as utilized for live imaging. Unfortunately, the microscope utilized for live imaging was taken out of use by the core facility after submission of the original draft of this manuscript, but we used a microscope with essentially a similar set-up. The data from fixed cells is now presented as Supplementary figure 1E and discussed in results on page 5. This analysis indicates that the change in minimal distance to the nuclear edge, reported in our study under serum starvation in live samples (0.32 and 0.5 micron), is more than one order of magnitude above the static error.

      (e) Please better explain what the various measured parameters mean in biological terms.

      We have amended Table 1 to provide better explanation of the measured parameters.

      (f) Please add a scale bar to Figure 6I.’

      Scale bar has been added to figure 6I.

      Reviewer #2 (Recommendations for the authors):

      Major points:

      (1) SHAP analysis identified nuclear area (MA) and its change (sA) as the most predictive features of starvation state, while motility features (MD, MaxD, TD) showed strong interactions with nuclear morphology. Discrete features, such as displacement outliers and homolog subclassification by speed/proximity, influenced classification, particularly in MLP models. Could the authors clarify why morphological and motility features act in combinatorial and context-dependent ways? A biological interpretation of this interdependence would strengthen the study.

      Unfortunately, we do not have a good biological interpretation for this. The fact that some interactions are context-dependent indicates that there could be subpopulations of cells/analyzed loci. For example, we found that the predictive value of nuclear area was high in a subgroup of low motility loci (Figure 4F and Supplementary figure 4D-F). We do not believe that adding more speculation would strengthen the study.

      (2) The manuscript shows that chromosome 1 moves toward the periphery within the first 20-40 minutes of serum withdrawal. However, it remains unclear whether the locus eventually "touches" the periphery and whether it subsequently stabilizes or retracts. It would be valuable to compute the time point of minimal nuclear distance and examine whether this is transient or sustained.

      With the experimental set-up utilized here, we imaged the loci for only two minutes at random time point within the first 40 minutes of the starvation. Hence extracting the time point of minimal nuclear distance is not meaningful from this dataset. As we discuss in the manuscript, following the dynamics of the same locus for longer periods of this would be very interesting in the future. However, this is not in the scope of the present manuscript.

      (3) The manuscript is at times difficult to follow, partly because methodological descriptions are highly detailed in the main text. Consider moving more of the methodological content into Supplementary Methods and emphasizing the main results and interpretations in the main text for clarity.

      We have carefully evaluated this point. Most methodological descriptions in the manuscript relate to the machine learning models and their explanation with SHAP. As we feel that this combination is an essential part of the manuscript, and likely the aspect that can have widest impact beyond chromatin dynamics studies, we feel that the background and our reasoning related to the chosen methods are important.

      (4) The distinction between the first 20 minutes and the latter 40-minute window is intriguing. Could these different time scales be paralleled with early versus delayed gene expression responses to serum starvation? A discussion of this temporal connection would add biological depth.

      This is an intriguing idea, and we have added a short note on this in the discussion (page 13). However, as we do not know how the U2OS cells utilized here respond transcriptionally to serum starvation, we are hesitant to speculate too much.

      (5) If the observed interpretations are robust, could this be demonstrated more explicitly through statistical principles or reproducibility tests across independent datasets?

      To provide further evidence of the robustness of our findings, we have 1) added new data to estimate the spatial resolution (Supplementary figure 1E) and 2) expand the statistical analysis as supplementary table 1. Regarding the spatial resolution (see also the response to reviewer 1), our experiments on fixed cells demonstrate that the change in minimal distance to the nuclear edge, reported in our study under serum starvation in live samples (0.32 and 0.5 micron), is more than one order of magnitude above the static error. Descriptive statistics data are shown on figures as kernel density estimation, confidence intervals and bootstrapped changes distributions. To make statistical analysis more comprehensive, we added a supplementary table, showing the results of inferential statistics, namely, two-sided Mann-Whitney (MW) U-test. MW test was used as a non-parametric statistic, with null hypothesis assuming the samples are coming from the same distribution. Null hypothesis was rejected at the p-value below 0.05. In most cases (bold font in table) MW test results were in accordance with the descriptive statistics confirming the conclusions in the study. In case of exceptions (MD, TD for telomeres and TDist), the results were reported, for example, as an “appearing trend” to reflect descriptive statistics while highlighting certainty levels.

      (6) Figure labeling is difficult to follow. Please include abbreviation explanations directly in the figure panels or legends for clarity.

      Abbreviations have been added to figure legends. Adding them to figures themselves would have made the figures too busy.

      (7) The manuscript reports higher displacement at the nuclear periphery. Can the authors explain why displacement amplitudes increase near the periphery and how this relates to nuclear architecture?

      We speculate in the manuscript (results, page 11; discussion, page 14) that actually the lower displacement observed with the central locus may, at least partially, result from anchoring this locus to the nucleolus (Fig 6I). Nevertheless, alternative explanations, such as differences in transcriptional and/or chromatin states may exist (see also the response to point 11), and this is now mentioned in the discussion (page 14).

      (8) How is the movement of chromosome 1 directed specifically toward the periphery, rather than being random fluctuations? This point requires clarification.

      This is an important question, but unfortunately our data does not provide an answer to this, and suggesting any mechanism would be pure speculation. Nevertheless, our results agree with previous studies utilizing fixed cells that also demonstrated movement of chromosome 1 towards nuclear periphery (Mehta et al., 2010), arguing against random fluctuation.

      (9) Only chromosome 1, and not the other tested chromosomes, undergoes this relocalization. Could the authors elaborate on why some chromosomes but not others display this behavior?

      Previous studies (Mehta et al., 2010) utilizing chromosome paints in fixed cells actually show the relocalization of several chromosomes upon serum starvation. The fact that we observed the relocalization of only chr1 loci is likely due to the labeling method and/or the cell model utilized in this study. This is quite explicitly discussed in the first paragraph of results (page 4).

      (10) The magnitude of these movements appears relatively small (0.5 micron). Can the authors discuss whether such small but reproducible displacements are likely to be biologically meaningful in terms of nuclear function or gene regulation?

      At the moment, our experimental set up allows us to analyze the dynamics of only a small portion of chr1, which indeed shows an average 0.5 micron displacement towards the nuclear periphery. Based on the chromosome painting data from fixed cells, the displacement at the level of whole chromosome is significantly larger. As mentioned in the discussion (page 13), the functional implications of radial repositioning of chromosomes upon serum starvation is not known. Therefore further discussion on the relevance of the magnitude reported here would be pure speculation.

      (11) Peripheral homologs of chromosome 1 became faster and more dynamic under starvation. Why might these loci be more prone to movement? Could this be linked to differences in transcriptional activity or chromatin state between central and peripheral homologs?

      At the moment we favour the idea that the central homolog is constrained by its anchorage to the nucleolus (Figure 6I). However, transcriptional activity and/or chromatin state may also play a role, and this possibility is now mentioned in the discussion on page 14.

      Minor points:

      (1) Figure legends use inconsistent capitalization and panel labels. These should be standardized across all figures for better readability.

      We apologize for these inconsistencies, and have aimed to standardize all labeling.

      References

      Mehta, I.S., Amira, M., Harvey, A.J., and Bridger, J.M. (2010). Rapid chromosome territory relocation by nuclear motor activity in response to serum removal in primary human fibroblasts. Genome Biol 11, R5.

    1. eLife Assessment

      This important study reports insights into how the caspase Dcp-1, best known for cell death, can also promote tissue growth in Drosophila, extending the authors' earlier work by identifying regulatory factors that shape this non-lethal activity. The compelling findings identify a physical and functional interaction between Dcp-1 and Bruce, as well as new Dcp-1-interacting proteins that function in autophagy: Sirt1, Fkbp59, Debcl, Buffy, Atg2, and Atg8a. This work helps broaden the understanding of the non-lethal roles of Dcp-1.

    2. Reviewer #1 (Public review):

      The authors clearly demonstrate that overexpressed Dcp-1, but not Drice, is activated without canonical apoptosome components.

      Using TurboID-based proximity labeling they revealed distinct proximal proteomes, among which Sirtuin 1, an Atg8a deacetylase, which promotes autophagy, was specifically required for Dcp-1 activation. Additionally, the show that autophagy-related genes, including Bcl-2 family members Debcl and Buffy, are required for Dcp-1 activation. Using structure-based prediction using AlphaFold3 they identified that Bruce, an autophagy-regulated inhibitor of apoptosis, as a Dcp-1-specific regulator acting outside the apoptosome-mediated pathway. Finally, they show that Bruce suppresses wing tissue growth. These findings indicate that non-lethal Dcp-1 activity is governed by the autophagy- Bruce axis, enabling distinct non-lethal functions independent of cell death.

      Comments on revised version.

      No further comments.

    3. Reviewer #2 (Public review):

      Summary:

      The Drosophila executioner caspase Dcp-1 has established roles in cell death, autophagy, and imaginal disc growth. This study reports previously unrecognized factors that work together with Dcp-1. Specifically, the authors performed a turboID-based proximal ligation experiment to identify factors associated Dcp-1 and Drice. Dcp-1-specific interactors were further examined for their genetic interaction. The authors report autophagy-related genes, including Debcl and Buffy, to be required for Dcp-1 activation. In addition, the authors present evidence of an interaction between Bruce and Dcp-1. Bruce expression blocks the Dcp-1 overexpression phenotype. Inhibition of effector caspases or overexpression of Bruce commonly reduced wing growth, suggesting a relationship between the two proteins.

      Strengths:

      The study identifies new Dcp-1-interacting proteins and provides a functional link between Dcp-1 and Sirt1, Fkbp59, Debcl, Buffy, Atg2, and Atg8a. During the revision, the authors have also added convincing new data supporting the interaction between Dcp-1 and Bruce. They further make a strong case regarding the quality of the turboID-proteomics data. Overall, this is a strong manuscript supporting an interesting discovery.

    4. Reviewer #3 (Public review):

      Summary:

      The present paper by Shinoda et al. from the Miura group builds upon findings reported in an earlier study by the same team (Shinoda et al., PNAS, 2019), which identified a non-apoptotic role for the Drosophila executioner caspase Dcp-1 in promoting wing tissue growth. That earlier work attributed this function primarily to Dcp-1 and to Decay, a caspase structurally related to executioner caspases, but not to DrICE, the principal apoptotic executioner caspase. The authors further proposed that this non-apoptotic caspase activity operates independently of the initiator caspase Dronc.

      In the current study, the authors both corroborate aspects of their previous findings and extend the investigation to mechanisms regulating Dcp-1 in this context. They identify roles for the giant IAP Bruce, two BCL-2 family members, and autophagy-related components in modulating non-apoptotic Dcp-1 activity. Moreover, they show that Bruce binds to a BIR-like peptide exposed upon Dcp-1 cleavage, but not to DrICE. The study further suggests that low levels of Dcp-1 activity promote wing tissue growth, whereas excessive activity induces cell death, as evidenced by impaired wing development following Dcp-1 overexpression. Overall, the manuscript provides several intriguing insights into the non-apoptotic regulation of the comparatively weak apoptotic executioner caspase Dcp-1 and complements the group's earlier work. However, several concerns remain regarding certain interpretations of the data and the experimental rigour of some of the results.

      Strengths:

      A major strength of the work is its systematic genetic and biochemical approaches, which combine tissue-specific manipulation with protein interaction mapping to explore how Dcp-1 is regulated. The identification of several regulatory factors, including an inhibitor of cell death protein and components linked to autophagy, provides a coherent framework for understanding how Dcp-1 activity might be tuned.

      Weaknesses:

      The evidence supporting some key claims remains incomplete. In particular, the type of cell death form induced when Dcp-1 is overexpressed is not clearly established, and additional tests would be needed to distinguish between the different cell death types.

      Likely impact:

      The study contributes to a growing body of work showing that proteins traditionally associated with cell death can have broader roles in tissue development. This conceptual advance is likely to be of interest to researchers studying growth control and tissue maintenance.

      Specific points:

      (1) Nature of the wing ablation phenotype<br /> A central concern is whether the wing ablation phenotype observed upon Dcp-1 overexpression truly reflects apoptotic cell death. The authors show in Fig. 1c that nuclei in cells overexpressing Dcp-1, but not DrICE, zymogens are highly condensed, which is suggestive of apoptosis. However, it is equally plausible that this phenotype reflects a form of non-apoptotic, Dcp-1-dependent cell death (e.g. autophagy-dependent cell death). This distinction could be readily addressed using TUNEL labelling and direct caspase activity assays. The latter would be particularly informative, as it remains unclear whether zymogen Dcp-1 is capable of cleaving standard effector caspase reporters in vivo. Does the anti-cleaved Dcp-1 antibody detect Dcp-1 activation following overexpression of the Dcp-1 zymogen?

      (2) Role of Decay<br /> In their earlier study, the authors identified Decay as another caspase influencing wing growth, albeit more modestly than Dcp-1. It is therefore unclear why this line of investigation was not pursued further in the current work. This omission is notable, as Decay is not implicated in apoptosis and, to date, no substantial physiological function has been assigned to this caspase in any system. At minimum, this point should be discussed explicitly.

      (3) Fig. 2: Proximity labelling analysis<br /> The authors use TurboID-mediated proximity labelling to reveal distinct Dcp-1- and DrICE-associated proteomes across tissues, with a particular focus on the wing disc. They further demonstrate that RNAi-mediated knockdown of the Dcp-1-associated proteins Sirt1 and Fkbp59 suppresses the wing ablation phenotype induced by Dcp-1 overexpression, suggesting that these factors are required for Dcp-1 activity. However, it should be clarified whether Bruce was identified as a Dcp-1 interactor in the proximity labelling dataset, given its proposed central regulatory role. In addition, further discussion of Fkbp59, its known functions and how it might mechanistically influence Dcp-1 activity, would be valuable.

      (4) Fig. 3: Autophagy-related factors<br /> Given that Sirt1 is known to promote autophagy, the authors next examine autophagy-related proteins and identify roles for Atg2, Atg8a, Debcl, and Buffy in Dcp-1 activation. Notably, these proteins do not promote cell death in the Hid-induced canonical apoptotic pathway. However, it is important to determine whether knockdown of Debcl, Buffy, Atg2, or Atg8a alone affects wing development in the absence of Dcp-1 overexpression, to exclude the possibility that these perturbations independently impair wing formation.

      (5) Evidence for canonical autophagy<br /> The involvement of autophagy would be more convincingly demonstrated by testing additional core autophagy genes, such as Atg7, Atg5, and Atg12, as well as performing a combined knockdown of Atg8a and Atg8b. Moreover, direct assessment of autophagy at the cellular level using established genetic reporters would substantially strengthen the conclusions.

      (6) Figs. 4-5: Functional consequences<br /> It would be informative to determine whether Synr, Debcl, or Buffy influence wing size on their own and whether their overexpression enhances wing growth.

      (7) Terminology and interpretation of cell death<br /> Taken together, the results suggest that Dcp-1 zymogen overexpression induces a form of non-apoptotic cell death, potentially autophagy-dependent or related. The reviewer does not understand the authors' insistence on referring to this process as apoptosis. The authors should be more cautious in their terminology: there is no canonical versus non-canonical apoptosis, there is simply apoptosis. Without stronger evidence, these effects should not be described as apoptotic cell death.

      Comments on revised version.

      In the revised manuscript, the authors addressed each of my concerns in good faith and, in my opinion, responded to them thoroughly and satisfactorily. I have no further concerns.

    5. Author response:

      The following is the authors’ response to the original reviews.

      We are grateful to all the reviewers for dedicating time to review our manuscript and for providing insightful comments and suggestions. We have revised our manuscript in line with the reviewers' feedback. The major revisions include characterization of Dcp-1 overexpression-induced cell death, demonstration of the involvement of autophagy in Dcp-1 activation, characterization of the interaction between full-length Bruce and cleaved Dcp-1. We have introduced new figures (Figure 1 – figure supplement 1, Figure 2 – figure supplement 2, Figure 3 – figure supplement 1, Figure 5 – figure supplement 1), new panels (Figures 1D, Figure 3C, Figure 4I, J) and a new table (Table S2). The previous Figure 5 – figure supplement 1 has been relocated to Figure 4 – figure supplement 2.

      With all concerns and suggestions from the reviewers addressed, our conclusion—that Bruce suppresses autophagy-regulated caspase activity and wing tissue growth in Drosophila— is now more robustly supported. We are confident that our revised manuscript makes a significant contribution to the fields of cell death, autophagy, and developmental biology, as it provides a new conceptual framework for understanding non-lethal caspase regulation. We remain hopeful that the reviewers will find it suitable for publication in eLife.

      Reviewer #1 (Public review):

      Summary:

      The authors clearly demonstrate that overexpressed Dcp-1, but not Drice, is activated without canonical apoptosome components. Using TurboID-based proximity labeling, they revealed distinct proximal proteomes, among which Sirtuin 1, an Atg8a deacetylase, which promotes autophagy, was specifically required for Dcp-1 activation. Additionally, the show that autophagy-related genes, including Bcl-2 family members Debcl and Buffy, are required for Dcp1 activation. Using structure-based prediction using AlphaFold3, they identified that Bruce, an autophagy-regulated inhibitor of apoptosis, acts as a Dcp-1-specific regulator acting outside the apoptosome-mediated pathway. Finally, they show that Bruce suppresses wing tissue growth. These findings indicate that non-lethal Dcp-1 activity is governed by the autophagy-Bruce axis, enabling distinct non-lethal functions independent of cell death.

      Strengths:

      This is an excellent paper with very good structure, excellent quality data and analysis.

      Weaknesses:

      This reviewer did not identify any weaknesses or recommendations for revision.

      We sincerely thank the reviewer for their highly positive evaluation of our work. We are pleased that the reviewer found the overall structure, data quality, and analyses to be strong, and that they clearly recognized the key findings of our study. No changes to the manuscript were required in response to this review.

      Reviewer #2 (Public review):

      Summary:

      The Drosophila executioner caspase Dcp-1 has established roles in cell death, autophagy, and imaginal disc growth. This study reports previously unrecognized factors that work together with Dcp-1. Specifically, the authors performed a turboID-based proximal ligation experiment to identify factors associated Dcp-1 and Drice. Dcp-1-specific interactors were further examined for their genetic interaction. The authors report autophagy-related genes, including Debcl and Buffy, to be required for Dcp-1 activation. In addition, the authors present evidence of an interaction between Bruce and Dcp-1. Bruce-expression blocks the Dcp-1 overexpression phenotype. Inhibition of effector caspases or overexpression of Bruce commonly reduced wing growth, suggesting a relationship between the two proteins.

      Strengths:

      On the positive side, the study identifies new Dcp-1-interacting proteins and provides a functional link between Dcp-1 and Sirt1, Fkbp59, Debcl, Buffy, Atg2, and Atg8a.

      Weaknesses:

      The data supporting the Dcp-1/Bruce interaction are not strong, even though the title of this manuscript highlights Bruce. For example, the authors' turboID data does not support Dcp1/Bruce interaction. The case for the interaction is based on a single experiment that overexpresses a truncated Bruce transgene in S2 cells.

      We sincerely thank the reviewer for their constructive and detailed evaluation of our manuscript. We appreciate the positive assessment that our study identifies new Dcp-1-associated factors and provides functional links between Dcp-1 and Sirt1, Fkbp59, and multiple autophagy-related genes, including Debcl, Buffy, Atg2, and Atg8a. We also thank the reviewer for clearly pointing out concerns regarding the strength and interpretation of the evidence connecting Bruce and Dcp-1. In the revised manuscript, we have addressed these concerns in two major ways. First, we provided additional experimental evidence explaining why TurboID-mediated labeling did not identify Bruce. Specifically, we showed that the majority of TurboID-tagged Dcp-1 expressed in wing imaginal discs remains in its full-length form, which is unlikely to engage Bruce. Second, and more importantly, we now demonstrated that endogenously expressed full-length Bruce interacts with cleaved Dcp-1 in wing imaginal discs. These new data provide strong support for a physiologically relevant interaction between Bruce and cleaved Dcp-1. Detailed descriptions of these experiments and results are provided in the point-by-point responses in the “recommendations for the authors” section. Together, these newly added data substantially strengthen the evidence for the Dcp-1/Bruce interaction and support the focus of the original manuscript title.

      Reviewer #2 (Recommendations for the authors):

      (1) The title of the manuscript highlights Dcp-1/Bruce interaction, even though the evidence there is not strong. The evidence for Dcp-1/Sirt1 and Dcp-1/Fkbp59 is stronger. How about changing the title to highlight these other Dcp-1 interactions?

      We thank the reviewer for the thoughtful suggestion. We agree that several Dcp-1-associated factors identified in our study, particularly Sirt1 and Fkbp59, are supported by functional evidence. Specifically, our data show that Sirt1 and Fkbp59 are required for Dcp-1 overexpression-mediated activation. However, Bruce differs from these factors in both the scope and the nature of its effects on Dcp-1. Bruce is not only shown to specifically suppress Dcp-1 activity, but also to suppress wing tissue growth, indicating a broader physiological role in modulating non-lethal Dcp-1 function. Importantly, we further demonstrate that Bruce can specifically physically interact with cleaved Dcp-1. In addition, in this revised manuscript, we show that using the endogenously mStayGold::V5-tag knock-in-tagged Bruce allele, cleaved Dcp-1, induced by overexpression of Dcp-1::VENUS in wing imaginal discs, can be co-immunoprecipitated with full-length Bruce (new Figure 4I, J). These results support a physical interaction between full-length Bruce and activated Dcp-1 in vivo, consistent with a direct inhibitory role. Based on these findings, we decided to retain Bruce in the manuscript title, as it is the only factor for which both physiological and functional interactions with Dcp-1 are supported by multiple independent lines of evidence.

      (2) The case for Dcp-1/Bruce interaction is not strong because the Dcp-1 turboID fails to identify Bruce. In fact, the Dcp-1 turboID approach may not have been effective, as it failed to detect many established interactions, including Diap1 (Wang et al. 1999 PMID 10481910; Tenev et al., 2006 PMID 15580265). The authors may want to comment on this.

      We thank the reviewer for raising this important point. We agree that Bruce, as well as DIAP-1, was not identified in our TurboID-MS labeling dataset (Figure 2C, Table S1). Previous studies have shown that DIAP1 interacts with Dcp-1 and Drice only after exposure of the IAP-binding motif (IBM) at the neo-N-terminus of the large executioner caspase subunit following cleavage (Tenev et al., 2005). Similarly, our co-immunoprecipitation analyses show that Bruce interacts specifically with cleaved Dcp-1, but not with full-length Dcp-1. In the revised manuscript, we confirmed by western blot that the majority of endogenously expressed Dcp-1 in wing imaginal discs is present in the full-length pro-form (new Figure 2 – figure supplement 2A). Thus, the failure to identify Bruce and DIAP1 by TurboID-MS using full-length Dcp-1 as bait is expected, as this approach primarily labels interactors of the inactive, full-length form of Dcp-1. To evaluate whether our proximity labeling approach was nevertheless effective, we compared our TurboIDMS dataset with a previously published immune-affinity purification (IAP)-MS dataset generated using catalytically inactive, C-terminally V5-tagged Dcp-1 overexpressed in Drosophila 1(2)mbn cells (Choutka et al., 2017). Although the experimental conditions differ in several respects, we observed a substantial overlap between the TurboID-MS-mediated and IAP-MS-mediated interaction lists (new Figure 2 – figure supplement 2B, new Table S2). Importantly, SesB, one of the best-characterized Dcp-1 interactors located in mitochondria (DeVorkin et al., 2014), was also identified in our mass spectrometry dataset (new Figure 2 – figure supplement 2B, new Table S2). Based on these analyses, we now more explicitly describe the experimental context and limitations of the TurboID approach, clarifying that it preferentially labels interactors of full-length Dcp-1 in the revised manuscript. We also incorporate comparisons with prior studies to further support the validity of our mass spectrometry experiments in the revised manuscript

      (3) The best experimental evidence for Bruce/Dcp-1 interaction can be found in Figure 4H. But here, they see a weak interaction only when a truncated Bruce construct is overexpressed in S2 cells. Whether Dcp-1 interacts with Bruce in a physiological setting remains unsupported.

      We thank the reviewer for the important comment. We agree that, in the original manuscript, the biochemical evidence for the Bruce/Dcp-1 interaction relied primarily on experiments using an overexpressed truncated Bruce construct in S2 cells and therefore did not sufficiently establish whether this interaction occurs in vivo, especially in wing imaginal discs. To address this concern, we performed additional experiments to examine the Bruce/Dcp-1 interaction. In the background of the mStayGold::V5-tag knocked-in Bruce allele, we overexpressed Dcp-1::VENUS using WPGal4 driver to induce Dcp-1 activation and tested whether full-length Bruce under endogenous expression interacts with cleaved Dcp-1 in wing imaginal discs. Following immunoprecipitation with anti-V5 antibody-conjugated magnetic agarose, we found that cleaved Dcp-1 signal was enriched by co-immunoprecipitation (new Figure 4I, J). These new data demonstrate that Bruce associates with cleaved Dcp-1 in vivo and thus support the physiological relevance of the Bruce/Dcp-1 interaction. We have clarified this point in the revised manuscript and included the corresponding data.

      (4) The genetic interaction between Bruce and Dcp-1 is interesting, but the interpretation becomes complicated because Bruce inhibits Reaper, and at the same time, Dcp-1 genetically interacts with Reaper, Hid, and Grim (Figures 1E, F, G). Thus, it remains unclear if the genetic interaction between Bruce/Dcp-1 is due to a direct interaction between Bruce/Dcp-1 or alternatively, because Bruce inhibits Reaper and Grim.

      We thank the reviewer for the comment. The primary function of Reaper, Hid, and Grim (RHG proteins), collectively referred to as IAP antagonists, is to directly interact with inhibitor of apoptosis proteins (IAPs), most notably DIAP-1 (Kornbluth and White, 2005; Ryoo and Baehrecke, 2010), leading to the inhibition of DIAP-1 function. RHG proteins have not been shown to directly inhibit caspases. Because inhibition of RHG proteins results in the stabilization of DIAP-1, it is likely that the effects observed upon RHG gene knockdown are mediated through DIAP-1. Consistent with this idea, overexpression of DIAP-1, while less potent than Bruce, can also suppress Dcp-1 activation (Figure 5B, C). However, we also acknowledge that Bruce suppresses Reaper- and Grim-dependent, but not Hid-dependent, cell death (Vernooy et al., 2002). In addition, Bruce directly targets Reaper through non-lysine ubiquitination, promoting its degradation (Domingues and Ryoo, 2012). Thus, it is possible that Bruce overexpression suppresses Reaper and thereby strengthens DIAP-1 function, which could indirectly contribute to the inhibition of Dcp-1 activation. Nevertheless, because the effect of Bruce overexpression is stronger than that of DIAP-1 overexpression (Figure 5B, C), and together with our physical interaction data of Bruce with cleaved Dcp-1, we propose that Bruce most likely inhibits Dcp-1 directly to attenuate its activation.

      (5) In general, the manuscript could benefit from highlighting the strong data on Sirt1 and Fkbp59, while clearly acknowledging the limitations of the Bruce/Dcp-1 interaction.

      We thank the reviewer for the comment. As described above, in the revised manuscript we now demonstrate that endogenously expressed full-length Bruce physically interacts with cleaved Dcp-1 in wing imaginal discs (Figure 4I, J). These new data provide strong support for a physiologically relevant interaction between Bruce and cleaved Dcp-1. Based on this evidence, we decided to highlight Bruce in the manuscript, as it is the only factor for which both physiological and functional interactions with Dcp-1 are supported by multiple independent lines of evidence.

      Reviewer #3 (Public review):

      Summary:

      The present paper by Shinoda et al. from the Miura group builds upon findings reported in an earlier study by the same team (Shinoda et al., PNAS, 2019), which identified a nonapoptotic role for the Drosophila executioner caspase Dcp-1 in promoting wing tissue growth. That earlier work attributed this function primarily to Dcp-1 and to Decay, a caspase structurally related to executioner caspases, but not to DrICE, the principal apoptotic executioner caspase. The authors further proposed that this non-apoptotic caspase activity operates independently of the initiator caspase Dronc.

      In the current study, the authors both corroborate aspects of their previous findings and extend the investigation to mechanisms regulating Dcp-1 in this context. They identify roles for the giant IAP Bruce, two BCL-2 family members, and autophagy-related components in modulating nonapoptotic Dcp-1 activity. Moreover, they show that Bruce binds to a BIR-like peptide exposed upon Dcp-1 cleavage, but not to DrICE. The study further suggests that low levels of Dcp-1 activity promote wing tissue growth, whereas excessive activity induces cell death, as evidenced by impaired wing development following Dcp-1 overexpression. Overall, the manuscript provides several intriguing insights into the non-apoptotic regulation of the comparatively weak apoptotic executioner caspase Dcp-1 and complements the group's earlier work. However, several concerns remain regarding certain interpretations of the data and the experimental rigour of some of the results.

      Strengths:

      A major strength of the work is its systematic genetic and biochemical approaches, which combine tissue-specific manipulation with protein interaction mapping to explore how Dcp-1 is regulated. The identification of several regulatory factors, including an inhibitor of cell death protein and components linked to autophagy, provides a coherent framework for understanding how Dcp-1 activity might be tuned.

      Weaknesses:

      The evidence supporting some key claims remains incomplete. In particular, the type of cell death form induced when Dcp-1 is overexpressed is not clearly established, and additional tests would be needed to distinguish between the different cell death types.

      Likely impact:

      The study contributes to a growing body of work showing that proteins traditionally associated with cell death can have broader roles in tissue development. This conceptual advance is likely to be of interest to researchers studying growth control and tissue maintenance.

      We sincerely thank the reviewer for their thoughtful and constructive evaluation of our study. In response to these concerns, we have performed additional experiments to clarify the nature of the cell death induced by Dcp-1 overexpression. Based on the detection of cleaved Dcp-1, the detection of executioner caspase activity, and TUNEL assay, we now conclude that excessive Dcp-1 expression induces typical executioner caspase activity-dependent apoptotic cell death. Detailed explanations and experimental results are provided in the point-by-point responses below. Overall, we believe that these additions strengthen the manuscript by clarifying the dual roles of Dcp-1 in promoting tissue growth at low activity levels while triggering apoptosis when excessively activated.

      Specific points:

      (1) Nature of the wing ablation phenotype

      A central concern is whether the wing ablation phenotype observed upon Dcp-1 overexpression truly reflects apoptotic cell death. The authors show in Figure 1c that nuclei in cells overexpressing Dcp-1, but not DrICE, zymogens are highly condensed, which is suggestive of apoptosis. However, it is equally plausible that this phenotype reflects a form of non-apoptotic, Dcp-1-dependent cell death (e.g. autophagy-dependent cell death). This distinction could be readily addressed using TUNEL labelling and direct caspase activity assays. The latter would be particularly informative, as it remains unclear whether zymogen Dcp-1 is capable of cleaving standard effector caspase reporters in vivo. Does the anti-cleaved Dcp-1 antibody detect Dcp-1 activation following overexpression of the Dcp-1 zymogen?

      We thank the reviewer for this important point regarding the nature of cell death. We agree that nuclear condensation alone is not sufficient to conclude apoptotic cell death, and we therefore performed additional experiments. First, we performed TUNEL staining and detected robust TUNEL-positive signals in wing imaginal discs upon Dcp-1 overexpression (new Figure 1D), supporting apoptotic DNA fragmentation. Second, to directly test whether Dcp-1 overexpression leads to executioner caspase activity in vivo, we used two independent executioner caspase activity probes, GC3Ai (Schott et al., 2017; Zhang et al., 2013) and CD8::PARP::VENUS (Williams et al., 2006). Both probes showed clear executioner caspase activity-positive signals in wing imaginal discs upon Dcp-1 overexpression (new Figure 1 – figure supplement 1C–F), demonstrating that Dcp-1 overexpression leads to executioner caspase activity capable of cleaving standard substrates in vivo. In addition, staining with an anti-cleaved Dcp-1 antibody was positive upon Dcp-1 zymogen overexpression (new Figure 1 – figure supplement 1B), indicating that the overexpressed Dcp-1 zymogen is converted into its active form. Consistent with this result, western blot analysis revealed that Dcp-1 zymogen overexpression results in the appearance of a cleaved Dcp-1 (new Figure 1 – figure supplement 1A). Importantly, consistent with our original observation that the wing ablation phenotype is suppressed by expression of the caspase inhibitor p35, we further showed that p35 overexpression completely abolished the appearance of cleaved Dcp-1 in western blot (new Figure 1 – figure supplement 1A), suggesting that Dcp-1 activation is mediated by self-cleavage. Taken together, these new results demonstrate that Dcp-1 zymogen overexpression induces typical executioner caspase activity-dependent apoptotic cell death. We have clarified this point in the revised manuscript and included the corresponding data.

      (2) Role of Decay

      In their earlier study, the authors identified Decay as another caspase influencing wing growth, albeit more modestly than Dcp-1. It is therefore unclear why this line of investigation was not pursued further in the current work. This omission is notable, as Decay is not implicated in apoptosis and, to date, no substantial physiological function has been assigned to this caspase in any system. At a minimum, this point should be discussed explicitly.

      We thank the reviewer for the comment regarding the role of Decay. In our previous study (Shinoda et al., 2019), we demonstrated that both Dcp-1 and Decay promote wing tissue growth in a non-lethal manner. In the present study, however, we focused our analysis on Dcp-1. This decision was based on both technical and biological considerations. From a technical perspective, we had established TurboID knock-in lines and UAS overexpression lines for Dcp-1, Drice, and Dronc, whereas corresponding genetic tools are not available for Decay. From a biological standpoint, Dcp-1 exerts a stronger effect on wing growth than Decay, as shown in our previous work, and exhibits a dual functional spectrum: Dcp-1 promotes tissue growth at low activity levels, whereas excessive activation induces overt cell death. By contrast, Decay has not been implicated in cell death in wing imaginal discs (Kondo et al., 2006). Given that a central aim of the present study was to dissect how executioner caspase activity is differentially regulated to support nonlethal functions versus apoptotic cell death, we therefore focused on the two executioner caspases that are known to participate in apoptosis, Dcp-1 and Drice. We agree with the reviewer that Decay remains an intriguing caspase with largely unexplored physiological roles, and further investigation into its regulation and function will be an important direction for future studies. Importantly, Decay has been shown to mediate Hid-induced cell death in the DIAP1- and apoptosome-independent manner in differentiating photoreceptors and accessory cells of the eye (Leulier et al., 2006). In addition, although not required for cell death, Decay accounts for most of the caspase activity during metamorphic midgut programmed cell death, which is executed by autophagy (Denton et al., 2009). Thus, similar to Dcp-1, Decay might be an executioner caspase that can be regulated independently of the canonical apoptosome-mediated pathway, potentially involving autophagy-Bruce axis, and thereby contributing to the regulation of tissue growth. We have now discussed this point in the revised manuscript.

      (3) Figure 2: Proximity labelling analysis

      The authors use TurboID-mediated proximity labelling to reveal distinct Dcp-1- and DrICEassociated proteomes across tissues, with a particular focus on the wing disc. They further demonstrate that RNAi-mediated knockdown of the Dcp-1-associated proteins Sirt1 and Fkbp59 suppresses the wing ablation phenotype induced by Dcp-1 overexpression, suggesting that these factors are required for Dcp-1 activity. However, it should be clarified whether Bruce was identified as a Dcp-1 interactor in the proximity labelling dataset, given its proposed central regulatory role. In addition, further discussion of Fkbp59, its known functions and how it might mechanistically influence Dcp-1 activity would be valuable.

      We thank the reviewer for the comment regarding the TurboID-based proximity labeling analysis and the interpretation of the identified Dcp-1-associated factors. With respect to Bruce, we clarify that Bruce was not identified as a Dcp-1 interactor in the TurboID proximity labeling dataset. Our co-immunoprecipitation analyses in S2 cells indicate that Bruce interacts specifically with cleaved Dcp-1, but not with the full-length, inactive form. In the revised manuscript, we confirmed by western blot that the majority of endogenously expressed Dcp-1 in wing imaginal discs exists in the full-length pro-form (new Figure 2 – figure supplement 2A). Therefore, the failure to detect Bruce in the TurboID experiment using full-length Dcp-1 as bait is expected, as this approach primarily labels proteins proximal to the inactive form of Dcp-1. To examine the Bruce/Dcp-1 interaction under more physiological conditions, we performed additional in vivo experiments. Using the mStayGold::V5-tag knock-in allele of Bruce, we overexpressed Dcp1::VENUS using WP-Gal4 driver to induce Dcp-1 activation and assessed whether endogenously expressed full-length Bruce associates with Dcp-1 in wing imaginal discs. Following immunoprecipitation with anti-V5 antibody-conjugated magnetic agarose, we found that cleaved Dcp-1 signal was enriched by co-immunoprecipitation (new Figure 4I, J). These new data demonstrate that Bruce associates selectively with the cleaved, active form of Dcp-1 in vivo, thereby supporting the physiological relevance of the Bruce/Dcp-1 interaction. We have clarified this point in the revised manuscript and included the corresponding data.

      FK506-binding proteins (FKBPs) are a conserved group of proteins known to bind FK506, an immunosuppressive drug. FKBPs contain FK domains, which correspond to peptidyl cis-trans isomerase (PPIase) domains. Drosophila Fkbp59 is an orthologue of the mammalian FKBP4 and FKBP5, both of which possess a C-terminal tetratricopeptide repeat (TPR) domain that functions independently of the PPIase domain by mediating protein-protein interactions. The mammalian orthologues of Drosophila Fkbp59 function as Hsp90 co-chaperones (GharteyKwansah et al., 2018). Importantly, loss of Fkbp59 results in pupal lethality (Iki et al., 2020), which precludes further mechanistic analysis on Dcp-1 activation using adult wing phenotypes. To date, the involvement of Fkbp59 in caspase regulation has not been reported. Given that Fkbp59 functions as a co-chaperone, it may facilitate Dcp-1 activation by promoting proper folding, stability, or subcellular positioning of Dcp-1 or its regulatory factors. Importantly, Dcp1 proximal proteins are enriched in chaperone-related factors, including CCT2, CCT8, Droj2, CG16817, Fkbp59, Sgt1, and nudC; seven out of sixteen identified proximal proteins are chaperone-related. These observations suggest that Dcp-1 activity may be regulated by chaperone proteins or that Dcp-1 activity may be spatially restricted to regions enriched in chaperone machinery. Further analysis of the relationship between Dcp-1 activity and chaperone-related proteins will be important to elucidate the mechanisms and functions underlying non-lethal Dcp1 activation.

      (4) Figure 3: Autophagy-related factors

      Given that Sirt1 is known to promote autophagy, the authors next examine autophagy-related proteins and identify roles for Atg2, Atg8a, Debcl, and Buffy in Dcp-1 activation. Notably, these proteins do not promote cell death in the Hid-induced canonical apoptotic pathway. However, it is important to determine whether knockdown of Debcl, Buffy, Atg2, or Atg8a alone affects wing development in the absence of Dcp-1 overexpression, to exclude the possibility that these perturbations independently impair wing formation.

      We thank the reviewer for the comment. To address whether knockdown of Debcl, Buffy, Atg2, or Atg8a independently affects wing development, we performed RNAi-mediated knockdown of each gene using the WP-Gal4 driver in the absence of Dcp-1 overexpression. Under these conditions, knockdown of Debcl, Buffy, Atg2, or Atg8a did not cause any detectable defects in wing morphology (new Figure 3 – figure supplement 1A), indicating that these autophagy-related factors specifically function to suppress Dcp-1-mediated cell death. We have clarified this point in the revised manuscript and included the corresponding data.

      (5) Evidence for canonical autophagy

      The involvement of autophagy would be more convincingly demonstrated by testing additional core autophagy genes, such as Atg7, Atg5, and Atg12, as well as performing a combined knockdown of Atg8a and Atg8b. Moreover, direct assessment of autophagy at the cellular level using established genetic reporters would substantially strengthen the conclusions.

      We thank the reviewer for the constructive comment regarding the involvement of canonical autophagy. To further strengthen the evidence that autophagy is required for Dcp-1 activation, we examined additional core autophagy-related genes that function at distinct steps of the autophagy process, in addition to the previously tested Atg2, which mediates autophagosomal membrane expansion, and Atg8a, a core component directly associated with autophagosomal membranes. Specifically, we performed knockdown of genes including FIP200/Atg17, which is required for the initiation of autophagosome formation; Atg9, which is required for autophagosomal membrane nucleation; Atg5, which is required for autophagosomal membrane expansion through Atg12-Atg5-Atg16 ubiquitin-like conjugation system; and Stx17, which is required for autophagosome-lysosome fusion (Umargamwala et al., 2024). Because Atg8b is known to be specifically expressed in the male germline and is dispensable for autophagy, at least in fat body cells (Jipa et al., 2021), we did not further examine Atg8b in wing imaginal discs. Using WPGal4 driver, knockdown of each of these genes significantly suppressed Dcp-1-induced wing ablation phenotype (new Figure 3 – figure supplement 1C), supporting a requirement for canonical autophagy components across multiple stages of autophagosome biogenesis in Dcp-1 activation. Importantly, knockdown of these autophagy-related genes alone did not affect wing morphology in the absence of Dcp-1 overexpression (new Figure 3 – figure supplement 1B), as observed previously for Atg2 and Atg8a, suggesting the suppressive effects are specific to Dcp-1 overexpression-dependent cell death. Together, these results indicate that inhibition of autophagy at any of several key steps can suppress Dcp-1-dependent cell death, demonstrating that intact canonical autophagy is required for Dcp-1 activation. In addition, to directly assess autophagy at the cellular level, we monitored autophagosome formation using mCherry::Atg8a reporter. Upon overexpression of Dcp-1::VENUS in the wing pouch region, we observed a clear accumulation of Atg8a-positive puncta in wing imaginal discs (new Figure 3C), demonstrating that Dcp-1 overexpression induces autophagy in vivo. Together, these results provide both genetic and cellular evidence that canonical autophagy is activated upon Dcp-1 overexpression and is required for Dcp-1-dependent cell death. We have clarified this point in the revised manuscript and included the corresponding data.

      (6) Figures 4-5: Functional consequences

      It would be informative to determine whether Synr, Debcl, or Buffy influence wing size on their own and whether their overexpression enhances wing growth.

      We thank the reviewer for the suggestion regarding the functional consequences of Synr, Debcl, and Buffy on wing size. As requested, we knocked down Debcl or Buffy using WP-Gal4 driver and found that this led to reduced wing size (new Figure 5 – figure supplement 1A), indicating that endogenous Debcl and Buffy promote wing growth potentially through regulating endogenous Dcp-1 activity. We have included the corresponding data in the revised manuscript. Because Synr RNAi did not show any detectable effect on the Dcp-1 overexpression-induced phenotype (Figure 3A, B), we did not further examine the effect of Synr knockdown on wing development alone. Overexpression of Synr was not examined in this study. However, Synr overexpression has previously been reported to induce cell death in wing imaginal discs, resulting in malformed adult wings (Ikegawa et al., 2023), suggesting that increased Synr expression is likely to have deleterious rather than growth-promoting effects. Because Debcl and Buffy are both required for Synr-induced cell death, overexpression of Debcl or Buffy may lead to similar phenotypes. Therefore, we did not test Debcl or Buffy overexpression in the wing imaginal discs.

      (7) Terminology and interpretation of cell death

      Taken together, the results suggest that Dcp-1 zymogen overexpression induces a form of nonapoptotic cell death, potentially autophagy-dependent or related. The reviewer does not understand the authors' insistence on referring to this process as apoptosis. The authors should be more cautious in their terminology: there is no canonical versus non-canonical apoptosis; there is simply apoptosis. Without stronger evidence, these effects should not be described as apoptotic cell death.

      We thank the reviewer for the important comment on terminology and interpretation of the cell death phenotype. As explained in our response to comment #1, we have performed additional experiments to clarify the nature of the cell death induced by Dcp-1 overexpression. Based on the detection of cleaved Dcp-1, the detection of executioner caspase activity, and TUNEL assay, we now conclude that excessive Dcp-1 expression induces typical executioner caspase activity-dependent apoptotic cell death. At the same time, as explained in our response to comment #5, we provide both genetic and cellular evidence that canonical autophagy is activated upon Dcp-1 overexpression and promotes Dcp-1 activation. We recognized that the phrase “Dcp1 activity-regulating alternative apoptosis signaling pathway” used in Figure 5L could be misleading, as it may imply the existence of an “alternative apoptosis”. To avoid this confusion, we have revised the figure legend to read “autophagy-facilitated alternative caspase activation pathway.”

      Reviewer #3 (Recommendations for the authors):

      Figure 1c should be annotated more clearly so that it is evident that the images shown are grouped by genotype.

      We thank the reviewer for the helpful suggestion. We have added lines to Figure 1C to improve clarity by indicating that the images are grouped by genotype.

      References

      Choutka C, DeVorkin L, Go NE, Hou Y-CC, Moradian A, Morin GB, Gorski SM. 2017. Hsp83 loss suppresses proteasomal activity resulting in an upregulation of caspase-dependent compensatory autophagy. Autophagy 13:1573–1589.

      Denton D, Shravage B, Simin R, Mills K, Berry DL, Baehrecke EH, Kumar S. 2009. Autophagy, not apoptosis, is essential for midgut cell death in Drosophila. Curr Biol 19:1741–1746.

      DeVorkin L, Go NE, Hou Y-CC, Moradian A, Morin GB, Gorski SM. 2014. The Drosophila effector caspase Dcp-1 regulates mitochondrial dynamics and autophagic flux via SesB. J Cell Biol 205:477–492.

      Domingues C, Ryoo HD. 2012. Drosophila BRUCE inhibits apoptosis through non-lysine ubiquitination of the IAP-antagonist REAPER. Cell Death Differ 19:470–477.

      Ghartey-Kwansah G, Li Z, Feng R, Wang L, Zhou X, Chen FZ, Xu MM, Jones O, Mu Y, Chen S, Bryant J, Isaacs WB, Ma J, Xu X. 2018. Comparative analysis of FKBP family protein: evaluation, structure, and function in mammals and Drosophila melanogaster. BMC Dev Biol 18:7.

      Ikegawa Y, Combet C, Groussin M, Navratil V, Safar-Remali S, Shiota T, Aouacheria A, Yoo SK. 2023. Evidence for existence of an apoptosis-inducing BH3-only protein, sayonara, in Drosophila. EMBO J 42:e110454.

      Iki T, Takami M, Kai T. 2020. Modulation of Ago2 loading by Cyclophilin 40 endows a unique repertoire of functional miRNAs during sperm maturation in Drosophila. Cell Rep 33:108380.

      Jipa A, Vedelek V, Merényi Z, Ürmösi A, Takáts S, Kovács AL, Horváth GV, Sinka R, Juhász G. 2021. Analysis of Drosophila Atg8 proteins reveals multiple lipidation-independent roles. Autophagy 17:2565–2575.

      Kondo S, Senoo-Matsuda N, Hiromi Y, Miura M. 2006. DRONC coordinates cell death and compensatory proliferation. Mol Cell Biol 26:7258–7268.

      Kornbluth S, White K. 2005. Apoptosis in Drosophila: neither fish nor fowl (nor man, nor worm). J Cell Sci 118:1779–1787.

      Leulier F, Ribeiro PS, Palmer E, Tenev T, Takahashi K, Robertson D, Zachariou A, Pichaud F, Ueda R, Meier P. 2006. Systematic in vivo RNAi analysis of putative components of the Drosophila cell death machinery. Cell Death Differ 13:1663–1674.

      Ryoo HD, Baehrecke EH. 2010. Distinct death mechanisms in Drosophila development. Curr Opin Cell Biol 22:889–895.

      Schott S, Ambrosini A, Barbaste A, Benassayag C, Gracia M, Proag A, Rayer M, Monier B, Suzanne M. 2017. A fluorescent toolkit for spatiotemporal tracking of apoptotic cells in living Drosophila tissues. Development 144:3840–3846.

      Shinoda N, Hanawa N, Chihara T, Koto A, Miura M. 2019. Dronc-independent basal executioner caspase activity sustains Drosophila imaginal tissue growth. Proc Natl Acad Sci U S A 116:20539–20544.

      Tenev T, Zachariou A, Wilson R, Ditzel M, Meier P. 2005. IAPs are functionally non-equivalent and regulate effector caspases through distinct mechanisms. Nat Cell Biol 7:70–77.

      Umargamwala R, Manning J, Dorstyn L, Denton D, Kumar S. 2024. Understanding developmental cell death using Drosophila as a model system. Cells 13:347.

      Vernooy SY, Chow V, Su J, Verbrugghe K, Yang J, Cole S, Olson MR, Hay BA. 2002. Drosophila Bruce can potently suppress Rpr- and Grim-dependent but not Hid-dependent cell death. Curr Biol 12:1164–1168.

      Williams DW, Kondo S, Krzyzanowska A, Hiromi Y, Truman JW. 2006. Local caspase activity directs engulfment of dendrites during pruning. Nat Neurosci 9:1234–1236.

      Zhang J, Wang X, Cui W, Wang W, Zhang H, Liu L, Zhang Z, Li Z, Ying G, Zhang N, Li B. 2013. Visualization of caspase-3-like activity in cells using a genetically encoded fluorescent biosensor activated by protein cleavage. Nat Commun 4:2157.

    1. eLife Assessment

      The authors describe a new member of the KCNE auxiliary subunits of potassium channels from a lamprey. This new subunit represents an early evolutionary member which confers new properties when expressed along with KCNQ channels. In the revised version of the manuscript, the authors present convincing evidence from several experimental approaches. The contents of this manuscript are important and should be relevant to understanding both the mechanism of modulation of KCNQ channels by KCNE subunits and the evolutionary history of these subunits, which this manuscript now extends to the divergence of early vertebrates.

    2. Reviewer #1 (Public review):

      [Editors' note: this version has been assessed by the Reviewing Editor without further input from the original reviewers. The authors have addressed the comments raised in the previous round of review.]

      Summary:

      In this study, the authors describe an early diverging vertebrate KCNE gene present in jawless lampreys that they denote KCNE0.

      Three forms of the protein are isolated from different lampreys, which have 95% homology to each other, but only moderate homology to KCNE1-6.

      Co-expression with lamprey KCNQ1 produced a non-inactivating current, whereas co-expression with mammalian KCNQ1 resulted in less modulation. Introduction of a tetra-leucine motif from KCNE4 into KCNE0 reduced current on co-expression with KCNQ1, conferring an inhibitory effect.

      Strengths:

      This is an interesting and uncontroversial report of a new KCNE isoform from lower vertebrates that gives insight into the evolutionary progression of the sequence and functional properties of the accessory protein.

    3. Reviewer #2 (Public review):

      Summary:

      This study functionally characterizes a single KCNE-like gene, kcne0, from a jawless vertebrate. The authors conducted multiple experiments, including TEVC, VCF, RT-PCR, and RNA-seq to show that KCNQ1 and kcne0 exhibited a broadly overlapping organ distribution in lamprey species, and KCNE0 produced a constitutively active current when co-expressed with lamprey KCNQ1, similar to the effects of human KCNE3 on KCNQ1. This modulation was species-specific, as co-expression of KCNE0 with other species' KCNQ1 was less effective. Moreover, the authors found that truncating the N-terminal had a more significant reduction of the modulatory effects than truncating the C-terminal of KCNE0. Interestingly, the introduction of the tetra-leucine motif from human KCNE4 into KCNE0 conferred KCNE0 with comparable effects of human KCNE4 on KCNQ1.

      Strengths:

      The authors clearly introduced an early-diverging member of the KCNE family, and convincingly demonstrated the function of this gene, KCNE0. The results are supported by experiments of multiple approaches and are clearly written. The work is significant and will interest readers from the extended research area.

      Weaknesses:

      No major concerns were identified with the manuscript in general.

    4. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      In this study, the authors describe an early diverging vertebrate KCNE gene present in jawless lampreys that they denote KCNE0.

      Three forms of the protein are isolated from different lampreys, which have 95% homology to each other, but only moderate homology to KCNE1-6.

      Co-expression with lamprey KCNQ1 produced a non-inactivating current, whereas co-expression with mammalian KCNQ1 resulted in less modulation. Introduction of a tetra-leucine motif from KCNE4 into KCNE0 reduced current on co-expression with KCNQ1, conferring an inhibitory effect.

      Strengths:

      This is an interesting and uncontroversial report of a new KCNE isoform from lower vertebrates that gives insight into the evolutionary progression of the sequence and functional properties of the accessory protein.

      Thank you for reviewing our manuscript and for your constructive comments. Our point-to-point responses are shown below.

      Weaknesses:

      (1) No error bars visible for lamprey Q1 isoforms (open symbols) in Figure 2G. No statistical comparison was provided to indicate whether lamprey Q1 isoform V1/2s are significantly different (nor in Supplementary Table 1).

      (2) There is the same issue in Figures 3 and 4. No appropriate statistical comparison is made between V1/2s for different truncations of PmKCNE0 (Figure 3), or between KCNQ1 species isoforms with and without PmE0.

      We thank you for these helpful comments. Based on your suggestions, we revised the presentation of error bars in Fig. 2G and in other panels showing G–V or F–V relationships (Figs. 2J, 3F, 3N, 4C, 4F, 4I, 4L, and 5E; Supplementary Fig. 5D) to make the SEM bars clearer. We also added statistical comparisons of V<sub>1/2</sub> values among the three lamprey KCNQ1 orthologs in Fig. 2G and among truncation-series constructs in Figs. 3F and 3N using one-way ANOVA followed by Tukey–Kramer multiple-comparison tests. For Fig. 4, we added statistical comparisons between KCNQ1 species isoforms expressed with or without PmKCNE0 (Figs. 4C, 4F, and 4I), and between PmKCNQ1 expressed alone or with human KCNE1 or KCNE3 (Fig. 4L), using unpaired two-tailed Welch’s t-tests. These statistical comparisons are included in Supplementary Table 1.

      Reviewer #2 (Public review):

      Summary:

      This study functionally characterizes a single KCNE-like gene, kcne0, from a jawless vertebrate. The authors conducted multiple experiments, including TEVC, VCF, RT-PCR, and RNA-seq to show that KCNQ1 and kcne0 exhibited a broadly overlapping organ distribution in lamprey species, and KCNE0 produced a constitutively active current when co-expressed with lamprey KCNQ1, similar to the effects of human KCNE3 on KCNQ1. This modulation was species-specific, as co-expression of KCNE0 with other species' KCNQ1 was less effective. Moreover, the authors found that truncating the N-terminal had a more significant reduction of the modulatory effects than truncating the C-terminal of KCNE0. Interestingly, the introduction of the tetra-leucine motif from human KCNE4 into KCNE0 conferred KCNE0 with comparable effects of human KCNE4 on KCNQ1.

      Strengths:

      The authors clearly introduced an early-diverging member of the KCNE family, and convincingly demonstrated the function of this gene, KCNE0. The results are supported by experiments of multiple approaches and are clearly written. The work is significant and will interest readers from the extended research area.

      Weaknesses:

      No major concerns were identified with the manuscript in general.

      We thank you for the positive assessment of our work and for the constructive suggestions. Our point-by-point responses are provided below.

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      (1) What is the physiological role of this KCNE0 and lamprey KCNQ1 in the lamprey species? While the authors mention that the physiological roles of KCNE0 are the next focus, it is preferable to discuss some of the potential functional significance of this newly characterised KCNE.

      We thank you for this helpful suggestion. We agree that discussing the potential physiological roles of KCNQ1–KCNE0 complexes in lamprey strengthens the manuscript. We have therefore expanded the Discussion to raise the possibility that, given the broad tissue distribution of kcne0 transcripts and the ability of KCNE0 to render lamprey KCNQ1 constitutively active, KCNQ1–KCNE0 complexes may contribute to general ion homeostasis, potentially analogous to the epithelial K<sup>+</sup> recycling function of mammalian KCNQ1–KCNE3, rather than to the highly specialized KCNQ1–KCNE1 function in the mammalian heart and inner ear (page 14, lines 263–267).

      (2) Human KCNQ1 has 676 amino acids, but LcKCNQ1 contains just 507 amino acids. The species-specific regulatory effects of KCNE0 may not only be attributed to KCNE0 itself but might also be influenced by the species of KCNQ1. Some discussion on this possibility will be helpful.

      We thank you for raising this important point. We agree that the species-specific regulatory effects observed in our cross-species pairing experiments are unlikely to be determined by KCNE0 alone and may also be influenced by species-specific features of the KCNQ1 α-subunit. To address this point, we expanded the Discussion to note that the KCNQ1 proteins used in this study vary in amino-acid length, largely reflecting differences in the cytoplasmic C-terminal region, which may affect KCNQ1–KCNE compatibility and thereby influence channel gating and coupling to KCNE subunits (pages 12–13, lines 227–238). We also updated Supplementary Fig. 4 to include LrKCNQ1 and LcKCNQ1, and clarified that the LcKCNQ1 construct used in this study encodes 644 amino acids.

      (3) In Supplementary Figure 3, bands corresponding to LcKCNQ1 (507 amino acids, Supplementary Figure 1) were not seen.

      We thank you for pointing out this potentially confusing point. Supplementary Fig. 3 shows RT-PCR products amplified from tissue cDNA, not full-length amplification of the LcKCNQ1 ORF. As stated in the Methods section (pages 19–20, lines 374–399), the primers used for RT-PCR in Fig. 1F and Supplementary Fig. 3 were different from those used for cloning the full-length LcKCNQ1 cDNA shown in Supplementary Fig. 1. Therefore, a band corresponding to the full-length LcKCNQ1 coding sequence was not expected in Supplementary Fig. 3.

      To clarify this point, we revised the figure legends and indicated the RT-PCR primer-binding sites with orange arrows in Supplementary Figs. 1 and 2.

    1. eLife Assessment

      This useful study addresses a timely question about semantic prioritisation in visual working memory, using behavioural manipulations and drift-diffusion modelling. However, the strength of evidence is incomplete for the broader claims about working-memory representations because the main interpretation relies on indirect inferences from non-decision time, which cannot uniquely identify memory access or retrieval.

    2. Reviewer #1 (Public review):

      Summary:

      This paper investigates whether semantic prioritization in visual working memory reflects pre-decisional access, evidence accumulation, or both, using drift diffusion modeling across a reanalysis of prior data and two new experiments. The core finding - that semantic information receives a robust pre-decisional access advantage that is amplified by attentional disruption rather than temporal delay alone - is novel and contributes meaningfully to ongoing debates about the format and accessibility of working memory representations.

      Strengths:

      The experimental approach is well-motivated, and the use of drift-diffusion modeling to decompose decision components adds analytical value beyond standard RT and accuracy measures. The two new experiments are pre-registered and address important questions. The broader theoretical conclusion - that working memory limits are shaped not only by storage capacity but by which representational formats remain accessible under attentional uncertainty - is an important and timely contribution to the field.

      Weaknesses:

      The central interpretive claims rely heavily on differences in non-decision time, a parameter that aggregates many processes unrelated to memory retrieval, making it rather difficult to uniquely attribute the observed effects to access or retrieval mechanisms specifically. Additionally, the characterization of the two memory conditions as genuinely perceptual versus semantic warrants further justification, as both may primarily require categorical rather than format-specific knowledge.

    3. Reviewer #2 (Public review):

      This manuscript aims to characterize how semantic information is prioritized relative to perceptual details in visual working memory. The central claim is that semantic judgements benefit from faster pre‑decisional access (shorter non‑decision time), and that advantages in evidence accumulation emerge under higher cognitive demands (e.g., when items are outside the focus of attention or must be maintained under interference). Based on this, the paper argues that unattended working‑memory contents are reformatted into more abstract, long‑term‑memory‑like semantic representations that remain more readily accessible than fine‑grained perceptual features.

      Strengths:

      (1) The question is timely and relevant to current research about the format of visual working memory.

      (2) Behaviorally, the semantic advantage is carefully documented in many conditions across datasets.

      (3) The use of hierarchical drift-diffusion modelling is helpful to decompose the semantic advantage into cognitive processes such as non‑decision time and drift‑rate components.

      Weaknesses:

      (1) The strong claims about visual working‑memory representation and "long‑term‑memory‑like" formats rest on an indirect inference from decision‑model parameters to representational content, and this link is not convincingly established. Non‑decision time, as implemented here, bundles many things, such as probe processing, cue processing, retrieval/access, and motor preparation, so reduced non‑decision time for semantic probes could reflect easier question reading, simpler response mapping, or more efficient decision preparation rather than a genuine advantage in accessing semantic memory representations. Although the manuscript acknowledges that non‑decision time includes multiple processes, it nonetheless treats this parameter as primary evidence for a retrieval‑stage semantic advantage, which overstates what the data can uniquely support.

      (2) The modelling approach is relatively constrained and does not fully address the underdetermination inherent in mapping latent drift-diffusion parameters onto specific psychological mechanisms. The preferred model that allows multiple parameters (non‑decision time, drift rate, threshold) to vary provides only modest improvements in predictive accuracy over simpler models, and several key drift‑rate effects are present only in particular load or lag conditions. As a result, the theoretical interpretation that semantic prioritization primarily reflects faster access and secondarily more efficient accumulation under high demand appears rather post hoc, and alternative accounts focused on generic task efficiency or strategy differences remain plausible.

      (3) The operationalization of "semantic" is narrow and largely categorical, focusing on animacy (animal/object) and a perceptual format dimension (photo/drawing), rather than richer semantic or associative relations among items. This makes it difficult to generalize the conclusions to broader claims about semantic structure and its integration into working‑memory representations. Important recent work on how semantic and associative relationships facilitate the formation, maintenance, and retrieval of visual working memory is not adequately integrated into the theoretical framing. Consequently, the discussion tends to generalize from a specific probe structure to a broader semantic prioritization theory without engaging fully with the existing literature on semantic facilitation and neural decoding of working‑memory content.

      (4) The paper contrasts its behavioral/model‑based results with prior neural decoding findings, but the comparison is not fair. Neural decoding provides complementary evidence about the content and format of working memory representations, whereas drift-diffusion parameters reflect downstream decision dynamics given a probe. Because the current work does not include any direct representational or neural measure, its conclusions about representational "reformatting" and long‑term‑memory‑like access remain speculative and, in places, feel like a stretch.

      (5) Overall, while the data show a semantic advantage in decision‑stage measures and the modelling provides an informative decomposition of this advantage, the manuscript does not fully achieve its stated aim of characterizing the representational format of visual working memory or demonstrating a mechanistic shift toward long‑term‑memory‑like semantic representations. The work primarily informs decision‑process analyses of the conditions under which semantic judgements are faster and more robust, rather than the nature of visual working‑memory representations themselves.

    4. Author response:

      We thank the reviewers for their constructive and careful assessment of our manuscript. We are encouraged that both reviewers recognised the value of the empirical contribution: the semantic advantage is robust across experiments, the two new experiments are pre-registered, and the drift-diffusion modelling provides an informative decomposition of behavioural performance. At the same time, both reviewers raise an important and convergent point: the manuscript currently places too much interpretive weight on non-decision time and sometimes moves too quickly from decision-model parameters to claims about the representational format of working memory.

      We agree that this aspect of the manuscript should be revised. In the next version, we will substantially soften claims about adaptive reformatting and long-term-memory-like formats. We will instead frame the central contribution more precisely: semantic-category judgements show a reliable advantage at stages preceding evidence accumulation and this advantage is modulated by attentional prioritisation and interference during the maintenance interval. Our data constrain the dynamics with which different kinds of information become available for WM-guided decisions, but they do not, on their own, provide a direct measure of representational format. This hypothesis should be tested in future experiments.

      At the same time, we think the data provide stronger constraints on alternative explanations than the current manuscript makes clear. The reviewers correctly note that non-decision time is not a pure retrieval parameter, as we also note in the discussion. It can include probe encoding, response preparation, motor execution, and other processes. We will therefore avoid more explicitly equating NDT directly with retrieval latency. However, many of the alternatives raised by the reviewers, such as easier question reading or simpler response mapping for semantic probes, predict a relatively fixed semantic–perceptual offset. In our experiments, the probes and response mappings are held constant across attentional conditions, while the semantic NDT advantage changes as a function of whether the relevant item can be prioritised in advance or must be selected/reactivated at test. We will restructure the Results and Discussion to make these condition × feature interactions central to the argument.

      We will also clarify the logic of Experiment 1. We agree with Reviewer 1 that a valid retro-cue likely triggers retrieval or reactivation of the cued item. Our original phrasing, which described the valid-cue condition as reducing retrieval demands, was imprecise. The critical manipulation is better described as shifting item prioritisation/retrieval earlier in the trial. Under valid cueing, the relevant item can be prioritised before the probe appears, whereas under neutral cueing, item selection and access must occur after probe onset. We will rewrite this section accordingly.

      We will also clarify our operationalisation of semantic and perceptual categories. The present contrast is specifically between semantic category information (animate versus inanimate) and perceptual-format information (photograph versus drawing). We agree that the perceptual judgement is still categorical and does not measure fine-grained perceptual fidelity. We will therefore avoid broad claims about semantic structure or perceptual detail in general. However, as pointed in the manuscript, we believe the contrast remains meaningful: the two dimensions are orthogonal within the same stimuli, and previous work using the same feature space showed the opposite ordering during perception (Linde-Domingo et al., 2019), where perceptual-format information was available before semantic-category information. We will move this argument earlier in the manuscript and present it as converging evidence for dissociable access dynamics, while acknowledging that it does not by itself prove representational format.

      In response to the modelling concerns, we will expand the model-validation section. Specifically, we plan to add posterior predictive checks for the reported models, report model comparisons more transparently, clarify when more complex models do or do not provide practically meaningful improvements, and include sensitivity analyses using alternative parameterisations where identifiable.

      We will also make several methodological clarifications. First, because the reanalysis of Kerrén et al. (2022) forms a substantial part of the manuscript, we will add a fuller description of the original task in the main text, including how the probed item was indicated at test. Second, we will rewrite the unclear sentence describing pseudo-random stimulus selection in Experiment 1 and add a control analysis testing whether performance differs when the probed item belongs to the majority versus minority category within the trial. Third, we will clarify the stimulus repetition scheme and discuss possible long-term-memory contributions. Importantly, because semantic and perceptual probes are applied to the same items from the same trials, any repetition history or proactive-interference contribution is shared across the two probe types, although we agree that this should be discussed explicitly.

      Finally, we will revise the broader theoretical framing. We will remove or substantially qualify claims linking the present data directly to episodic memory and imagery. We will also integrate the recent literature suggested by Reviewer 2 on semantic structure, associative relations, long-term-memory contributions to working memory, and boundary conditions for semantic labelling effects. This will allow us to position the study as one piece of a broader literature on how semantic information influences WM performance, rather than as direct evidence for a general representational reformatting mechanism.

      In summary, the revised manuscript will make a narrower but stronger claim: semantic-category information shows a robust pre-accumulation advantage during WM-guided decisions, and this advantage is shaped by attentional prioritisation and interference during maintenance. We will present this as evidence about WM access dynamics and decision components, not as direct evidence that WM representations are transformed into long-term-memory-like formats.

    1. eLife Assessment

      The manuscript by von Velsen et al. offers valuable structural insights into the mitogen-activated protein kinase (MAPK) pathway by providing cryo-EM structures of stabilized MEK1-ERK2 kinase-substrate complexes in inactive, active, and nucleotide-free states, complemented by HDX-MS, SAXS, ITC, crystallography, and molecular dynamics. The work provides solid evidence for the overall architecture of the complex and identifies interaction sites that help explain MAPK pathway specificity. However, some mechanistic conclusions are not yet fully supported, particularly the designation of one state as an active phosphoryl-transfer configuration, the claim that substrate binding releases the MEK1 catalytic machinery, the proposed link to processive phosphorylation, and the extrapolation to disease-associated mutations.

    2. Reviewer #1 (Public review):

      Summary:

      This manuscript describes three conformers derived from a complex between ERK2-T185V, a variant of MEK1-DD with the KIM sequence replaced by the KIM from the p38 activator, GRA24, ADP, and AlF4-. The goal was to try to capture the complex in its active state. The results show contacts between the kinases between their N-lobe and their C-lobes that resemble MKK6-p38 complexes previously reported by the authors. Two MEK1-ERK2 conformers (States 1,3) are deemed inactive based on the lack of access of ERK-Y187 to the MEK1 active site, and the absence of ADP bound to MEK1 in State 3, while one conformer (State 2) is deemed active, but not fully active due to disorder in MEK1 activation loop (A-loop) and an essential salt bridge between strand beta3 and helix aC. HDX-MS and SAXS solution measurements and all-atom MD simulations are used to model the mutant complex and variants with WT ERK2. The study concludes that substrate recognition involves low-energy contacts with MEK, allowing substantial protein flexibility within the complex in a manner that may accommodate processive phosphorylation of ERK2.

      Strengths:

      The strengths of the work are that the findings provide important structural insights for MEK-ERK signaling and protein phosphorylation in general. These are valuable given that atomic resolution structures of kinase-substrate complexes are still limited in number. The authors succeeded in showing key contacts between subunits and conformational variations within the complex.

      Weaknesses:

      Weaknesses were that some of the conclusions about activity state, dynamics, and effects of ligand binding were less convincing. For example, that State 2 truly represents an active configuration seemed ambiguous, given the absence of Mg2+ and AlF4- in the cryoEM structure and disorder in the activation loop and the K97-E114 salt bridge. Conclusions by SAXS that ADP-AlF4 binding increases active site compaction while increasing local flexibility were not rigorously supported by HDX data, given that the latter were performed without ligand. Sections of the narrative and figures throughout were often confusing, and many assertions were made without clear explanation. Data shown in the supplementary materials were not always described in the Results, even those important for the conclusions. Figure legends and text lacked clear descriptions of specific complexes analyzed. Substantial changes are recommended to improve the readability and clarity of the work.

    3. Reviewer #2 (Public review):

      Summary:

      The authors used Cryo-EM to obtain a complex between MEK1 and ERK2. They used the same method as previously used by the same authors to form a stable complex between MKK6 and p38, an extra-strong KIM replacing the wild-type KIM in MEK1. Three conformers were resolved, with the highest resolution of 3.0 Å. The multiple conformers indicate more flexibility in MEK2 than in ERK2. These data suggest that nucleotide exchange is possible while maintaining MEK1-ERK2 interactions. SAXS and HDX data reinforce the idea of flexibility. They point to interactions between the two N-terminal domains between histidines at the N-terminus of helix C and between the G helices that are maintained in each of the 3 conformers, and sequence and structure suggest these histidines may be a source of specificity in MEK1-ERK2 versus MKK6-p38 interactions. A 2.2 Å structure of a complex between ERK1 (88% identical to ERK2) and the docking peptide used was also presented. Molecular dynamics simulations suggest that the MEK1-ERK2 complex can assume a fully active configuration of MEK1.

      Strengths:

      This is the first structure of a MEK1-ERK2 complex. The structural data are valuable additions to our understanding of MAP2K-MAPK interactions. The discussion points offered in the results section are palatable. These include the origins of specificity and the idea of flexibility in the MAP2K in support of a processive mechanism for the dual phosphorylation activity of MAP2Ks.

      Weaknesses:

      (1) This reviewer considers that the abstract is overstated. Specifically, this paper does not reveal the molecular details of phosphoryl transfer, nor does it demonstrate that substrate binding releases the catalytic machinery.

      (2) The discussion is in some places not supported by evidence and in others has superfluous text. Examples follow:

      - "Once the αG-helix is docked, and the C-lobe histidine triad is in place, the N-lobe interactions must then be fulfilled." The data in this paper does not suggest an order of events.<br /> - "If the substrate MAPK is incorrect, the N-lobe interaction will not be stabilised, preventing alignment of the MAPK A-loop with the MAP2K active site." This statement could be described as obvious.

      (3) Much of the discussion is embedded in the results, such that it is difficult to separate new facts offered by the paper from speculation.

    4. Author response:

      eLife Assessment

      The manuscript by von Velsen et al. offers valuable structural insights into the mitogen-activated protein kinase (MAPK) pathway by providing cryo-EM structures of stabilized MEK1-ERK2 kinase-substrate complexes in inactive, active, and nucleotide-free states, complemented by HDX-MS, SAXS, ITC, crystallography, and molecular dynamics. The work provides solid evidence for the overall architecture of the complex and identifies interaction sites that help explain MAPK pathway specificity. However, some mechanistic conclusions are not yet fully supported, particularly the designation of one state as an active phosphoryl-transfer configuration, the claim that substrate binding releases the MEK1 catalytic machinery, the proposed link to processive phosphorylation, and the extrapolation to disease-associated mutations.

      We would like to counter the final statement. We were very careful in our description of state 2, while we describe it as ‘active’ we clearly explain that the resolution of the reconstruction is not sufficient to define all the classical indicators of a kinase active state; however, the map is consistent with the active conformation, the complex is active in vitro, and the MEK1 variant used is the well known DD mutant that is constitutively active. While the A-loop is not observed, this is in agreement with many crystal structures of other DD mutants. We therefore decided to define this state as ‘active’ as the A-loop of ERK2 approaches the active site, the alpha-C helix has moved in and the A-loop of MEK1 no longer occludes the active site - to clarify the state we refer to the classically active confirmation as ‘fully active’. Our supporting data also show that the complex is highly dynamic during turnover, meaning we have captured MEK1 in a number of conformations on the landscape of an active state – we feel that rather than a limitation, this is an important observation in MAP2K studies. Finally, the determination of an 80 kDa complex by cryoEM to resolutions well below 4 Å is a huge technical achievement allowing the first snapshots of the MEK1-ERK2 complex to be visualised.

      Regarding the A-helix release – our observation is that the helix becomes less folded on binding of substrate. There are many studies, which we cite, that show that destabilising this helix leads to release of the catalytic machinery, see Mansour et al, 1996, Biochemistry, 35, 15529-15536 and Jindal et al 2017 J. Biol. Chem. 292, 18814-18820 for initial studies. Our observation shows that this is linked to substrate binding – a very relevant new insight that demonstrates the importance of this helix, in addition to many previous studies, but links unfolding to substrate recognition for the first time.

      For the mechanism of processive phosphorylation – it has been well established that both processive and distributive mechanisms exist. While the way that a distributive mechanism could work is obvious (complete dissociation of the two proteins), it has not been clear how a MAP2K can remain bound to its substrate and exchange nucleotides. While caution should be employed in interpreting our state 3 structure, it clearly shows what nucleotide exchange when bound to substrate can look like and that this low nucleotide affinity state is linked to disorder in the P-loop, the A-helix and substrate binding via the KIM. We would love to perform experiments that could demonstrate this but cannot at present think of an appropriate method – the reviewers did not suggest a route either.

      Finally, for the cancer-causing mutations – there are many studies demonstrating that the mutations lead to a destabilisation of the A-helix. Our study links this to substrate recognition. While this is inference, it seems justified to describe a link between substrate recognition, A-helix unfolding and disease mutations given the large body of literature describing these events.

      We are currently performing a series of in-cell activity assays that should strengthen our claims regarding the A-helix and other observations in the structure - the histidine interactions in particular.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This manuscript describes three conformers derived from a complex between ERK2-T185V, a variant of MEK1-DD with the KIM sequence replaced by the KIM from the p38 activator, GRA24, ADP, and AlF4-. The goal was to try to capture the complex in its active state. The results show contacts between the kinases between their N-lobe and their C-lobes that resemble MKK6-p38 complexes previously reported by the authors. Two MEK1-ERK2 conformers (States 1,3) are deemed inactive based on the lack of access of ERK-Y187 to the MEK1 active site, and the absence of ADP bound to MEK1 in State 3, while one conformer (State 2) is deemed active, but not fully active due to disorder in MEK1 activation loop (A-loop) and an essential salt bridge between strand beta3 and helix aC. HDX-MS and SAXS solution measurements and all-atom MD simulations are used to model the mutant complex and variants with WT ERK2. The study concludes that substrate recognition involves low-energy contacts with MEK, allowing substantial protein flexibility within the complex in a manner that may accommodate processive phosphorylation of ERK2.

      Strengths:

      The strengths of the work are that the findings provide important structural insights for MEK-ERK signaling and protein phosphorylation in general. These are valuable given that atomic resolution structures of kinase-substrate complexes are still limited in number. The authors succeeded in showing key contacts between subunits and conformational variations within the complex.

      Weaknesses:

      Weaknesses were that some of the conclusions about activity state, dynamics, and effects of ligand binding were less convincing. For example, that State 2 truly represents an active configuration seemed ambiguous, given the absence of Mg2+ and AlF4- in the cryoEM structure and disorder in the activation loop and the K97-E114 salt bridge. Conclusions by SAXS that ADP-AlF4 binding increases active site compaction while increasing local flexibility were not rigorously supported by HDX data, given that the latter were performed without ligand. Sections of the narrative and figures throughout were often confusing, and many assertions were made without clear explanation. Data shown in the supplementary materials were not always described in the Results, even those important for the conclusions. Figure legends and text lacked clear descriptions of specific complexes analyzed. Substantial changes are recommended to improve the readability and clarity of the work.

      We thank reviewer #1 for in-depth comments and analysis of our manuscript. However, there is a misunderstanding regarding HDX-MS and SAXS data. First, the HDX-MS data were performed on the ADP.AlF<sub>4</sub><sup>-</sup> inhibited complex - this was not made sufficiently clear in the text, and we will amend this. Secondly, we are not trying to support local flexibility observed in the SAXS data with the HDX data. The HDX data support the interactions observed in the cryoEM structure and demonstrate flexibility in the proline-rich loop, the ERK A-loop and unfolding of the MEK1 A-helix. The SAXS data demonstrate that when the transition state complex is formed, the complex is more compact but flexibility within the complex increases - as observed in the dimensionless Kratky plot, supporting our observations in the cryoEM maps. Therefore, the HDX data and SAXS data are separate observations. We thank reviewer #1 for all the comments and will rewrite the manuscript in order to increase clarity as suggested.

      Reviewer #2 (Public review):

      Summary:

      The authors used Cryo-EM to obtain a complex between MEK1 and ERK2. They used the same method as previously used by the same authors to form a stable complex between MKK6 and p38, an extra-strong KIM replacing the wild-type KIM in MEK1. Three conformers were resolved, with the highest resolution of 3.0 Å. The multiple conformers indicate more flexibility in MEK2 than in ERK2. These data suggest that nucleotide exchange is possible while maintaining MEK1-ERK2 interactions. SAXS and HDX data reinforce the idea of flexibility. They point to interactions between the two N-terminal domains between histidines at the N-terminus of helix C and between the G helices that are maintained in each of the 3 conformers, and sequence and structure suggest these histidines may be a source of specificity in MEK1-ERK2 versus MKK6-p38 interactions. A 2.2 Å structure of a complex between ERK1 (88% identical to ERK2) and the docking peptide used was also presented. Molecular dynamics simulations suggest that the MEK1-ERK2 complex can assume a fully active configuration of MEK1.

      Strengths:

      This is the first structure of a MEK1-ERK2 complex. The structural data are valuable additions to our understanding of MAP2K-MAPK interactions. The discussion points offered in the results section are palatable. These include the origins of specificity and the idea of flexibility in the MAP2K in support of a processive mechanism for the dual phosphorylation activity of MAP2Ks.

      Weaknesses:

      (1) This reviewer considers that the abstract is overstated. Specifically, this paper does not reveal the molecular details of phosphoryl transfer, nor does it demonstrate that substrate binding releases the catalytic machinery.

      (2) The discussion is in some places not supported by evidence and in others has superfluous text. Examples follow:

      - "Once the αG-helix is docked, and the C-lobe histidine triad is in place, the N-lobe interactions must then be fulfilled." The data in this paper does not suggest an order of events.

      - "If the substrate MAPK is incorrect, the N-lobe interaction will not be stabilised, preventing alignment of the MAPK A-loop with the MAP2K active site." This statement could be described as obvious.

      (3) Much of the discussion is embedded in the results, such that it is difficult to separate new facts offered by the paper from speculation.

      We thank reviewer #2 for comments and thorough analysis of our manuscript. We agree that perhaps the abstract should be toned down in terms of claims of an active conformation even though we feel that the combination of the first structure of the MEK1-ERK2 complex combined with MD simulation studies clearly demonstrate how phosphoryl transfer will occur. We are now also performing in-cell assays to support our theory of A-helix regulation. For point 2 we based the order of events on data from Juyoux et al 2023 Science, 381, 1217-1225, where in long-timescale MD simulations and experimentally validated adaptive Markov state model simulations the KIM interaction was the last to dissociate after the alpha-G helix interaction. In our MD simulations of the MEK1-ERK2 complex, the interactions formed by the N-lobe were weaker than those formed by the alpha-G. Indeed, dissociation of the N-lobe was observed in various independent simulations, whereas dissociation of the alpha-G was observed only once.

      Assuming that, as in the MKK6-p38a complex, the association proceeds along the reverse of the dominant dissociation pathway, the simulations suggest that the KIM interaction forms first, followed by the alpha-G and finally the N-lobes. While alternative association pathways are possible, this interpretation is consistent with the MD and in line with the main association and dissociation pathway observed for the MKK6-p38a complex. This is additionally supported by the observation that there is no catalytic activity if the KIM is removed, demonstrating this as the first essential recruitment event. We will expand this section to include our arguments.

      For the second example, we feel this is rather unfair. The statement that if the His-His interaction is absent, catalysis will be prevented is only obvious if one knows about the His-His interaction - this is the first structure showing pathway-specific interactions in the variable loop regions of a MAPK. If it is obvious, why has no one described these residues as important before?

      We have taken on board the comments on the style of the manuscript and will make significant changes as suggested by both reviewers.

    1. eLife Assessment

      The authors use a novel patch-leaving task to reveal a reward-reset strategy when mice choose to leave a depleting resource, and find that accumulated step-like activity in the dorsomedial striatum is correlated with the timing of these decisions. These important behavioral and neural findings are supported by substantial and convincing data. Additional control analyses would strengthen the evidence that these signals are specifically related to timing and patch-leaving decisions, rather than alternative task-related processes.

    2. Reviewer #1 (Public review):<br /> <br /> Summary:

      In this study, Shuler and colleagues record neurons from the DMS in mice performing a patch foraging task. In this task, mice had the choice between harvesting rewards from 2 ports - one the time-investment port where the rate of reward declined over time and the other a context port where the rate of reward was either high or low. Mice performed the task appropriately, switching between ports as the rate of reward declined in the time-investment port and switching more rapidly when the context port delivered high versus low rewards. The behavior of the mice was also strongly driven by time since the most recent reward receipt, in conflict with normative accounts of patch foraging. Individual DMS neurons showed bistable firing patterns, transitioning to high rates of activity at various times from reward. Overall, the population tiled the temporal space, and the accumulation of the number of neurons in the high firing state was predictive of patch exit. The rate of accumulation varied with things that also affected behavior.

      Strengths:

      Overall, the aims of the study were clear and important, the experiment directly addresses them, and the results are clear and provide compelling support for the authors' conclusions.

      Weaknesses:

      I have only a few comments and questions to consider, none of which are criticisms of what was done, really.

      (1) Probably my chief question, alluded to in the discussion, is what the evidence is that DMS plays a causal role in generating these correlates and the resulting behavior, in light of the lack of causal evidence here. What are other options? Could such information depend on upstream areas such as OFC or mPFC, with DMS just a pass-through? And while I would not ask for causal data, is there a specific prediction? That is, if the area were inactivated, would mice stay longer or shorter? Not do the task? If I wanted to do a causal test of the authors' idea regarding the contribution of DMS to this behavior, what would be predicted, and what result would invalidate the hypothesis? Speculating on this a bit, beyond just saying DMS is involved, would be useful.

      (2) Not much is said about the suboptimal strategy. Would DMS continue to play the same role if the mice showed no effect of recent reward and instead performed appropriately? Or is some other area doing that job? Or is this not important? I thought it was interesting that the mice basically did not treat the game quite like they were supposed to. Is it important to go back and look at what is happening in DMS under normative conditions to really know how this area contributes to proper foraging?

      (3) Do these neurons also track time in the context port? Or do they only exhibit this behavior in the port where rewards are depleting? This seems like an interesting question. Do they show the same profile in different ports, if so?

    3. Reviewer #2 (Public review):

      Summary:

      Here, Sutlief et al. use a novel patch-foraging task to investigate the role of dorsomedial striatal (DMS) neurons in determining when animals disengage from a resource. They show that mice, contrary to canonical optimal-foraging predictions, adopt a strategy in which reward receipt resets timing behavior, with decisions further shaped by both cumulative time spent in a patch and the overall quality of the environment. The authors further demonstrate that a subset of DMS neurons exhibits step-like activity patterns during task performance. Importantly, the accumulation of these state transitions across the neuronal population predicts the timing of patch-leaving decisions on a trial-by-trial basis, providing a potential neural mechanism underlying decisions about when to abandon a currently exploited resource.

      Strengths:

      This study addresses an important question using a well-designed, interesting behavioral task. The finding that mice employ a reward-triggered exit-timing policy is particularly interesting, as it is pertinent to the many patch foraging-style tasks that have been developed for use in mice, where rewards are delivered as discrete events. The identification of step-like activity in DMS neurons is mostly compelling, and the authors' trial-by-trial analysis linking this activity to behavior provides some support for its relevance to patch-leaving decisions.

      Weaknesses:

      A key interpretational issue is whether the DMS signal reflects timing specifically, rather than movement initiation or other task-related factors. The authors argue that once a sufficient number of neurons transition, the animal exits the time-investment port. However, it remains unclear whether this population threshold reflects a timing computation that determines when to leave in the more abstract sense, or a signal more directly related to movement onset (that may also be initiated after some proportion of the population has changed its activity). An important control would be to examine neural activity while animals are engaged at the context port. In this epoch, animals presumably do not need to time their departure in the same way, but they still eventually initiate movement. If the DMS signal reflects timing rather than movement, one would not expect the same accumulation-to-threshold pattern of step-like transitions at the context port.

      It would also be helpful for the authors to clarify the behavioral definition of the leaving decision. Can mice return to the time-investment port after exiting it if they do not subsequently enter the context port? How exactly is "exit" defined: as withdrawal from the time-investment port, entry into the context port, or some other behavioral event? Is there variability in the latency between time-investment port exit and context-port entry, and if so, is this latency related to DMS step-like activity? These details are important for interpreting whether the neural activity is aligned with a timing decision, movement initiation, or the execution of a transition between task states.

      The classification approach for identifying step-like activity seems generally reasonable, and the low false-positive rate against homogeneous Poisson controls is reassuring. However, one potential issue is that the identification of trial-by-trial state transitions is not independent of the session-level characterization of each neuron. The algorithm first fits a sigmoid to the pooled session data and then uses the resulting high- and low-firing-rate states to constrain interval-level fits. This may bias the analysis toward finding step-like transitions in neurons whose activity is only approximately step-like at the session level, effectively reducing the space of alternative solutions available to the interval-level fits. As implemented, the approach therefore functions more as a detector of consistency with a session-defined step model than as an unbiased test of whether individual intervals are better described by discrete state transitions versus alternative dynamics such as ramps or gradual drifts. This concern could be addressed by comparing the constrained sigmoid model against alternatives, such as constant-rate or ramping models, on held-out intervals, or by deriving state parameters from an independent subset of trials and testing classification on the remaining trials.

      The inclusion threshold for the accumulation analysis is difficult to evaluate. Sessions were included if they contained at least seven simultaneously recorded step-like units, but this number is hard to interpret without knowing the total number of recorded units per session and the fraction classified as step-like. Seven units may be sufficient for fitting a population accumulation trajectory, but because the cutoff is based on an absolute number rather than a proportion of the recorded population, it is unclear whether included sessions reflect robust population-level step-like dynamics or a relatively small selected subset of DMS activity. Reporting the number and fraction of step-like units per session, as well as the sensitivity of the accumulation results across different inclusion thresholds, would help clarify this point.

    4. Reviewer #3 (Public review):

      Sutlief and colleagues report behavioral and neural results from mice performing a patch foraging task. Behaviorally, they argue that time since last reward is a major determinant of when mice decide to leave a patch. In the brain, they find neurons in the dorsomedial striatum that show step-like changes in their firing rate at a range of times following reward. Population analyses show that the cumulative fraction of neurons that have undergone such a step-like change in firing rate can be used to predict patch-leaving times with impressive accuracy.

      Overall, this is an interesting set of results that has been analyzed in a principled way. The manuscript is well written, the results are explained clearly, and the evidence supporting the authors' conclusions is strong. The manuscript is therefore a potentially valuable contribution to the growing literature assessing how the brain solves stopping problems like the patch foraging scenario. I have suggestions for the authors to consider that might further increase the rigor of their results, and a few suggestions for improving the clarity of the work for readers.

      (1) I don't quite understand how the behavioral task works. Are mice rewarded for making discrete nose poke responses in the investment and context ports? Or are they required to nose poke and hold? Is reward given with some probability per response (which decreases with time in the patch), or is the reward probability a function of elapsed time in the patch, time since last response, or dwell time in the port? Also, exactly what equation defines how reward probability changes over time for the high- and low-value contexts? I couldn't find these details anywhere in the manuscript, and they would be helpful for better understanding the behavior and the later neural results.

      (2) How was the optimal strategy determined? Several features of the author's task violate the assumption of the marginal value theorem, so computing the optimal residence time is not a straightforward application of the classic model. There's a diagram in Figure 1h that depicts an MVT-like graphical solution, but the conventions of the plot are not familiar to me, and there's no description of how it works in the results or methods. More detail here would be much appreciated. In a similar vein, the authors report that mice generally exceeded optimal residence times in patches, but no statistical comparison is provided to back up that statement. There should be some formal test of this if it is to be included in the results.

      (3) The authors argue that time since last reward is the predominant determinant of patch leaving time. However, as the authors note, time since last reward is correlated with other task variables (patch reward rate, time in patch, etc.). I don't trust that SVM coefficients can be interpreted as straightforward measures of a variable's importance for classification performance in the case of correlated predictors. A better approach would be to assess how well the model performs as subsets of variables are added or removed from the model.

      (4) For the SVM analysis, I'm not quite understanding how or why the authors are using 5 s after mice left the patch as additional "Leave" examples. For instance, is time since entry computed for the investment patch, or the context patch that mice enter after they leave the investment patch? Similarly, is the time since the last reward relative to the investment patch, or the reward the mouse is likely to receive at the context patch? Moreover, I'm not sure it's safe to assume that because the mouse left at time t, time t+1 necessarily reflects conditions on which the mouse would definitely leave again. If we're thinking about the stay/leave decision as something that is being repeated sequentially on a fast time scale to determine how long mice stay in the patch, it doesn't follow that observing a mouse leave means that any patch conditions after that would necessarily result in the same decision. If that were the case, it would mean that seeing a mouse leave a patch after 2 s would preclude ever observing a residence time longer than 2 s, which is clearly not compatible with the authors' data. Ultimately, it's only possible to observe one decision to leave per trial; including data points beyond that as additional leave examples seems overly speculative to me.

      (5) The authors validate their approach for quantifying step-like changes in firing rate using simulations of constant-rate Poisson spiking and observe a low false positive rate. This is encouraging, but it doesn't seem like the only way in which their method could go awry, or even the most concerning way. I would be much more interested in seeing the false positive rate for continuous, ramp-like changes in firing rate, which would be much more likely to trip up the authors' approach and are also the major relevant alternative hypothesis to step-like changes in firing rate. Random walks in firing rate might also be worth testing.

      (6) The finding that cumulative "transitioned" neurons is predictive of patch leaving is interesting. However, I can't help but wonder how truly informative this variable is for predicting patch leaving. It seems as though neurons can only transition firing rates one time. That means that as time in the patch increases, the fraction of transitioned neurons naturally increases. Similarly, all visits must eventually end with the mouse leaving the patch, so the hazard rate of leaving increases with time in the patch. Given that, can the authors be certain that the cumulative transitioned neurons are really what's predicting patch leaving time, or would any generically increasing function perform roughly the same? An interesting test would be to mismatch the neural predictor and behavior at the level of trials. If this mechanism is really specific, rather than something that captures the general structure of an increasing hazard rate of leaving, then prediction of leaving time should work substantially better when the neural predictor is correctly matched to behavior on the trial for which it was recorded.

    5. Author response:

      Reviewer #1

      (1) Causality and the role of DMS; a specific, falsifiable prediction. We agree that the paper should not leave the causal question implicit, and we will expand the Discussion to state a concrete prediction rather than a general claim of involvement. Briefly, if the accumulation signal we describe carries the animal's intended departure time, then suppressing DMS during patch occupancy should not simply shift exit times in one direction but should degrade their structure: exit-time variability should increase, and exit timing should lose its systematic dependence on reward-rate context and on the time of the most recent reward. A plausible alternative outcome is disengagement from the task altogether, which would be uninformative and would need to be controlled for. The result that would falsify our hypothesis is the one we will state explicitly: exit timing that remains as predictable, and as sensitive to context and reward history, under DMS suppression, as without it. We will also discuss the alternative the reviewer raises, that these signals are inherited from cortical inputs such as OFC or mPFC with DMS acting as a relay, and note that our data cannot presently distinguish this from a locally generated signal.

      (2) Behavior under a normative strategy. This is an interesting question and we will address it in the Discussion. Our expectation, which we will frame as a prediction rather than a result, is that an animal timing from patch entry rather than resetting at each reward would show accumulation that begins at entry and proceeds to a context-dependent threshold at the reward-rate-optimal time, rather than the reward-triggered resets we observe. In this view, the reset structure of the neural signal is a reflection of the behavioral policy rather than a property of the region. We will make clear that this is a testable prediction that our current dataset does not address.

      (3) Do these neurons also track time at the context port? We intend to answer with new analysis, and it converges with Reviewer #2's suggested analysis (below), so we treat the two together there.

      Reviewer #2

      (a) Timing versus movement initiation: activity at the context port. We take this to be a central interpretational concern. We will examine whether the step-like DMS activity extends to the context port, testing the interval between the final context-port reward and departure for the same step-like transitions and accumulation we observe in the time-investment port. We will apply the same comparison to the context-port inter-reward intervals, which addresses Reviewer #1's third point about whether these neurons also track time at the context port.

      We want to flag one feature of the task that bears on how the outcome should be read. The context port is not a timing-free epoch. Its four rewards are delivered at predictable, regularly spaced intervals, and the interval between the final reward and the animal's departure is self-timed. Departure from the context port is therefore also a self-timed action, and observing accumulation there would not by itself indicate that the signal reflects movement initiation rather than timing. What the comparison can inform is whether the accumulation is specific to a decision about when to disengage from a depleting resource, or is a more general feature of self-timed departures. This is a meaningful distinction either way, and one we will report and interpret whichever direction the result falls.

      (b) Operational definition of leaving; the exit-to-entry latency. We agree these details are necessary for interpretation and their absence is our omission. Exit is the final withdrawal from the time-investment port preceding the next context-port visit, and we will make that clear in the revised methods. Mice can and occasionally do re-enter the time-investment port without an intervening context-port visit (especially early in training). Such re-entries are not counted as exits. We will also examine whether the latency between time-investment-port exit and context-port entry relates to the accumulation slope on the corresponding interval, to test whether the neural signal relates to the decision or to the execution of the transition.

      (c) Independence of interval-level fits from the session-level model. This is a fair characterization of the procedure, and we accept the distinction the reviewer draws between a detector of consistency with a session-defined step model and an unbiased test of discrete versus continuous dynamics. We will address it with a held-out validation: estimating each unit's state parameters and transition time from one half of its intervals and testing whether the transition times recovered from the withheld half agree. The discrete-versus-continuous comparison is addressed directly by the ramp simulations under Reviewer #3's point (5) below.

      (d) The inclusion threshold for the accumulation analysis. We will add a supplementary figure reporting the total number of recorded units per session and the fraction classified as step-like, so that the seven-unit criterion can be evaluated against the recorded population rather than in the abstract. Yield varied substantially across sessions, from a handful of units to roughly one hundred, and we will show this distribution directly. We will also report the accumulation results across a range of inclusion thresholds spanning approximately five to eight simultaneously recorded step-like units, so that readers can assess sensitivity to the choice.

      Reviewer #3

      (1) Specification of the task. We agree the task description was insufficient, and we will correct this at the front of the Results and in the Methods. The time-investment port operates on a poke-and-hold basis: the mouse maintains its head in the port and rewards are delivered stochastically over time for as long as it remains, with no requirement to withdraw and re-poke. Reward delivery follows an exponentially decaying rate in time since port entry, with a time constant of eight seconds, integrating to an expected eight rewards of one microliter each (8 µL total) for indefinite occupancy; we will give the explicit function. The reward probability function in the time-investment port is identical across blocks. The high- and low-reward-rate contexts are properties of the context port alone (four rewards over five seconds versus four rewards over ten seconds), and we will make this contrast unambiguous, since it is the manipulation on which the design rests.

      (2) Derivation of the optimum, Figure 1h, and a formal test of overstaying. We appreciate this comment. The optimal residence time in our task is not obtained by the classical Charnov tangent construction. It is computed by explicit maximization of the overall reward rate over the full cycle, following the framework in Sutlief et al. (2025) and shown graphically in Figure 1h. We will expand the legend of Figure 1h so its conventions are stated explicitly, give the reward-rate-maximizing derivation as an explicit equation in the Methods, and reframe the surrounding text around reward-rate maximization as the normative principle, with MVT identified as the special case it is. We will also add the formal statistical comparison of observed residence times against the computed optimum, which the reviewer correctly notes was asserted rather than tested.

      (3) Interpretation of SVM coefficients with correlated predictors. We accept this criticism. We will not rest the ordering of predictors on coefficient magnitudes alone. We will add a variable inclusion-and-ablation analysis, reporting cross-validated classification performance as each predictor is added to and removed from the model, so that the contribution of time since last reward is assessed by its effect on performance rather than by its normalized weight. We will additionally add a complementary analysis of the leave hazard that estimates the contribution of each variable without requiring the classification framework.

      (4) The five-second post-exit window. The reviewer is right that we did not explain this choice, and right that it rests on an assumption. Our reasoning was that a single exit moment per trial leaves the decision boundary badly under-constrained, and that treating the moments immediately following an exit as conditions under which the animal would also have left is licensed by the fact that within-patch reward rate declines monotonically with time, so conditions in the counterfactual continued visit would have been strictly less favorable than those already rejected. We accept that this is an assumption rather than an observation and will state it as such. We will also report the analysis across a range of window durations so that the independence of the result on this choice is visible. To the reviewer's specific questions: both time since entry and time since last reward are computed with respect to the time-investment port throughout, and we will state this explicitly.

      (5) False positive rate against ramps and random walks. We agree this is the more informative validation, and that continuous ramping is the most relevant alternative to ours. We will generate simulated units with continuous ramp-like rate changes, matched to the firing rates and interval structure of our recorded units, and pass them through the identical classification pipeline to obtain false positive rates comparable to the Poisson analysis already reported. We will retain the flat-rate Poisson simulation and present the ramp results as additional panels of the same supplement. We will also explore random-walk dynamics. Together with the held-out validation of transition times described under Reviewer #2(c), this converts the step characterization from a single-null validation into a comparison against the relevant continuous alternatives.

      (6) Specificity of the accumulation signal versus a generic increasing function. This is a valuable challenge and we will address it directly. We will implement the trial-mismatch control the reviewer proposes, randomly reassigning accumulation trajectories to reward-to-exit intervals within session and showing the extent to which predictive performance degrades relative to the correctly matched case.

      We would also note two features of the existing results that speak to this concern, and which we will bring forward in the revision because we did not make them salient enough. First, a signal that merely tracked elapsed time would be expected to shift its starting level as well as its rate across trials with different exit times; instead, the accumulation slope is strongly related to exit time (mean r = -0.551) while the intercept is not (mean r = 0.013), indicating a variable rate from a stable origin. Second, and more to the point, the accumulation arrives at a common level at the moment of exit whether the animal leaves early or late. The rate of accumulation shifts with the animal’s policy on that trial such that the threshold is met at the intended time. 

      What makes this predictor non-trivial is its trial-by-trial correspondence to behavior, not simply that it increases over time. We will make this argument explicitly alongside the shuffle control.

      Summary

      To summarize the planned additions: (i) analysis of step-like activity and its accumulation at the context port, with the interpretive caveat noted above; (ii) validation of step detection against ramping alternatives, together with held-out estimation of transition times; (iii) a trial-mismatch control for the specificity of the accumulation predictor; (iv) an inclusion-and-ablation analysis of the behavioral predictors and a complementary hazard model of the leave decision; (v) reporting of unit yield, step-like fraction, and sensitivity of the accumulation results to the inclusion threshold; (vi) analysis of the exit-to-context-entry latency in relation to the neural signal; (vii) a formal statistical test of overstaying relative to the computed optimum; and (viii) substantial clarification of the task specification, the operational definition of exit, and the derivation of the optimal residence time, including an expanded Figure 1h legend.

      Our aim in the revision is to meet the specificity concern raised in the assessment as directly as the existing data allow, and we hope the revised manuscript will warrant reconsideration of the strength-of-evidence characterization.

      We are grateful to the reviewers for the care evident in their reports, and to you both for handling the manuscript.

      References

      Charnov, E. L. (1976). Optimal foraging, the marginal value theorem. Theoretical Population Biology, 9(2), 129–136.

      Sutlief, E., Walters, C., Marton, T., & Hussain Shuler, M. G. (2025). The value of initiating a pursuit in temporal decision-making. eLife. https://doi.org/10.7554/eLife.99957.2.

    1. eLife Assessment

      This important study uses longitudinal EEG to chart how neural tracking of syllables and word-level statistical structure develops over the first two years of life in infants at high and low likelihood for autism and links these measures to verbal outcomes at 18-20 months. The strength of evidence is convincing: the prospective longitudinal design, careful data-quality handling, and partial least squares analyses are appropriate and well executed, though some interpretations of the group differences in syllable tracking, along with the possible contributions of multilingual exposure and sleep state during recording, warrant caution. The work will be of interest to developmental cognitive neuroscientists studying language acquisition and early neural markers of neurodevelopmental conditions.

    2. Reviewer #1 (Public review):

      Summary:

      This manuscript reports a prospective longitudinal study examining whether infants with high likelihood (HL) for autism differ from low-likelihood (LL) infants in two levels of word learning: brain-to-speech cortical entrainment and implicit word segmentation. The authors report reduced syllable tracking and post-learning word recognition in the HL group relative to the LL group. Importantly, both the syllable-tracking entrainment measure and the word recognition ERP measure are positively associated with verbal outcomes at 18-20 months, as indexed by the Mullen Verbal Developmental Quotient. Overall, I found this to be a thoughtfully designed and carefully executed study that tackles a difficult and important set of questions. With some clarifications and modest additional analyses or discussion on the points below, the manuscript has strong potential to make a substantial contribution to the literature on early language development and autism.

      Strengths:

      This is an important study that addresses a central question in developmental cognitive neuroscience: what mechanisms underlie variability in language learning, and what are the early neural correlates of these individual differences? While language development has a relatively well-defined sensitive period in typical development, the mechanisms of variability-particularly in the context of neurodevelopmental conditions-remain poorly understood, in part because longitudinal work in very young infants and toddlers is rare. The present study makes a valuable contribution by directly targeting this gap and by grounding the work in a strong theoretical tradition on statistical learning as a foundational mechanism for early language acquisition.

      I especially appreciate the authors' meticulous approach to data quality and their clear, transparent description of the methods. The choice of partial least squares correlation (PLS-c) is well motivated, given the multidimensional nature of the data and collinearity among variables and the manuscript does a commendable job explaining this technique to readers who may be less familiar with it.

      The results reveal interesting developmental changes in syllable tracking and word segmentation from birth to 2 years in both HL and LL infants. Simply mapping these trajectories in both groups is highly valuable. Moreover, the associations between neural indices of brain-to-speech entrainment and word segmentation with later verbal outcomes in the LL group support a critical role for speech perception and statistical learning in early language development, with clear implications for understanding autism. Overall, this is a rich dataset with substantial potential to inform theory.

      Comment on revised version.

      The revised manuscript has provided additional analyses that lead to critical clarification of the main findings, including the longitudinal nature of the relationship between neural tracking of speech and language, the role of sleep, and the potential modulation effect of stream structure on syllable-level neural tracking. The overall results highlight the robustness of the findings as well as the specific relevance of the structured speech tracking to verbal outcomes of infants with high likelihood (HL) of autism.

    3. Reviewer #2 (Public review):

      Summary:

      This article looks at differences in how the brain entrains to, or tracks, the rhythmic presentation of syllables and words in speech in infants at increased likelihood versus low likelihood for autism. The authors first sought to characterize how brain responses are modulated by learning the statistical probability of a given syllable following the one before it over the first two years of life. They then sought to identify at which stages of word learning infants at increased likelihood for autism showed difficulties, and whether those difficulties worsened over time. Finally, they sought to indicate whether infants' statistical learning and word learning abilities could predict later verbal skills. The authors found similar developmental trajectories of neural entrainment to syllables in infants at high and low likelihood for autism, but infants at high likelihood for autism had overall weaker syllable-level entrainment. Infants at high versus low likelihood for autism showed different developmental trajectories for word entrainment. Lower syllable entrainment in high-likelihood infants corresponded with poorer verbal outcomes, but word entrainment was not associated with verbal outcomes. Event-related potential responses to words and part words were positively associated with verbal outcomes, however, but only in low-likelihood infants.

      Strengths:

      Overall, the article provides rigorous statistical analysis of longitudinal EEG data to provide strong support for the claims that neural entrainment to syllable and word features of speech may be a useful marker for language development difficulties, particularly in infants at increased likelihood for neurodevelopmental disorders. The EEG data collection and preprocessing procedures are well within standards within the field. Readers should take care to note that authors indexed neural entrainment to speech using phase-locking values instead of spectral power.

      Comments on revised version.

      While the statistical analyses are rigorous, there are a few potential confounds to the results. The authors now do a nice job addressing these limitations to the work. For example, sleep status may modulate some of the biomarkers relevant for language learning. Exposure to additional languages may influence performance on the verbal assessment, though the authors do clarify that participants came from majority French-speaking households. As a result, readers should be encouraged to interpret that neural entrainment to speech features is likely a useful mechanism to explain differences in language development, while taking this interpretation with some caution.

    4. Author response:

      The following is the authors’ response to the current reviews.

      Reviewer #1 (Recommendations for the authors):

      (1) Interpretation of Syllable-Tracking in the RND Condition:

      The finding of greater syllable-tracking in the LL group compared to the HL group in the RND condition warrants cautious interpretation. Currently, there is no direct statistical evidence demonstrating greater PLV at 4 Hz in the Structured versus Random conditions for either group; readers must infer this solely from numeric differences in Figure S5 B and D. Therefore, while the interpretation on Page 14 (Lines 443-446) "successful segmentation may enhance syllable tracking via top-down predictions of the next syllable" is an interesting speculation, it feels somewhat far-reaching. Additionally, the authors should discuss whether this upregulated syllable tracking in the structured condition (which is specific to the HL group) represents an adaptive or maladaptive response.

      The reviewer correctly highlights the lack of direct  comparison between conditions (RND versus STR). We tempered our claims in the cited paragraph and insisted on the speculative nature of this part of the discussion. We also clarified that, to us, it may represent an adaptive compensatory strategy:

      Page 14, line 441: “Interestingly, our supplementary analyses (Supplementary Material Figure S4-5) suggest that syllable entrainment may be differentially affected in HL versus LL infants, depending on the statistical structure of the input stream (RND versus STR). However, as our experiment was not explicitly designed to test stream effects, these results should be interpreted with caution. Future studies could explore how successful segmentation may enhance syllable tracking via top-down predictions of the next syllable in both LL and HL infants. If confirmed, such a mechanism may improve alignment to syllable onsets, potentially constituting a compensatory process allowed by preserved segmentation abilities.”

      (2) Preservation of Statistical Learning in HL Infants:

      The text added on Pages 17-18 (Lines 562-566) regarding a "heightened dependence on bottom-up mechanisms (in autism)" does not appear to be supported by the data or by theories of implicit statistical learning. Because greater syllable-level entrainment was observed in the LL group than the HL group across both the random and structured conditions, the data actually point toward impaired bottom-up processes. Furthermore, implicit statistical learning typically involves an interplay of both bottom-up and top-down mechanisms; the implicit nature of a task does not guarantee a strictly bottom-up process. Consequently, this interpretation is not entirely convincing.

      We agree with the reviewer that the concepts of “top-down” and “bottom-up” were not fully appropriate to support our point in the cited paragraph. We should have used the concepts of implicit versus explicit learning instead, in line with previous literature suggesting increased reliance on preserved implicit learning in autism to compensate for altered explicit processes. The paragraph was slightly modified.

      Page 18, line 564: “According to these studies, autistic impairments in explicit attentional processes, such as social orienting - which are critical for bootstrapping language acquisition (70) - may result in a heightened dependence on implicit mechanisms, including statistical learning. As previously discussed, preserved word segmentation abilities may further compensate for alterations in lower-level implicit processes, such as syllable tracking.”

      Reviewer #2 (Recommendations for the authors):

      Potential typo on line 199 - I think an apostrophe is needed here.<br /> Potential typo on line 255 - do you mean Central electrodes?

      We addressed the typos spotted by reviewer.

      Line 199: variables’

      Line 255: Centro-frontal electrodes


      The following is the authors’ response to the original reviews

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This manuscript reports a prospective longitudinal study examining whether infants with high likelihood (HL) for autism differ from low-likelihood (LL) infants in two levels of word learning: brain-to-speech cortical entrainment and implicit word segmentation. The authors report reduced syllable tracking and post-learning word recognition in the HL group relative to the LL group. Importantly, both the syllable-tracking entrainment measure and the word recognition ERP measure are positively associated with verbal outcomes at 18-20 months, as indexed by the Mullen Verbal Developmental Quotient. Overall, I found this to be a thoughtfully designed and carefully executed study that tackles a difficult and important set of questions. With some clarifications and modest additional analyses or discussion on the points below, the manuscript has strong potential to make a substantial contribution to the literature on early language development and autism.

      Strengths:

      This is an important study that addresses a central question in developmental cognitive neuroscience: what mechanisms underlie variability in language learning, and what are the early neural correlates of these individual differences? While language development has a relatively well-defined sensitive period in typical development, the mechanisms of variability - particularly in the context of neurodevelopmental conditions - remain poorly understood, in part because longitudinal work in very young infants and toddlers is rare. The present study makes a valuable contribution by directly targeting this gap and by grounding the work in a strong theoretical tradition on statistical learning as a foundational mechanism for early language acquisition.

      I especially appreciate the authors' meticulous approach to data quality and their clear, transparent description of the methods. The choice of partial least squares correlation (PLS-c) is well motivated, given the multidimensional nature of the data and collinearity among variables, and the manuscript does a commendable job explaining this technique to readers who may be less familiar with it.

      The results reveal interesting developmental changes in syllable tracking and word segmentation from birth to 2 years in both HL and LL infants. Simply mapping these trajectories in both groups is highly valuable. Moreover, the associations between neural indices of brain-to-speech entrainment and word segmentation with later verbal outcomes in the LL group support a critical role for speech perception and statistical learning in early language development, with clear implications for understanding autism. Overall, this is a rich dataset with substantial potential to inform theory.

      Weaknesses:

      (1) Clarifying longitudinal vs. concurrent associations

      Because the current analytical approach incorporates all time points, including the final visit, it is challenging to determine to what extent the brain-language associations are driven by longitudinal relationships vs. concurrent correlations at the last time point. This does not undermine the main findings, but clarifying this issue could significantly enhance the impact of the individual-differences results. If feasible, the authors might consider (a) showing that a model excluding the final visit still predicts verbal outcomes at the last visit in a similar way, or (b) more explicitly acknowledging in the discussion that the observed associations may be partly or largely driven by concurrent correlations. Either approach would help readers interpret the strength and nature of the longitudinal claims.

      We thank the reviewer for this insightful comment. We agree that distinguishing between longitudinal predictive power and concurrent correlations at the final visit is crucial for clarifying the nature of these brain-language associations. Following the reviewer’s suggestion (a), we re-ran the two critical Partial Least Squares Correlation (PLS-c) analyses by excluding all EEG and behavioral data from the final 18–21 month visit (n = 54 recordings kept) to test whether earlier trajectories still predict the final verbal outcome.

      (1) Syllable entrainment (4 Hz) (original analysis on Figure 2C–D): The PLS-c restricted to the 3- to 15-month visits still identified a single significant component (p=.001, r=.56, 64.0% explained covariance, Figure 2 -figure supplement 3). Bootstrap ratios (BSR) were: contrast (low vs. high autism likelihood) 4.1; mean age −1.3; contrast*mean-age −2.1; delta-age 10.5; contrast*delta-age 2.0; age<sup>2</sup> −6.0; contrast*age<sup>2</sup> −5.8; and notably verbal outcome 7.1; contrast*verbal-outcome −6.3.

      The latent component and its spatial electrode configuration remain highly consistent with the original analysis (Figure 2C–D). This confirms that excluding the final visit preserves the model’s predictive validity: lower syllable entrainment correlates with poorer verbal outcomes at 18–21 months, particularly in the high-likelihood group.

      (2) Late evoked response to novel words (original analysis on Figure 6): The PLS-c analysis on the ERP late time window (1500–3000 ms), excluding the final visit, also revealed one significant component (p=.002, r=.74, 33.5% explained covariance, Figure 6 -figure supplement 1). Bootstrap ratios (BSR) were: contrast (part-word versus word) 14.5; mean age -5.2; contrast*mean-age 6.2; delta-age -0.8; contrast*delta-age 3.5; age<sup>2</sup> -0.8; contrast*age<sup>2</sup> -12.1; verbal-outcome 20.1; contrast*verbal-outcome -7.2. The latent component closely mirrored the original analysis (Figure 6), with frontal electrodes contributing negatively and posterior electrodes positively. Minor divergences in age-related parameter contributions were observed, likely due to the absence of 18-21 month timepoints, which previously contributed to the convex/concave shapes of the group age trajectories in figure 6B (left panel).

      Crucially, both models (with and without the final visit) positively predicted verbal outcomes (Figure 6B, right panel, and Figure 6 -figure supplement 1B). However, excluding the final visit reversed the direction of the group*verbal-outcome interaction (from 3 to -7.2): This indicates that after ruling out cross-sectional correlations at 18–21 months, the early predictive value of the late ERP to word novelty is more prominently observed in high-likelihood infants, suggesting that the original result was influenced by concurrent cross-sectional correlations at the final visit. This aligns with the syllable entrainment findings (Figure 2 -figure supplement 3), as both 4 Hz neural tracking and late ERP responses to novelty predominantly predict verbal outcomes in infants at high likelihood for autism.

      We reported these supplementary analyses in the revised manuscript as follows:

      We added Figure 2 -figure supplement 3 and Figure 6 -figure supplement 1. In general, most of figures that were present in Supplementary materials were moved as figure supplements to enhance readability.

      Page 8 lines 234-240 (pages and lines refer to the reviewed uploaded manuscript): “To rule out the possibility that the association between syllable entrainment and verbal outcome was driven by concurrent measures taken at 18–21 months, we re-ran the PLS-c analysis excluding EEG data from the final visit (n = 54 recordings kept). The resulting latent component remain significant (p = .001) and showed contributions from behavioral and EEG variables that were highly similar to those observed in the previous analysis, with a verbal outcome BSR of 7.1 and a group’verbal-outcome interaction BSR of −6.3 (Figure 2 -figure supplement 3).”

      Page 12 lines 387-394: “As we did for neural entrainment to syllables, we conducted a new analysis on late ERP to word novelty, excluding EEG data from the final visit. This PLS-c yielded one significant latent component (p = .002, r=.74, 33.5% explained covariance, Figure 6 -figure supplement 1) with globally similar EEG parameter contributions and age trajectory modelling. Verbal outcome still significantly contributed to the latent component (BSR=20.1), with a negative verbal outcome*group interaction (BSR=-7.2). These results suggest that, after ruling out cross-sectional correlations at 18–21 months, the late ERP to word novelty predominantly predicts verbal outcomes in high-likelihood infants for autism.”

      Page 17 lines 547-548: “As with syllable entrainment, the late ERP to novel words primarily predicted verbal outcomes in high-likelihood (HL) infants.”

      Page 18 lines 588-590: “Likewise, the absence of a late ERP orientation response in HL participants may represent an early neural signature of altered attention to novelty that can be used both as a non-invasive predictor of language development and as a potential target for early intervention.’

      (2) Incorporating sleep status into longitudinal models

      Sleep status changes systematically across developmental stages in this cohort. Given that some of the papers cited to justify the paradigm also note limitations in speech entrainment and word segmentation during sleep or in patients with impaired consciousness, it would be helpful to account for sleep more directly. Including sleep status as a factor or covariate in the longitudinal models, or at least elaborating more fully on its potential role and limitations, would further strengthen the conclusions and reassure readers that these effects are not primarily driven by differences in sleep-wake state.

      The reviewer is highlighting here a limitation of our study design that comprised sleeping status that varied from one timepoint to another among participants. To rule out any confounding effect of wake status (coded as a binary variable: sleeping or awake during recording) on analyses comparing groups, a linear mixed-effect model with repeated measures was fitted finding no significant difference between high- and low-likelihood participants (p=.769, reported at page 20, lines 646-647). However, as rightly suggested by the reviewer, this doesn’t prevent from a sleep bias on age trajectories, especially given that sleeping status significantly decreases with age in our sample.

      Including sleep status as a covariate in our analyses, as suggested by the reviewer, would be difficult to implement in our PLS-c methods, since a categorical behavioral parameter that varies within participants is not possible in the models provided by myPLS toolbox.

      As an alternative option, we re-ran all analyses that explored the condition effect on the whole sample within the sleeping participants only (n=25 recordings) to confirm that the same age-trajectories of EEG parameters were highlighted. However, negative results should be interpreted with caution since the sample is small for such a multivariate approach, resulting in modest statistical power.

      (1) Syllable entrainment (4 Hz) (original analysis on Figure 2A–B): The PLS-c identified one significant component (p <.001, r = .78, 85.1% explained covariance, Figure 2 -figure supplement 2 and Figure 3 -figure supplement 1). Bootstrap ratios (BSR) were: contrast (4hz vs. adjacent frequencies) 30.3; mean age -2.9; contrast*mean-age -2.4; delta-age 3.8; contrast*delta-age 3.4; age<sup>2</sup> -1.1; contrast* −2.5. The spatial distribution of contributing electrodes globally matched that shown in Figure 2A. The high contrast BSR (30.3) confirms robust syllable entrainment in sleeping infants. Critically, the contrast*age<sup>2</sup> parameter contributed negatively to the latent component (BSR = −2.5), confirming that the convex age trajectory of syllabic entrainment (Figure 2B) is also present in the sleeping subsample.

      (2) Word entrainment (1.3 Hz) (original analysis on Figure 3A–B): The PLS-c identified one significant component (p <.001, r = .63, 37.3% explained covariance, Author response image 1). Bootstrap ratios (BSR) were: contrast (1.3hz vs. adjacent frequencies) 24.9; mean age -7.5; contrast*mean-age -3.2; delta-age 2.9; contrast*delta-age 0.0; age<sup>2</sup> 1.5; contrast* 1.1. The spatial distribution of significant electrodes partially overlaps with the ones in the original analysis, primarily showing fronto-central positive contribution to the latent component. The high contrast BSR confirms a robust word entrainment in sleeping participants, in line with previous studies (e.g., Flò et al, Sci Rep, 2022). However, the lack of a significant contrast* age<sup>2</sup> suggests that the U-shape age trajectory illustrated on Figure 3 might be modulated by wakefulness or due to a lack of power in the present analysis. A non-significant trend towards a U-shape pattern with a 12-month nadir is visible in sleeping participants, but additional data from sleeping 18-21 months sleeping infants would be required to confirm or refute this trend.

      (3) Early evoked response to novel words (original analysis on Figure 4): The PLS-c analysis on the ERP early time window (0–1000 ms) in sleeping participants revealed no significant component. The absence of early response to word novelty in sleeping participant might account for the lack of response observed in the whole sample, illustrated on Figure 4. To test this hypothesis, we conducted the same PLS-c in awake participants (n=58 recordings), which also yielded no significant latent component. This suggests that the lack of a measurable early response to word novelty observed in the whole sample is consistent across both sleeping and awake infants, and not driven by any of the two subsamples.

      (4) Late evoked response to novel words (original analysis on Figure 5): The PLS-c analysis on the ERP late time window in sleeping participants revealed no significant component. This suggests that sleeping participants might present a reduced or even absent late response to novel words. Given this identified effect of sleep on late ERP response, we reran the PLS-c on the late ERP window using group as contrast (original analysis on figure 6), excluding the sleeping participants to avoid any confounds. This PLS-c revealed one significant component (p = .006, r = .68, 30.5% explained covariance, Author response image 1). Bootstrap ratios (BSR) were: contrast (low versus high likelihood) 7.6; mean age -5.9; contrast*mean-age -3.8; delta-age 2.5; contrast*delta-age -3.5; age<sup>2</sup> -2.5; contrast*age<sup>2</sup> -4.9; verbal-outcome 8.5; contrast*verbal-outcome -1.0. Behavioral parameters contribute to this latent component with similar magnitude and polarity as in the original analysis. Electrode contributions are also highly consistent, with frontal negative and posterior positive contributions. This confirms that sleeping participants, despite their potentially reduced late response, did not significantly bias the results presented in Figure 6.

      Author response image 1.

      Late evoked response potential (ERP) to word novelty in awake participants. A. Design and brain saliences derived from the significant latent component. Brain topographies of bootstrap ratios (BSR) are displayed at 250ms intervals. Black dots indicate BSR > 2.3. B. Participants’ brain scores for part-word and word conditions, as a function of age (left panel) and verbal DQ (right panel). For details on brain scores, see Figure 6 -figure supplement 1. Linear fitting is used for illustrative purposes only. HL: high likelihood for autism; LL: low likelihood for autism.

      We reported these analyses in the revised manuscript as follows:

      We added Figure 2 -figure supplement 2A and Figure 3 -figure supplement 1.

      Page 7, lines 218-224: “Because some infants were asleep during the recording session, particularly at younger ages, we performed a supplementary control analysis restricted to this sleeping subsample (n = 25 recordings, Figure 2 -figure supplement 2). This PLS-c also identified a significant latent component (p < .001, r = .78, 85.1% explained covariance), with a significant contrast effect (BSR = 30.3) and a significant negative contrast*age<sup>2</sup> interaction (BSR = −2.5). These findings confirm that the convex age trajectory observed in the main analysis remains present and observable even in sleeping infants.”

      Page 8, lines 252-259: “We further investigated word entrainment in sleeping participants (n=25), which yielded one significant latent component (p<.001, r=.63, 37.3% explained covariance, Figure 3 -figure supplement 1). Centro-frontal electrode contributed to this component, with a high contrast BSR (24.9), confirming a similar word entrainment pattern in the sleeping subsample. The contrast*age<sup>2</sup> was also positive but not significant (1.1), suggesting a trend toward a U-shape age trajectory with a 12-month nadir in sleeping infants. Additional 18-21 month recording would be required to confirm this trend.”

      Page 11, lines 347-349: “The same PLS-c, conducted separately in sleeping (n=25) and awake subsamples (n = 58), yielded no significant latent component, indicating a consistent absence of early response to word novelty in both sleeping and awake infants.”

      Page 11-12 lines 368-370: “The same PLS-c in the sleeping subsample yielded no significant latent component, suggesting that sleep may reduce or even abolish the late response to word novelty.”

      Page 12 lines 382-385: “Given that no late response was detected in sleeping participants, we re-ran the PLS-c analysis using group as a contrast in the awake subsample (n=58). This yielded one significant latent component (p=.006, r=.68, 30.5% explained covariance), with behavioral and electrode contributions highly overlapping with those in Figure 6.”

      Page 18 lines 590-592: “This potential biomarker might nevertheless be modulated by participants’ sleep status, warranting careful consideration of vigilance state in future studies.”

      (3) Use of PLS-c and potential group × condition interactions

      I am relatively new to PLS-c. One question that arose is whether PLS-c could be extended to handle a two-way interaction between group and condition contrasts (STR vs. RND). If so, some of the more complex supplementary models testing developmental trajectories within each group (Page 8, Lines 258-265) might be more directly captured within a single, unified framework. Even a brief comment in the methods or discussion about the feasibility (or limitations) of modeling such interactions within PLS-c would be informative for readers and could streamline the analytic narrative.

      The reviewer raises a valid concern regarding the capacity of PLS-c to accommodate multi-way interactions among categorical and continuous variables. While PLS-c has no inherent theoretical constraints on the number of predictor terms (they can even exceed the sample size in number), practical limitations arise from model stability and interpretability when the ratio of predictors to sample size becomes excessive. As noted by Geladi and Kowalski (1986), exceeding ~10% of the sample size with predictors increases noise sensitivity and overfitting.

      In our study, the PLS-c analyses already reach this ~10% limit, with a maximum of nine predictors for a sample size of n=83. Attempting to integrate both group and condition as contrasts — along with necessary age parameters to account for developmental trajectories — would result in 12 predictors (or 15 if verbal outcome is included). Specifically, the model would require behavioral terms for Group, Condition, Group*Condition, Mean-age, Group*Mean-age, Condition*Mean-age, Delta-age, Group*Delta-age, Condition*Delta-age, Age<sup>2</sup>, Group*Age<sup>2</sup>, Condition*Age<sup>2</sup>, Verbal-outcome, Group*Verbal-outcome, and Condition*Verbal-outcome.

      Although a unified multivariate model capturing the complex dynamics at play in our sample is theoretically appealing, the substantial risk of overfitting precludes its feasibility. Therefore, we opted to use only one categorical predictor per PLS-c analysis to maintain model parsimony and reliability. However, a larger sample could overcome this limitation, allowing a stable and unified model of longitudinal EEG data that simultaneously captures age trajectories, group, clinical outcome, and condition.

      Reference:

      Geladi, P., & Kowalski, B. (1986). Partial least-squares regression: A tutorial. Analytica Chimica Acta, 185, 1–17. https://doi.org/10.1016/S0003-2670(00)82582-3

      We added the following comment in the method section:

      Page 24, lines 773-777: “We limited the number of behavioral variables to nine to mitigate noise sensitivity and overfitting risks associated with exceeding the 10% sample size threshold (Geladi & Kowalski, 1986). This limitation precluded the implementation of a single PLS-c model incorporating group, condition (STR vs. RND), age, and their interactions.”

      (4) STR-only analyses and the role of RND

      Page 8, Lines 241-245: This analysis is conducted only within the STR condition. The lack of group difference observed here appears consistent with the lack of group difference in word-level entrainment (Page 9, Lines 292-294), suggesting that HL and LL groups may not differ in statistical learning per se, but rather in syllabic-level entrainment. As a useful sanity check and potential extension, it might be informative to explore whether syllable-level entrainment in the RND condition differs between groups to a similar extent as in Figure 2C-D. In other work (e.g., adults vs. children; Moreau et al., 2022), group differences can be more pronounced for syllable-level than for word-level entrainment. Figure S6 seems to hint that a similar pattern may exist here. If feasible, including or briefly reporting such an analysis could help clarify the asymmetry between the two learning measures and further support the interpretation of syllabic-level differences.

      The reviewer points to the interesting pattern highlighted in supplementary figure S6, suggesting that group differences in syllabic entrainment might be modulated by the structure of the stream (STR versus RND). Such modulatory effect of stream structure on entrainment to syllables has been suggested by many studies, like Moreau et al (2022), as pointed by the reviewer, and seems at play in our sample, as illustrated on supplementary figure S5 (decline in the 4hz PLV that exceeds the size of confidence intervals, ~90 s after STR onset).

      Following the reviewer’s suggestion, we ran a PLS-c testing group effect on 4hz PLVs in each stream:

      (1) in the RND stream: the analysis yields one significant component (p<.001, r=.49, 52.7% explained covariance, Author response image 2A-B). Bootstrap ratios (BSR) are: contrast (low versus high likelihood) 6.9; mean age -1.0; contrast*mean-age 0.1; delta-age 11.8; contrast*delta-age 0.1; age<sup>2</sup> -4.9; contrast*age<sup>2</sup> 4.6; verbal-outcome 9.7; contrast*verbal-outcome -0.6. Interestingly, the model still highlights a strong link between syllable tracking and group, suggesting that RND also discriminate between HL and LL. However, RND syllable tracking doesn’t appear to be linked to group x verbal-outcome as we observed in Figure 2C-D.

      (2) In the STR stream, we obtained one significant latent component (p=.002, r=.51, 57.8% explained covariance, Author response image 2C-D). Bootstrap ratios (BSR) are: contrast 2.2; mean age -1.6; contrast*mean-age -1.1; delta-age 6.1; contrast*delta-age -0.7; age<sup>2</sup> -4.4; contrast*age<sup>2</sup> 0.5; verbal-outcome 8.2; contrast*verbal-outcome -7.3. Here, the strong association between syllable tracking and group x verbal-outcome is similar to the model presented in Figure 2C-D.

      Taken together, these results suggest that the apparent STR/RND dissociation illustrated in Figure S6 might primarily reflect a Group*Verbal-outcome divergence, with syllable tracking in the STR stream being related to verbal outcome mainly in high likelihood for autism.

      Author response image 2.

      Syllable entrainment within RND (A-B) and STR (C-D).

      These results were reported in the revised manuscript in the Result section (Time course of the entrainment along experiment subheader), implying a slight reframing of the result presentation of supplementary analysis S6. Author response image 2 was added in supplementary material as Figure S5.

      Page 10, lines 307-319: “The group, age and verbal outcome parameters were mainly correlated (BSR>2.3) with the neural entrainment occurring~90 seconds after the onset of the STR stream, coinciding with the time participants began tracking word boundaries (Supplementary material, S3). This result suggests that the group differences in syllable entrainment, as shown in Figure 2C-D, as their associations with verbal outcome, are modulated by the structure of the stream (STR versus RND). We ran one additional PLS-c for each stream separately, using group as contrast. In both streams, the PLS-c yielded a significant LC (p<.001 for RND and p=.002 for STR), with a positive group effect (BSR>2.3) in both LC (Supplementary material, S5). Most strikingly, the group*verbal outcome parameter reached significance exclusively within the STR latent component (BSR:-7.3). These results suggest that while syllable tracking is generally decreased in HL infants across both streams, its association with verbal outcome is prominently driven by the stream containing words (STR).”

      Page 14, lines 443-446: “This temporal overlap suggests that successful segmentation may enhance syllable tracking via top-down predictions of the next syllable, improving alignment to syllable onsets in LL infants as well as in HL with better verbal outcome.’

      (5) Multi-speaker input and voice perception (Page 15, Lines 475-483)

      The multi-speaker nature of the speech input is an interesting and ecologically relevant feature of the design, but it does add interpretive complexity. The literature on voice perception in autism is still mixed: for example, Boucher et al. (2000) reported no differences in voice recognition and discrimination between children with autism and language-matched non-autistic peers, whereas behavioral work in autistic adults suggests atypical voice perception (e.g., Schelinski et al., 2016; Lin et al., 2015). I found the current interpretation in this paragraph somewhat difficult to follow, partly because the data do not directly test how HL and LL infants integrate or suppress voice information. I think the authors could strengthen this section by slightly softening and clarifying the claims.

      We acknowledge the reviewer’s concern regarding the potential ambiguity in the cited paragraph. To address this, we have revised the text to explicitly clarify the aims of our study and its design. Furthermore, we now emphasize the speculative and post-hoc nature of the hypotheses and interpretations presented, thereby ensuring transparency regarding the limitations of our findings.

      Page 16 lines 520-530), as follows: “HL infants, on the other hand, did not show this transient disruption. In this group, word entrainment remained stable over time. To account for this unexpected finding, we followed up on the post-hoc hypothesis proposed above: a reduced sensitivity to social and vocal cues observed in HL infants may have spared segmentation abilities by limiting the interference introduced by speaker variability. If this post-hoc hypothesis holds true, LL and HL infants would differ not in their intrinsic ability to learn statistical regularities per se, but rather in how they integrate or suppress competing cues (such as speaker changes) during the segmentation process. It is important to note, however, that the present study was not designed to isolate and evaluate the specific impact of speaker changes on word segmentation. Consequently, this interpretation remains speculative, and additional research is required to further address this question.”

      (6) Asymmetry between EEG learning measures

      Page 16, Lines 502-507 touches on the asymmetry between the two EEG learning measures but leaves some questions for the reader. The presence of word recognition ERPs in the LL group suggests that a failure to suppress voice information during learning did not prevent successful word learning. At the same time, there is an interesting complementary pattern in the HL group, who show LL-like word-level entrainment but does not exhibit robust word recognition. Explicitly discussing this asymmetry - why HL infants might show relatively preserved word-level entrainment yet reduced word recognition ERPs, whereas LL infants show both - would enrich the theoretical contribution of the manuscript.

      We concur with the reviewer’s observation that our findings imply a theoretically significant double dissociation between HL and LL groups, specifically concerning the asymmetries between word-level neural entrainment and word recognition mechanisms. We believe this point was partly addressed in the subsequent paragraph, where we stated that “in contrast” to LL, HL infants “showed no clear ERP difference between novel and familiar triplets”, while “both groups showed similar word neural entrainment during learning”. We further explored potential explanations for this apparent dissociation, such as a possible deficit in novelty orientation that may be specific to HL infants and unrelated to statistical learning itself. We cited Liu et al (2023) as a reference showing the dissociation between mechanisms underlying implicit versus explicit traces of statistical learning. We acknowledge that we can discuss more in depth the potential preservation of statistical learning in HL infants. We have incorporated the following discussion in the reviewed manuscript, supported by relevant references:

      Pages 17-18, lines 562-566: “Interestingly, this dissociation between spared implicit versus impaired explicit statistical learning in autism has been previously discussed in the literature (Zwart et al, 2018, Kissine, 2021). According to these studies, autistic impairments in top-down attentional processes, such as social orienting — which are critical for bootstrapping language acquisition (Kuhl, 2007) — may result in a heightened dependence on bottom-up mechanisms, including implicit statistical learning.”

      References:

      Zwart, F.S., Vissers, C.T.W.M., Kessels, R.P.C. and Maes, J.H.R. (2018), Implicit learning seems to come naturally for children with autism, but not for children with specific language impairment: Evidence from behavioral and ERP data. Autism Research, 11: 1050-1061. https://doi.org/10.1002/aur.1954

      Kissine, M. (2021). Autism, constructionism, and nativism. Language 97(3), e139-e160. https://dx.doi.org/10.1353/lan.2021.0055.

      Kuhl, P.K. (2007), Is speech learning ‘gated’ by the social brain?. Developmental Science, 10: 110-120. https://doi.org/10.1111/j.1467-7687.2007.00572.x

      References:

      (1) Moreau, C. N., Joanisse, M. F., Mulgrew, J., & Batterink, L. J. (2022). No statistical learning advantage in children over adults: Evidence from behaviour and neural entrainment. Developmental Cognitive Neuroscience, 57, 101154. https://doi.org/10.1016/j.dcn.2022.101154

      (2) Boucher, J., Lewis, V., & Collis, G. M. (2000). Voice processing abilities in children with autism, children with specific language impairments, and young typically developing children. Journal of Child Psychology and Psychiatry, 41(7), 847-857. https://doi.org/10.1111/1469-7610.00672

      (3) Schelinski, S., Borowiak, K., & von Kriegstein, K. (2016). Temporal voice areas exist in autism spectrum disorder but are dysfunctional for voice identity recognition. Social Cognitive and Affective Neuroscience, 11(11), 1812-1822. https://doi.org/10.1093/scan/nsw089

      (4) Lin, I.-F., Yamada, T., Komine, Y., Kato, N., Kato, M., & Kashino, M. (2015). Vocal identity recognition in autism spectrum disorder. PLOS ONE, 10(6), e0129451.https://doi.org/10.1371/journal.pone.0129451

      Reviewer #2 (Public review):

      Summary:

      This article looks at differences in how the brain entrains to, or tracks, the rhythmic presentation of syllables and words in speech in infants at increased likelihood versus low likelihood for autism. The authors first sought to characterize how brain responses are modulated by learning the statistical probability of a given syllable following the one before it over the first two years of life. They then sought to identify at which stages of word learning infants with increased likelihood of autism showed difficulties, and whether those difficulties worsened over time. Finally, they sought to indicate whether infants' statistical learning and word learning abilities could predict later verbal skills. The authors found similar developmental trajectories of neural entrainment to syllables in infants at high and low likelihood for autism, but infants at high likelihood for autism had overall weaker syllable-level entrainment. Infants at high versus low likelihood for autism showed different developmental trajectories for word entrainment. Lower syllable entrainment in high-likelihood infants corresponded with poorer verbal outcomes, but word entrainment was not associated with verbal outcomes. Event-related potential responses to words and part words were positively associated with verbal outcomes, however, but only in low-likelihood infants.

      Strengths:

      Overall, the article provides rigorous statistical analysis of longitudinal EEG data to provide strong support for the claims that neural entrainment to syllable and word features of speech may be a useful marker for language development difficulties, particularly in infants at increased likelihood for neurodevelopmental disorders. The EEG data collection and preprocessing procedures are well within standards in the field. Readers should take care to note that authors indexed neural entrainment to speech using phase-locking values instead of spectral power.

      Weaknesses:

      While the statistical analyses are rigorous, a few of the components of the models are not clearly defined, and some corrections and thresholds for significance warrant further justification. Further, a few stimuli and participant details that could influence results are not specified. It is not clear whether all participants came from majority French-speaking families; differences in the amount of French language exposure (compared to other languages that may be spoken by a participant's family) could influence results. The standardized volume of the stimuli is also not included. As a result, readers should be encouraged to interpret that neural entrainment to speech features is likely a useful mechanism to explain differences in language development, while taking this interpretation with some caution.

      We thank the reviewer for these remarks.

      Regarding the amount of French exposure: while all participants were raised in primarily French-speaking environments (i.e., French as the dominant language at home and daycare), the parental questionnaire at intake indicated that 45% of the sample was exposed to additional languages, reflecting Geneva’s highly multicultural demographics. We did not quantify the extent of this exposure, which could range from very occasional exposure to situations close to true bilingualism. The structural sensitivity hypothesis (Weiss et al., 2020) posits that additional language exposure may enhance detection of statistical structures in artificial language input, even when these structures differ from those in native languages. Yet, empirical support is mixed: Yim & Rudoy (2013) found no bilingualism effect in a paradigm close to ours (triplet segmentation via auditory statistical learning, n=112 children), whereas most studies reporting bilingual advantages for statistical learning involved tasks very distinct from ours, like artificial grammar and phonotactic rule learning, or multi-cue integration for segmentation (Weiss et al., 2020).

      Regarding the volume of stimuli, they were played at 50cm distance with an intensity of 75dB. Both considerations have been included in the new version of the manuscript. In general, we moved most of the figures present in Supplementary material to figure supplements to improve readability.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Minor Comments:

      Figure 6: The figure caption is not complete (there is no description for the right half of panel B).

      We thank the reviewer for this observation, figure 6 caption has been completed.

      Reviewer #2 (Recommendations for the authors):

      Broadly speaking, I would recommend reducing the number of abbreviations in this article, and I would recommend that the authors take care as to where these abbreviations are being introduced. Many of the abbreviated terms are defined in the Materials and Methods section, which is presented after the abbreviations are used in the main results.

      We acknowledge that our extensive use of abbreviations compromises the readability of the manuscript. Consequently, we have removed the following abbreviations:

      - SL (replaced by statistical learning)

      - LC (replaced by latent component)

      - ASD (replaced by autism)

      - TP (replaced by transition probability)

      - MEG (replaced by magneto-encephalogram)

      - MSEL (replaced by Mullen Scale of Early Learnings)

      The remaining abbreviations are:

      HL (high likelihood for autism), LL (low likelihood for autism), EEG (electroencephalogram), PLS-c (partial least square correlation), ERP (event-related potential), RND (random), STR (structured), BSR (bootstrap ratio), AIC (Akaike Information Criterion), PLV (phase locking value), DQ (developmental quotient), APSI (Autism Parent Screen for Infants).

      Moreover, we carefully reviewed how abbreviations were introduced and identified that PLS-c, STR and RND were not defined prior to the Method section. This oversight has been corrected in the reviewed manuscript.

      I would also recommend that the authors be careful with the structuring of the Introduction, particularly with their research questions and hypotheses. The article initially makes clear that the research questions are focused on the developmental trajectory of statistical learning, the levels of word learning that may differentiate high-likelihood versus low-likelihood infants, and the stability of those differences, and associations between statistical learning and various levels of word learning with verbal outcomes. The use of acoustic variability across syllables, while a valuable methodological tool, is somewhat presented as an additional research question, but not clearly stated or tested as such.

      We acknowledge that the introduction (particularly the paragraph from lines 173 to 184) may have implied that speaker variability across syllables was one of our primary research aims. We clarify here that speaker variability was introduced as a mean to increase task difficulty, particularly for high-likelihood (HL) participants, with the aim of amplifying the effect sizes in our analyses.

      To address this, we have removed the theoretical discussion on speaker variability in autism and typical development (lines 173–184) and explicitly stated that speaker variability was not a research question in this study. Crucially, our experimental design did not include a control condition without speaker variability, and thus we could not test its specific effects on statistical learning across age trajectories and groups.

      Page 6, lines 173-176 (pages and lines refer to the reviewed uploaded manuscript): “It is worth noting, however, that our study was not designed to isolate or quantify the specific impact of speaker variability on statistical learning, as the experimental design did not include a baseline control condition omitting this acoustic variation.”

      The authors do a nice job in the Materials & Methods explaining PLS-c and defining the latent components and bootstrapped ratios that will be shared in the Results. An additional brief iteration defining these statistical elements is needed at the beginning of the Results section.

      We thank the reviewer for their appreciation of our Method section. We agree that an additional iteration in the result section would improve readability. We added the following paragraph at the very beginning of the Result section, briefly defining PLS-c and its main statistical output (latent components and bootstrap ratios):

      Pages 6-7, lines 193-202: “Briefly, PLS-c is a data-driven multivariate modelling approach designed to identify significant patterns of electrode clusters (from a brain data matrix containing electrophysiological measures, here PLV) and their associations with “behavioral” variables (from a behavioral design matrix, here age-related parameters). Patterns of brain x behavior associations are called latent components, and their statistical significance is evaluated using permutation testing (n=1000, Bonferroni correction for number of components tested, alpha=.006). Brain and behavioral variables respective contributions to any significant latent component are tested with bootstrapping (500 random samples and replacement), with bootstrap ratios (BSR) greater than 2.3 indicating a stable contribution (for details, see the Materials and Methods section).”

      (1) Page 18 Line 576. The authors need to clarify whether participants were required to be in primarily French-speaking environments and whether there was a minimum amount of French language exposure that participants were required to have if they were exposed to additional languages besides French in their everyday life.

      The reviewer raises a valid concern regarding participants’ language exposure. In this study, all participants were raised in primarily French-speaking environments, with French as the dominant language at home and daycare. The parental questionnaire at intake indicated that 45% of the sample was exposed to additional languages, reflecting Geneva’s highly multicultural demographics. However, we did not quantify the extent of this exposure, which could range from very occasional exposure to situations close to true bilingualism.

      The structural sensitivity hypothesis (Weiss et al., 2020) posits that additional language exposure may enhance detection of statistical structures in artificial language input, even when these structures differ from those in native languages. Yet, empirical support is mixed: Yim & Rudoy (2013) found no bilingualism effect in a paradigm close to ours (triplet segmentation via auditory statistical learning, n=112 children), whereas most studies reporting bilingual advantages for statistical learning involved tasks very distinct from ours, like artificial grammar and phonotactic rule learning, or multi-cue integration for segmentation (Weiss et al., 2020).

      To include these considerations, Limitations and Material and methods sections were modified as follows:

      Page 19, lines 604-607: “Second, although all participants were primarily exposed to French, we did not quantify additional language exposure, precluding any analysis of its potential moderator effects on statistical learning in our groups and age-trajectories. However, prior work has reported no effect of bilingualism on auditory triplet segmentation in children (Yim & Rudoy, 2013).”

      Page 20, lines 630-631: “All participants were raised in primarily French-speaking environments, with French as the dominant language at home and daycare.”

      References:

      Weiss DJ, Schwob N, Lebkuecher AL. Bilingualism and statistical learning: Lessons from studies using artificial languages. Bilingualism: Language and Cognition. 2020;23(1):92-97. doi:10.1017/S1366728919000579

      Yim D, Rudoy J. Implicit statistical learning and language skills in bilingual children. J Speech Lang Hear Res. 2013 Feb;56(1):310-22. doi: 10.1044/1092-4388(2012/11-0243). Epub 2012 Aug 15. PMID: 22896046.

      (2) Page 18 Line 588. Further, the authors should clarify whether the 7 infants in the HL group, due to early parental concerns were defined by the 18-21-month APSI scores or by parental report prior to study enrollment.

      These 7 infants were recruited based on early parental concerns prior to intake. The APSI score at 18-21 months is only reported to provide an illustration of the amount of early autistic signs that were present in these 7 infants, and to provide an estimation of their probability to develop autism later on based on Sacrey et al., 2018 longitudinal study on the APSI predictive value. We agree with the reviewer that our phrasing suggests that the APSI was used as an inclusion criterion. We rephrased the page 20 lines 642-646 as follows:

      “The 7 other HL infants presented with early parental concerns for autism, based on parental report prior to enrollment. Their Autism Parent Screen for Infants (APSI) total score at their 18-21 months visit was 15.6±6.4, [8-22] range – a score greater than 8 reflecting a 63% positive predictive value for autism in HL populations.”

      (3) Page 20 Line 641. The authors should specify the volume of the stimuli.

      The volume of stimuli was reported in the main text (page 22, lines 695-696) as follows:

      “Stimuli were played on a Bose® Companion 2 Series III at a 50cm distance with an intensity of 75dB.”

      (4) I'd prefer Figure 1 to be reorganized slightly - at present, the placement of the arrows explaining the analysis steps is not intuitive.

      We addressed the reviewer’s comments (4) and (5) together as they both refer to Figure 1B.

      (5) Page 23 lines 718-719. I think it would be helpful to explicitly define each of the interaction variables included in the behavioral design matrix. Further, this matrix should be labeled consistently in both Figure 1B and in the main text.

      We refined figure 1B and its corresponding main text (in Methods section) for clarity. The arrows are now simpler and more parsimonious, labels (e.g., participant i, visit n, behavior design matrix and its parameters) are now standardized between the figure and the main text, and the interaction terms at lines 718-719 are explicitly defined.

      (6) Page 23 lines 726-731: It would be helpful to know whether applying a Bonferroni correction in addition to completing permutation testing is standard when evaluating latent components derived from PLS-c. The authors should also cite justification for a bootstrap ratio cutoff of 2.3 for defining stability.

      In PLS-c analyses, multiple comparisons correction across latent components and bootstrap ratio (BSR) thresholding at 2.3 are commonly adopted practices.

      - Correction for multiple comparisons in PLS-c: PLS-c performs singular decomposition of the data into latent components equal in number to the variables included in the behavior design matrix (7-9 in our study, depending on the inclusion of Verbal outcome as an input variable). Each latent component’s statistical significance is assessed through permutation testing, generating a null distribution for its singular value (Krishnan et al., 2011). Given the multiple tests (one permutation test per latent component), Type I error inflation must be addressed. Recent PLS-c studies commonly applied Bonferroni correction (default procedure in the myPLS toolbox, used by Zoeller et al., 2017, and Delavari et al, 2021), though FDR correction has also been used (Lombardo et al, 2018).

      - Stability threshold for bootstrap and replacement: Within each latent component, saliences’ stability (brain/behavior parameter contributions to each latent component) are evaluated using bootstrapping (Krishnan et al., 2011). The bootstrap ratio (BSR) of each parameter, calculated as the saliency divided by its bootstrap-derived standard error, functions analogously to a z-score under normality assumptions. The BSR can then be used to assess the stability of the saliency (i.e., how stable is its contribution to the latent component). BSR thresholds in the literature typically range from 1.96 to 3.0. Krishnan et al (2011) state that when BSR are “larger than 2 the corresponding saliences are considered significantly stable”. Delavari et al (2021) and our study used a 2.3 thresholding, corresponding to a 99.0% bootstrap confidence interval not crossing the zero line – roughly equivalent to a two-tailed p<.001. Lombardo et al (2018) used a looser threshold of 1.96, corresponding to a 95% confidence interval not crossing the zero line (~two-tailed p<.05), while Zöller et al (2017) used a more stringent 3.0 thresholding (~p<.001, or 99.9% confidence interval not crossing the zero line).

      Thus, our application of Bonferroni correction for multiple comparisons and our 2.3 BSR threshold aligns with established conventions.

      We added following lines in the manuscript:

      Page 25 lines 784-785: “Bonferroni correction was applied to account for multiple comparisons across the 9 tested latent components in the PLS-c, yielding an adjusted alpha of .006 (Zoeller et al, 2017; Delavari et al, 2021).”

      Page 25 lines 789-792: “BSR are analogous to Z-scores and can be used to assess the stability of the saliency. We considered BSR > 2.3 as stable, corresponding to a 99.0% bootstrap confidence interval not crossing zero – roughly equivalent to a two-tailed p<.001 (Delavari et al., 2021; Krishnan et al., 2011).”

      References:

      Delavari F, Sandini C, Zöller D, Mancini V, Bortolin K, Schneider M, Van De Ville D, Eliez S. Dysmaturation Observed as Altered Hippocampal Functional Connectivity at Rest Is Associated With the Emergence of Positive Psychotic Symptoms in Patients With 22q11 Deletion Syndrome. Biol Psychiatry. 2021 Jul 1;90(1):58-68. doi: 10.1016/j.biopsych.2020.12.033. Epub 2021 Jan 18. PMID: 33771350.

      Lombardo, M.V., Pramparo, T., Gazestani, V. et al. Large-scale associations between the leukocyte transcriptome and BOLD responses to speech differ in autism early language outcome subtypes. Nat Neurosci 21, 1680–1688 (2018). https://doi.org/10.1038/s41593-018-0281-3

      Daniela Zöller, Marie Schaer, Elisa Scariati, Maria Carmela Padula, Stephan Eliez, Dimitri Van De Ville. Disentangling resting-state BOLD variability and PCC functional connectivity in 22q11.2 deletion syndrome. NeuroImage, Volume 149, 2017, Pages 85-97, ISSN 1053-8119, https://doi.org/10.1016/j.neuroimage.2017.01.064

      Anjali Krishnan, Lynne J. Williams, Anthony Randal McIntosh, Hervé Abdi, Partial Least Squares (PLS) methods for neuroimaging: A tutorial and review, NeuroImage, Volume 56, Issue 2, 2011, Pages 455-475, ISSN 1053-8119, https://doi.org/10.1016/j.neuroimage.2010.07.034

      (7) I have a few minor grammar/formatting recommendations for the authors as well:

      (a) Should the Geneva Autism Cohort be capitalized? At present, it is not.

      We agree with the reviewer’s suggestion, and we capitalized the Geneva Autism Cohort in the main text (page 18, line 571)

      (b) Page 24, line 750. Do the authors mean that the data was re-referenced to average?

      The preprocessed data is not average-referenced (see section Data pre-processing). Therefore, both for neural entrainment computation and ERPs, the data were average-referenced.

      (c) It would be nice to have a figure of the actual ERP for each condition and age group.

      We agree that PLS-c can be difficult to interpret without the raw actual ERPs on which it was modelled. We direct the reviewer to supplementary figure S6 at page 59, which displays the raw ERPs for each condition (part-word, word, and their subtraction) per age group. Supplementary figures S7-8 at pages 60-61 further illustrate topographical ERPs for each group (high and low likelihood for autism). We deemed these figures too extensive for the main text. Instead, the most relevant ERP topographies are presented in Figures 4-6 to facilitate PLS-c interpretation.

    1. eLife Assessment

      In this valuable study, the authors identified a rare population of Nestin-expressing cells within the external granule layer of the early postnatal mouse cerebellum. They demonstrated that these cells are distinct from Sox2+ progenitors and can give rise to medulloblastoma. Collectively, the findings provide convincing evidence that this unique Nestin+ population is susceptible to oncogenic transformation and may underlie the preferential emergence of Sonic hedgehog-driven medulloblastomas.

    2. Reviewer #1 (Public review):

      Summary:

      GCPs, which drive postnatal cerebellar growth and can give rise to SHH-MB, are not uniform. The authors show that GCPs include a rare Nestin-expressing subpopulation with distinct molecular features. This subpopulation is spatially restricted, enriched for stem cell-like properties, and shows a high competency for tumor formation comparable to larger GCP pools, with tumors preferentially arising in the posterior-lateral cerebellum. Overall, the findings indicate that SHH-MB might originate preferentially from this small, tumor-competent Nestin-expressing GCP subset.

      Strengths:

      (1) The authors use a breadth of approaches from histology, mouse genetics, and single-cell RNA sequencing.

      (2) Throughout, this paper uses very elegant genetic approaches, such as the double Nes-FlpoER; Atoh1-FSF-Cre; LSL-Smo-M2, to generate tumors only from Atoh1+; Nes+ double-positive cells. This intersectional genetic experiment makes for a very clear answer.

      (3) The findings reported in this manuscript are valuable since they reveal a novel GCP subpopulation defined by spatial and molecular identity. Some of their experiments suggest that these cells could represent the main cell-of-origin of SHH MB. The experiments are carefully performed, and the evidence is convincing.

      Weaknesses or elements that could be improved:

      (1) A transgenic Nestin-CFP mouse is used in this study. However, it is not clear whether CFP accurately reflects the Nestin protein. Figure 1: After the promoter is turned off, these cells might remain positive for CFP for longer than they are positive for Nestin, due to CFP protein stability. Is the Nestin protein present in these cells? Nestin double immunofluorescence with CFP and Sox2 and Barhl1 could be performed to address this. Related to this comment, it is also important to note that this is a rat promoter transgene. So the transgene might not reflect exactly the endogenous Nestin expression.

      (2) Could the posterior restriction of Nestin-CFP be due to the timing (P1) at which the authors looked? In other words, if they look earlier, would the authors see Nestin-CFP cells more anterior?

      (3) Since only one medulloblastoma mouse model (Smo-M2) is used to conclude that "the Nes-expressing GCP population in the normal cerebellum is transcriptionally closer to SHH MB tumor cells than the remainder of the GCPs", the findings might not apply to other SHH-MB models. This should be mentioned.

    3. Reviewer #2 (Public review):

      Summary:

      In this manuscript, the authors studied transgenic reporter mice to profile Nestin expression in the postnatal mouse cerebellum. They discovered a small population of Nestin+; Atoh1+ granule neuron precursors (GNPs) in the external granule cell layer (EGL). Using immunostaining, qPCR, and RNA-sequencing, the authors showed that these Nestin+ cells are not identical to Sox2+ cells (e.g., the majority of Nestin+ cells are Sox2-). Using various mouse genetic strategies, including an elegant intersectional strategy that specifically targets Nestin+; Sox2+ cells, the authors showed that Nestin+ cells are capable of initiating Sonic hedgehog (SHH) medulloblastoma when they express the SmoM2 allele that drives constitutively active SHH signaling. Lastly, the authors profiled the transcriptomes of these cells and showed that they display enriched stem cell genes and are closer to the transcriptomes of GNP-like cells in medulloblastoma compared to Nestin- GNPs in the developing cerebellum.

      Strengths:

      (1) The comprehensive mouse genetics experiments, in combination with immunostaining, lineage tracing, and RNA-seq studies, provided compelling evidence that rare Nestin+ cells are present in the EGL, predominantly at the posterior lateral cerebellum in early postnatal mice.

      (2) The intersectional genetics experiment unequivocally show that Nestin+; Atoh1+ cells can be oncogenically transformed by SmoM2, leading to SHH medulloblastoma.

      (3) The more stem cell-like transcriptomic features of the Nestin+ GNPs compared to Nestin- GNPs provide support for the heterogeneity of this transient progenitor cell population, with implications for development, congenital diseases, and tumors from the cerebellum.

      Weaknesses:

      Main comments:

      My main concern relates to whether these Nestin+; Atoh1+ cells are restrictively localized in the EGL. Both the title "A Rare Nestin-Expressing Granule Cell Precursor Subpopulation Underlies SHH Medulloblastoma Formation" and what the authors described throughout the manuscript propose that Nestin+; Atoh1+ cells in the EGL are the cell-of-origin of SHH medulloblastoma. To definitively conclude this, the authors need to comprehensively analyze all regions of the developing cerebellum.

      Most importantly, are Nestin+; Atoh1+ cells present in the rhombic lip? Are there any rhombic lip cells genetically labeled in their intersectional mouse mutants (e.g., the Atoh1Frt-Cre/+; Nes-FlpoER; R26LSL-SsmoM2/+ mice)?

      If Nestin+; Atoh1+ cells are present at non-EGL regions in the developing cerebellum, the authors would have to reconsider many of their conclusions and also the title of this paper.

      Additional comments:

      (1) To investigate Nestin expression, the authors used Nes-CFP transgenic mice expressing CFP from promoter/enhancer sequences from the rat Nes gene (Encinas et al., 2006). Given that Nestin expression is of central importance for this study, it is important to validate that these reporter mice faithfully report Nestin protein expression (e.g., by co-labeling CFP with Nestin antibody and systemically comparing signals throughout the cerebellum, ideally in several developmental stages).

      (2) The authors mostly presented immunostaining data of the cerebellum from P1 mice. It is important to systematically profile the appearance and disappearance of these Nestin+, Atoh1+ cells in mouse cerebellum across developmental stages (e.g., embryonic, early, and late postnatal stages).

      (3) How different is the proliferative ability of the Nestin+ versus Nestin- GNPs at various developmental stages? Also, the difference between EdU+; Barhl1+; Nestin+ and EdU+; Barhl1+; Nestin- cells is quite small despite statistical difference (Figure 1N). Do the authors think this very small EdU incorporation difference can translate into a biological difference (in developmental and/or disease context)?

      (4) Lines 145-147: "Compared to double-negative cells, Atoh1 and Nes were significantly higher in the double-positive fraction, supporting the identity of the cells as a previously unrecognized rare population of GCPs at P1 that expresses both the GCP marker Atoh1 and ventricular zone marker Nes." Nestin is not a ventricular zone marker. This should be rephrased.

      (5) In Figure 3, the authors showed mouse survival data and concluded that Nes-driven and Atho1-driven SHH medulloblastoma models show similar tumor penetrance. This is not an entirely accurate description of their data. The Nes-SmoM2 mice displayed significantly longer survival compared to the Atoh1-SmoM2 mice (Figure 3B). This conclusion needs to be revised.

      (6) In Figure 5, the authors showed that genes enriched in cluster 10 included Sox2, Nes, Wls, and Wnt1, while Neurod1 and Rbfox3 were preferentially expressed in the other GCP clusters. They conclude that cluster 10 represents a less differentiated, more stem-like GCP state, potentially positioned upstream in the lineage hierarchy. While these few markers are useful, it is more informative to formally support this conclusion by comparing the stem cell transcriptomic signature (using a larger gene list) between cluster 10 and other GCPs.

      (7) In Figure 5, the authors performed gene ontology analysis and showed that cluster 10 is enriched for biological processes linked to WNT signaling and proposed that this molecular profile supports their identity as a transient, developmentally plastic population within the GCP lineage related to the rhombic lip. The authors are recommended to use an orthogonal approach (i.e., immunostaining to compare nuclear localization of beta-Catenin) to validate their transcriptome-based finding.

    1. eLife Assessment

      This manuscript presents important new findings showing that the transcription factor and regulator of lipid and glucose metabolism PPARγ is methylated by the enzyme SETD6. The data convincingly demonstrate that methylation of PPARγ by SETD6 regulates its function in controlling transcription of lipid storage and metabolism genes and thereby modulates lipid accumulation in liver cells. This work uncovers a new role for post-translational modulation by lysine methylation in controlling transcription factor activity and lipid metabolism in a relevant physiological context.

    2. Reviewer #1 (Public review):

      Summary:

      In this manuscript from the Levy lab, the authors investigate whether SETD6 regulates hepatic lipid accumulation through direct methylation of PPARγ. They show that SETD6 binds and mono-methylates PPARγ at K170 and provide evidence that this modification enhances PPARγ occupancy at target promoters, promotes expression of lipid metabolism genes, as well as facilitates lipid droplet accumulation in HepG2 cells. The authors also find a positive feedback loop or circuit in which PPARγ activates SETD6 transcription in a methylation-dependent manner, thereby reinforcing this lipogenic program. Overall, the work presents a novel SETD6-PPARγ regulatory axis linking lysine methylation to transcriptional control of lipid storage genes, with possible relevance to NAFLD-associated biology.

      In all, I find this to be an important paper that describes and advances a new regulatory pathway that has significance to human health and disease. It would also be of interest to a broad audience. That said, there are also some concerns that the authors should address, as outlined below.

      Major concerns (pertains to rigor - highest priority)

      (1) Overall, the work presented is of high quality and the data nicely support the conclusions; however, a few panels should be strengthened that have missing controls or information:<br /> a. The co-IP panel in Fig. 1B lacks a lane where HA SETD6 is expressed without PPARγ. This control is needed to verify that the SEDT6-HA signal depends on PPARγ.<br /> b. In Fig. 1C, the authors should show that the co-IP works in both directions (include IP for PPARγ/blot for SETD6). I am a bit confused also over the labeling with IP on the left and on top of the panel next to the beads label. More importantly, the data would be stronger if the authors take advantage of a deletion line to validate the co-IP is specific to the presence of both.<br /> c. The same IP labeling issue exists for Fig 3B (label is on the same and on top).<br /> d. Antibody information (e.g., where the pan-methyl Ab comes from and at what dilutions they are used at) is missing.

      Nice to have experiments (medium priority - strongly consider)

      (2) A missing gap is how K170me1 contributes to DNA binding and gene transcription. One possibility is that methylation enhances the DNA binding activity of PPARγ. Given the authors have all of the reagents, it would be possible to perform a gel shift assay (or other approach) with and without SETD6-mediaetd methylation. Is DNA binding affected/enhanced?

      (3) Along these lines, I wonder if there is another possibility: could SETD6-mediated methylation of PPARγ drive SETD6-PPARγ interaction? In other words, in the K170R, is SETD6 still even associated with PPARγ, and this interaction is required for promoter recruitment? Alternatively, would a catalytic dead version of SETD6 fail to associate with PPARγ? Currently, no experiments test the impact of an unmethylatable version of PPARγ or catalytic dead version of SETD6 on SETD6-PPARγ interaction or SETD6 recruitment to promoters.

      Minor concerns (text and figure display)

      (4) The text has multiple typos and grammatical errors.

      Comments on revised version.

      Great job on addressing the comments. It is a nice study.

    3. Reviewer #2 (Public review):

      Summary:

      In this work, the authors investigated the regulation of the transcription factor PPARγ by the post-translational modification lysine methylation The data demonstrate that the lysine methyltransferase SETD6 targets PPARγ for methylation using biochemical and cell-based assays. Methylation of PPARγ occurs in its DNA binding domain, and the authors demonstrate that loss of methylation limits PPARγ chromatin binding, particularly to lipid storage and metabolism genes promoters. As a physiological output, the authors demonstrate that deletion of SETD6 and loss of PPARγ methylation also disrupt lipid droplet accumulation in hepatocytes. In addition, the authors uncover a positive feedback loop in which SETD6 methylation of PPARγ also regulates its binding to the SETD6 promoter and expression of the gene.

      Strengths:

      One of the key strengths of this manuscript is the novelty of the findings in terms of identifying a new mode of regulation of PPARγ that modulates its chromatin association in cells and thereby regulating lipid metabolism genes. The authors nicely combine biochemical studies of SETD6 activity with cell-based assays investigating PPARγ and SETD6 function in regulating lipid storage. Data supporting this conclusion is largely convincing and frequently, multiple assays are used to provide sufficient support to the conclusions. This work therefore expands regulatory modes of PPARγ and identifies a new target for SETD6, an enzyme that targets a number of other transcription factors. Furthermore, the regulatory loop that controls SETD6 expression via PPARγ methylation is likely important for understanding SETD6 function in different cell types that have high levels of lipid accumulation or regulation. The gene expression and lipid accumulation assays are useful for testing the physiological outcome of loss of SETD6 activity or PPARγ methylation directly. In the revised manuscript, the authors have added useful structural modeling to better define potential roles of methylation of PPARγ in regulating its function, particularly relative to DNA binding, and to better define the physical interaction between PPARγ and SETD6.

      Weaknesses:

      The revised manuscript substantially improved on the presentation of the data and broadened the discussion to provide more context to both the role of SETD6 and to elaborate on potential mechanisms by which methylation impacts PPARγ function and under what physiological conditions this interaction and regulation is important. This improves and strengthens the manuscript and its impact overall.

      Comments on revised version.

      The authors addressed all of my major concerns following this round of review and I do not have additional recommendations. The presentation of the manuscript including text and figures is improved compared to the previous version. I have updated my public review to reflect these changes.

    4. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      In this manuscript from the Levy lab, the authors investigate whether SETD6 regulates hepatic lipid accumulation through direct methylation of PPARγ. They show that SETD6 binds and monomethylates PPARγ at K170, and provide evidence that this modification enhances PPARγ occupancy at target promoters, promotes expression of lipid metabolism genes, as well as facilitates lipid droplet accumulation in HepG2 cells. The authors also find a positive feedback loop or circuit in which PPARγ activates SETD6 transcription in a methylation-dependent manner, thereby reinforcing this lipogenic program. Overall, the work presents a novel SETD6PPARγ regulatory axis linking lysine methylation to transcriptional control of lipid storage genes, with possible relevance to NAFLD-associated biology.

      In all, I find this to be an important paper that describes and advances a new regulatory pathway that has significance to human health and disease. It would also be of interest to a broad audience. That said, there are also some concerns that the authors should address, as outlined below.

      We are grateful to the reviewer for the positive feedback and appreciation of our work.

      Major concerns (pertains to rigor - highest priority)

      (1) Overall, the work presented is of high quality, and the data nicely support the conclusions; however, a few panels should be strengthened that have missing controls or information:

      (a) The co-IP panel in Figure 1B lacks a lane where HA SETD6 is expressed without PPARγ. This control is needed to verify that the SEDT6-HA signal depends on PPARγ.

      We thank the reviewer for this valuable suggestion. The overexpression co-immunoprecipitation experiment referred to by the reviewer has been moved to Supplementary Figure S1 in the revised manuscript. In this experiment, immunoprecipitation was performed using an anti-FLAG antibody to pull down FLAG-tagged PPARγ. In the absence of FLAG-PPARγ, the anti-FLAG immunoprecipitation does not recover a bait protein, and therefore HA-SETD6 is not expected to be specifically immunoprecipitated. Thus, an HA-SETD6-only condition would primarily serve as a negative control for the anti-FLAG pull-down rather than provide additional information regarding the specificity of the interaction.

      Importantly, in the revised manuscript we have substantially strengthened the evidence supporting the SETD6–PPARγ interaction by adding two independent complementary experiments. First, we included a reciprocal endogenous co-immunoprecipitation (new Figure 2B), demonstrating that endogenous PPARγ co-immunoprecipitates with endogenous SETD6. Second, we added an independent proximity ligation assay (PLA) (new Figure 2D), which further confirms the interaction between SETD6 and PPARγ in cells. Together with the in vitro binding assay presented in Figure 2A, these orthogonal approaches provide compelling evidence for the specificity of the SETD6–PPARγ interaction. Therefore, we believe that the requested HA-SETD6-only control would not provide additional mechanistic insight beyond the comprehensive validation now included in the revised manuscript.

      (b) In Figure 1C, the authors should show that the co-IP works in both directions (include IP for PPARγ/blot for SETD6). I am a bit confused also over the labeling with IP on the left and on top of the panel next to the beads label. More importantly, the data would be stronger if the authors took advantage of a deletion line to validate that the co-IP is specific to the presence of both.

      We thank the reviewer for this helpful suggestion. We have revised the manuscript to strengthen the evidence supporting the endogenous interaction between SETD6 and PPARγ. Specifically, we now include a reciprocal endogenous co-immunoprecipitation (new Figure 2B), demonstrating that endogenous SETD6 co-immunoprecipitates with endogenous PPARγ and, conversely, that endogenous PPARγ co-immunoprecipitates with endogenous SETD6. These reciprocal experiments independently validate the specificity of the interaction.

      In addition, we have revised the figure layout and labeling to more clearly distinguish the immunoprecipitating antibody from the bead control, thereby addressing the reviewer's concern regarding the presentation of the co-immunoprecipitation data.

      Although we did not perform the co-immunoprecipitation in a depletion/knockout background, we believe that the combination of reciprocal endogenous co-immunoprecipitation (New Figure 2B), the independent proximity ligation assay (New Figure 2D), and the direct in vitro binding assay (Figure 2A) provides multiple orthogonal lines of evidence supporting a specific interaction between SETD6 and PPARγ.

      (c) The same IP labeling issue exists for Figure 3B (label is on the same and on top).

      We have revised the labeling in Figure 3B to clearly distinguish the immunoprecipitating antibody from the bead control, thereby improving the clarity of the figure.

      (d) Antibody information (e.g., where the pan-methyl Ab comes from and at what dilutions they are used at) is missing.

      We thank the reviewer for pointing this out. We have now added the missing information regarding the pan-methyl antibody to the Materials and Methods section, including the supplier, catalogue number, and experimental conditions used. Specifically, the pan-methyl antibody used in this study was purchased from Abcam (ab23366) and was used at a 1:500 dilution for western blot analysis and 2 μg per reaction for immunoprecipitation experiments.

      Nice to have experiments (medium priority - strongly consider)

      (2) A missing gap is how K170me1 contributes to DNA binding and gene transcription. One possibility is that methylation enhances the DNA-binding activity of PPARγ. Given that the authors have all of the reagents, it would be possible to perform a gel shift assay (or other approach) with and without SETD6-mediated methylation. Is DNA binding affected/enhanced?

      We thank the reviewer for raising this important point. To investigate whether K170 methylation could directly affect PPARγ binding to DNA, we performed structural modeling based on the available co-crystal structure of PPARγ bound to DNA (PDB: 3DZU). As shown in the new Figure 6F, K170 is positioned near the DNA-binding region; however, the modeled K170me1 side chain is predicted to face away from the DNA interface and does not appear to sterically interfere with the PPARγ–DNA interaction. In addition, modeling of multiple K170me1 rotamers did not suggest any major disruption of the DNA-bound conformation.

      These observations suggest that K170 methylation is unlikely to directly alter the intrinsic DNA-binding affinity of PPARγ. In contrast, our ChIP-qPCR experiments demonstrate that K170 methylation positively regulates PPARγ occupancy at target promoters in cells. Together, these findings support a model in which K170 methylation promotes PPARγ chromatin association and transcriptional activity through mechanisms other than direct modulation of DNA binding, such as altered cofactor recruitment or protein–protein interactions.

      We agree with the reviewer that future biochemical approaches, including EMSA/gel shift assays or quantitative DNA-binding measurements, will be valuable to directly determine whether K170 methylation affects the intrinsic DNA-binding affinity of PPARγ. We have incorporated this new structural analysis and the corresponding discussion into the revised manuscript.

      (3) Along these lines, I wonder if there is another possibility: could SETD6-mediated methylation of PPARγ drive SETD6-PPARγ interaction? In other words, in the K170R, is SETD6 still even associated with PPARγ, and this interaction is required for promoter recruitment? Alternatively, would a catalytic dead version of SETD6 fail to associate with PPARγ? Currently, no experiments test the impact of an unmethylatable version of PPARγ or a catalytic dead version of SETD6 on SETD6-PPARγ interaction or SETD6 recruitment to promoters.

      We thank the reviewer for this insightful suggestion. To address whether SETD6 catalytic activity is required for its association with PPARγ, we performed an additional PLA experiment comparing SETD6 WT and the catalytic mutant SETD6 Y285A. As shown in the revised New Figure 2D, both SETD6 WT and SETD6 Y285A showed comparable proximity to PPARγ in cells, indicating that SETD6 catalytic activity is not required for the physical association between SETD6 and PPARγ.

      These findings support a model in which SETD6 first recognizes and binds PPARγ independently of its catalytic activity. Subsequent methylation of PPARγ at K170 is therefore likely to regulate the downstream functional consequences of this interaction, including enhanced chromatin occupancy and transcriptional activation, rather than the initial SETD6–PPARγ association itself.

      In addition, we generated using AlphaFold a structural model of the SETD6–PPARγ complex as a supportive visualization (New figure S2). Given the limited confidence of the prediction, we interpret this model cautiously and include it in the Supplementary Information rather than the main figures. We agree with the reviewer that future studies examining SETD6 recruitment to PPARγ target promoters and the effect of the PPARγ K170R mutant on SETD6–PPARγ association will further refine the molecular mechanism.

      Minor concerns (text and figure display)

      (4) The text has multiple typos and grammatical errors, and there are some issues with the figure display.

      We thank the reviewer for this comment. We carefully revised the manuscript to correct typographical and grammatical errors throughout the text and also addressed the figure display issues noted by the reviewer.

      Reviewer #2 (Public review):

      Summary:

      In this work, the authors investigated the regulation of the transcription factor PPARγ by the post-translational modification lysine methylation. The data demonstrate that the lysine methyltransferase SETD6 targets PPARγ for methylation using biochemical and cell-based assays. Methylation of PPARγ occurs in its DNA binding domain, and the authors demonstrate that loss of methylation limits PPARγ chromatin binding, particularly to lipid storage and metabolism gene promoters. As a physiological output, the authors demonstrate that deletion of SETD6 and loss of PPARγ methylation also disrupt lipid droplet accumulation in hepatocytes. In addition, the authors uncover a positive feedback loop in which SETD6 methylation of PPARγ also regulates its binding to the SETD6 promoter and expression of the gene.

      Strengths:

      One of the key strengths of this manuscript is the novelty of the findings in terms of identifying a new mode of regulation of PPARγ that modulates its chromatin association in cells and thereby regulates lipid metabolism genes. The authors nicely combine biochemical studies of SETD6 activity with cell-based assays investigating PPARγ and SETD6 function in regulating lipid storage. Data supporting this conclusion is largely convincing, and frequently, multiple assays are used to provide sufficient support to the conclusions. This work therefore expands regulatory modes of PPARγ and identifies a new target for SETD6, an enzyme that targets a number of other transcription factors. Furthermore, the regulatory loop that controls SETD6 expression via PPARγ methylation is likely important for understanding SETD6 function in different cell types that have high levels of lipid accumulation or regulation. The gene expression and lipid accumulation assays are useful for testing the physiological outcome of loss of SETD6 activity or PPARγ methylation directly.

      We thank the reviewer for his/her positive feedback on the manuscript.

      Weaknesses:

      The data presented in the manuscript are largely convincing in support of the authors' conclusions; however, there are some errors in the presentation of the figures and some issues in the text that would benefit from editing. Furthermore, there are some important questions not fully addressed in the results or discussion. 

      It would be great if the authors could speculate more on the diverse roles of SETD6 in methylated transcription factors and/or provide more context regarding the conditions that are likely to support methylation of PPARγ by SETD6. 

      We thank the reviewer for this important suggestion. In the revised Discussion, we expanded the manuscript to better place our findings within the broader context of SETD6-mediated regulation of transcription factors. Previous studies from our group and others demonstrated that SETD6 methylates multiple chromatin-associated transcriptional regulators, including RelA, E2F1, TWIST1, and BRD4, thereby modulating transcriptional selectivity, chromatin occupancy, and cofactor recruitment. We now discuss the possibility that SETD6 functions as a context-dependent signalling integrator that selectively regulates transcription factor activity through lysine methylation under distinct physiological conditions.

      In addition, we expanded the Discussion regarding potential cellular contexts that may favour PPARγ methylation by SETD6. Because PPARγ activity is strongly induced during lipid overload and fatty acid exposure, conditions associated with steatosis and metabolic stress may enhance the functional importance of SETD6-dependent methylation. We also discuss the possibility that chromatin accessibility, ligand-dependent activation of PPARγ, and metabolic signaling pathways may collectively influence the formation and stability of the SETD6–PPARγ complex.

      Also, while a potential cross-talk between methylation and phosphorylation is described in the discussion, it would be great to provide more structural insight into how this might regulate DNA binding of PPARγ and/or discuss whether there are other possibilities given the location of the target lysine in the DNA binding domain.

      We thank the reviewer for this valuable suggestion. To provide additional structural insight into the potential interplay between methylation and phosphorylation within the PPARγ DNA-binding domain, we performed structural modeling based on the published PPARγ–DNA co-crystal structure (PDB: 3DZU). As shown in the new Figure S6, the modeled K170me1 side chain is predicted to face away from the DNA interface and does not introduce steric clashes with DNA, suggesting that K170 methylation is unlikely to directly alter the DNA-binding interface through steric effects. In contrast, phosphorylation of the neighboring residue T166 is predicted to introduce multiple intramolecular steric clashes within the DNA-binding domain. These structural changes could influence the local conformation or dynamics of the DNA-binding domain and thereby indirectly modulate PPARγ DNA binding or transcriptional activity.

      In addition, as suggested by the reviewer, we expanded the Discussion to consider alternative mechanisms by which K170 methylation may regulate PPARγ function. While our ChIP-qPCR experiments demonstrate that K170 methylation positively regulates PPARγ chromatin occupancy at target promoters, the structural modeling suggests that this effect is unlikely to arise from direct steric modulation of the DNA interface. Instead, K170 methylation may influence chromatin occupancy by regulating protein–protein interactions, cofactor recruitment, local conformational dynamics, or other chromatin-associated mechanisms. We have incorporated these new structural analyses and the expanded discussion into the revised manuscript.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) The KEGG panels are too low resolution to read.

      We thank the reviewer for this comment. We improved the resolution of the KEGG pathway enrichment panels and, as suggested by the reviewer (see below), we separated the upregulated and downregulated gene sets to improve clarity and readability. These changes are now reflected in the revised new Figures 5C, 5D, 6B and 6C.

      (2) Figure 3 panel D is hard to visualize; a better version or repeat experiment is needed.

      We thank the reviewer for this comment. To address this concern, we replaced the original Figure 3D with a new independent experiment that more clearly demonstrates the methylation of endogenous PPARγ by SETD6. We believe that the new data provide substantially stronger evidence and improve the clarity of the revised manuscript.

      (3) There is a typo in Figure 2A (line present in "S" of Signal).

      We thank the reviewer for pointing out this typo. The error in Figure 2A has now been corrected in the revised manuscript.

      Reviewer #2 (Recommendations for the authors):

      Overall, the experiments and analyses presented are sufficient to support the conclusions and interpretations of the work. However, there are some issues of presentation and writing that are worth addressing. These comments are listed below:

      (1) My only substantial recommendation is to improve the discussion to provide more context to the findings in terms of both the larger role of SETD6 in methylating transcription factors (some of whom also regulate its expression) and the potential modes through which PPARγ DNA binding activity could be regulated by methylation.

      We thank the reviewer for this insightful suggestion. In the revised manuscript, we substantially expanded the Discussion to better place our findings within the broader context of SETD6mediated regulation of transcription factors. We now discuss previous studies demonstrating that SETD6 methylates multiple chromatin-associated transcriptional regulators, including RelA, E2F1, TWIST1, and BRD4, thereby modulating chromatin occupancy, cofactor recruitment, and transcriptional selectivity. We further propose that SETD6 functions as a context-dependent signaling regulator that integrates distinct cellular pathways through the selective lysine methylation of transcription factors.

      In addition, we expanded the Discussion regarding the potential mechanisms by which PPARγ K170 methylation regulates transcriptional activity. We incorporated new structural modeling based on the published PPARγ–DNA co-crystal structure (PDB: 3DZU), which suggests that K170 methylation is unlikely to directly alter the DNA-binding interface through steric effects, whereas phosphorylation of the neighboring residue T166 may induce intramolecular steric clashes within the DNA-binding domain. We also expanded the Discussion to consider alternative mechanisms by which K170 methylation may regulate PPARγ function, including modulation of chromatin occupancy, local conformational dynamics, protein–protein interactions, and recruitment of transcriptional cofactors or chromatin-associated proteins. Finally, we discuss that future biochemical studies will be important to determine whether K170 methylation also influences the intrinsic DNA-binding affinity of PPARγ.

      Can the authors incorporate any other published structural data to speculate on the role of methylation or describe more about how it is expected that the methylation-phosphorylation crosstalk modulates DNA binding? This type of discussion would better highlight the potential importance of this new modification on PPARγ.

      We thank the reviewer for this important suggestion. In the revised manuscript, we expanded the Discussion and incorporated additional structural analyses (new Figure 6F and new figure S6) based on the published PPARγ–DNA co-crystal structure (PDB: 3DZU). K170 is positioned within the DNA-binding domain, between the two zinc-finger motifs that mediate DNA recognition and stabilization on PPRE-containing DNA. Our structural modeling predicts that the K170me1 side chain is oriented away from the DNA interface and does not introduce steric clashes with DNA, suggesting that methylation is unlikely to directly alter the DNA-binding interface through steric effects. Nevertheless, lysine methylation can influence protein surface properties, protein– protein interactions, and recognition by regulatory binding partners, raising the possibility that K170 methylation modulates PPARγ chromatin occupancy or promoter selectivity through indirect mechanisms.

      In contrast, structural modeling predicts that phosphorylation of the neighboring residue T166 introduces multiple intramolecular steric clashes within the DNA-binding domain. These clashes could alter the local conformation or dynamics of the DNA-binding domain and thereby indirectly influence PPARγ–DNA interactions and transcriptional activity. Together, these observations suggest that K170 methylation and T166 phosphorylation may represent a regulatory crosstalk that fine-tunes PPARγ function through distinct structural mechanisms.

      We further expanded the Discussion to consider additional, non-mutually exclusive mechanisms by which K170 methylation may regulate PPARγ function, including modulation of chromatin occupancy, protein–protein interactions, cofactor recruitment, and stabilization of transcriptional complexes at target genes. While these possibilities require further mechanistic investigation, we agree with the reviewer that these structural considerations highlight the potential regulatory importance of this newly identified PPARγ modification.

      (2) In the gene expression experiments presented in Figure 5, it would be useful if the GO terms were described as enriched in either the up- or down-regulated gene sets. From the way it is presented, it is not clear if specific categories of genes are found enriched in those upregulated or downregulated upon KO of SETD6. This is also true for the experiments presented in Figure 6 regarding the PPARγ mutant.

      We thank the reviewer for this important suggestion. In the revised manuscript, we separated the pathway enrichment analyses into upregulated and downregulated gene sets for both the SETD6 knockout RNA-sequencing experiments (Figure 5) and the PPARγ WT versus K170R mutant analysis (Figure 6). This revision provides improved clarity regarding which biological pathways are positively or negatively associated with SETD6 depletion or disruption of PPARγ K170 methylation. The updated KEGG enrichment analyses are now presented in the revised new Figures 5C, 5D, 6B and 6C.

      In addition, if there is a significant overlap in genes misregulated in both mutants, this could be shown in the figure.

      We thank the reviewer for this suggestion. We compared the differentially expressed genes identified in the SETD6 KO cells and the PPARγ K170R mutant cells to evaluate the extent of overlap between the two datasets. However, we did not observe a substantial or statistically significant overlap in misregulated genes under the thresholds used in our analysis. Therefore, we decided not to include this comparison in the revised figure. Nevertheless, both datasets consistently showed enrichment for pathways associated with lipid metabolism and PPAR signaling, supporting a functional connection between SETD6 and PPARγ-mediated transcriptional regulation.

      (3) There are some typos and grammatical errors throughout the work, and it should be carefully edited. One example is the following heading: PPARγ K170 methylation by SETD6 regulates mediates lipid droplets formation.

      We thank the reviewer for this comment. The manuscript was carefully revised to correct typographical and grammatical errors throughout the text. In particular, the heading mentioned by the reviewer was corrected in the revised manuscript.

      Minor errors in the figures:

      (1) Formatting issue in the labeling for the x-axis of Figure 1E.

      We thank the reviewer for pointing out this formatting issue. The labeling of the x-axis in Figure 1E has been corrected in the revised manuscript.

      (2) Figure 2B - HA-SETD6 should be labeled as minus for the first lane.

      We thank the reviewer for pointing this out. We corrected the labeling in Figure 2B (now Figure S1), and the first lane is now properly indicated as negative for HA-SETD6.

      (3) Figure 2D - Labeling needs improvement. Is this FLAG-PPARγ? "NC" was not defined in the legend. If negative control, what type?

      We thank the reviewer for this comment. We improved the labeling and figure legend of Figure 2D for clarity. Specifically, we now clearly indicate that the experiment was performed using Flag-PPARγ, and we defined “NC” in the legend as the negative control condition. In addition, during the revision process we noticed that the original PLA experiment was performed in HeLa cells rather than HepG2 cells, as previously indicated. This has now been corrected throughout the revised manuscript.

      (4) Figure 3 - The CRSIPR control should be described somewhere in the legend or the methods.

      We thank the reviewer for this comment. We clarified the description of the CRISPR control cells in the Materials and Methods section. Specifically, we now explicitly state that the CRISPR control (CT) cells were generated using the empty lentiCRISPR vector without SETD6-targeting sgRNAs.

      (5) Figure 5B and 6A - The legend is not labeled nor defined in the text- fold-change, log2 fold change, z score?

      We thank the reviewer for pointing this out. We revised the figure legends for Figures 5B and 6A to explicitly define the heatmap scale and normalization method. Specifically, we now indicate that the heatmaps represent normalized gene expression values displayed as Z-scores.

    1. eLife Assessment

      This study presents valuable data suggesting that ATP-induced modulation of alveolar macrophage (AM) functions is associated with NLRP3 inflammasome activation and enhanced phagocytic capacity. While the in vivo and in vitro data reveal an interesting phenotype, the evidence provided is incomplete and does not fully support the paper's conclusions. Additional investigations would be of value in complementing the data and strengthening the interpretation of the results. This study should be of interest to immunologists and the mucosal immunity community.

    2. Reviewer #1 (Public review):

      Summary:

      Alveolar macrophages (AMs) are key sentinel cells in the lungs, representing the first line of defense against infections. There is growing interest within the scientific community in the metabolic and epigenetic reprogramming of innate immune cells following an initial stress, which alters their response upon exposure to a heterologous challenge. In this study, the authors show that exposure to extracellular ATP can shape AM functions by activating the P2X7 receptor. This activation triggers the relocation of the potassium channel TWIK2 to the cell surface, placing macrophages in a heightened state of responsiveness. This leads to the activation of the NLRP3 inflammasome and, upon bacterial internalization, to the translocation of TWIK2 to the phagosomal membrane, enhancing bacterial killing through pH modulation. Through these findings, the authors propose a mechanism by which ATP acts as a danger signal to boost the antimicrobial capacity of AMs.

      Strengths:

      This is a fundamental study in a field of great interest to the scientific community. A growing body of evidence has highlighted the importance of metabolic and epigenetic reprogramming in innate immune cells, which can have long-term effects on their responses to various inflammatory contexts. Exploring the role of ATP in this process represents an important and timely question in basic research. The study combines both in vitro and in vivo investigations and proposes a mechanistic hypothesis to explain the observed phenotype.

      Weaknesses:

      Although these findings are convincing and intrinsically interesting, they do not support the conclusion that ATP induces trained immunity. By definition, trained immunity refers to long-lasting metabolic and epigenetic reprogramming initiated by a primary stimulus. Importantly, some of these changes persist after the cells have returned to a basal activation state, thereby generating an altered response upon secondary stimulation (https://doi.org/10.1038/s41590-020-00845-6). In the present study, the data demonstrate a sustained increase in inflammasome activation and enhanced microbicidal activity for up to seven days following ATP exposure. While this sustained activation is noteworthy as well as metabolic shift, it does not demonstrate the existence of trained immunity. The terms priming or sustained activation would therefore be more appropriate than trained immunity.

      Similarly, the observation of increased chromatin accessibility at inflammasome-related genes is expected given the robust activation of this pathway. The presence of open chromatin at these loci does not, by itself, constitute evidence for long-term trained immunity. The authors should therefore be cautious with their terminology and avoid overinterpreting their findings.

      The authors have revised the manuscript to address the comments raised during the first rounds of review. However, several figures, figure legends, and methodological sections still require additional adjustments and clarification.

      The Methods section remains incomplete and requires substantial revision. For instance, the methodology used to quantify immune cell populations presented in Figure 2 is still not described. It is not stated how immune cells were isolated and identified (e.g. flow cytometry from lung tissue). No information is provided regarding tissue digestion, cell isolation procedures, or gating strategy (presumably by flow cytometry). These details are essential and should be included, together with the corresponding gating strategy and absolute cell numbers.

      There are inconsistencies throughout the manuscript. For example, the authors report n = 3 in the figure legend 2 and 3 independent experiments, whereas 3 or 4 points are represented in the graphs. This discrepancy is unclear and should be clarified.

      Overall, while the study addresses an interesting biological question, the manuscript would benefit from substantial revision prior to publication. In particular, clarifications and improvements regarding the methodology, data presentation, and interpretation are required to strengthen the rigor and reproducibility of the conclusions. Several of the conclusions extend beyond what is directly supported by the data. In particular, the interpretation that these findings demonstrate trained immunity should be revised, and additional methodological clarifications and corrections are required.

    3. Author response:

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #1 (Public review):

      Please include an uptake control (early time point) or time‑course to distinguish phagocytosis from intracellular killing.

      We agree that distinguishing uptake from intracellular killing is important for interpreting bactericidal assays. Due to a recent transition, we are unable to conduct additional early‑time‑point assays. To address this transparently, we have revised the manuscript to clarify that our measurements represent overall bacterial load reduction, reflecting the combined effects of uptake and killing.

      The normalization as ‘fold killing’ is non‑standard; please report absolute CFU (log scale).

      We have retrieved the raw data and re‑expressed all bactericidal activity measurements as absolute CFU. All relevant figures, legends, and text have been updated accordingly.

      Authors report quantification of cytokine concentrations, yet no information is provided regarding how these measurements were performed.

      Cytokine concentrations were quantified by ELISA. We have now added details regarding assay kits, sample preparation, and detection parameters to the Methods section.

      While the choice of IL‑1β and IL‑6 is straightforward, the focus on IL‑18 requires explicit justification.

      IL‑18 is a macrophage‑associated pro‑inflammatory cytokine with established links to inflammasome activation and trained immunity pathways, providing a clear justification for its inclusion.

      The methodology used to quantify immune cell populations presented in Figure 2 is not described.

      We have added a detailed description of the flow cytometry methodology, including staining and gating strategies.

      Immune cell quantification would be expected in the context of the challenge experiment as well.

      While we agree that such data would be valuable, additional mouse experiments cannot be performed because the animals used in the challenge model are not currently available during the transition period. If feasible, we are exploring ex vivo flow cytometry data from TWIK2 mutant versus wild‑type macrophages.

      AMs are not considered recruited immune cells; this should be corrected.

      We have corrected this in the figure legend and throughout the manuscript.

      The authors report n = 5 for the survival curves in the figure legend, whereas n = 7 is stated in the Methods section.

      We have corrected the sample size to ensure consistency between the figure legend and Methods.

      ATAC‑seq peaks are referred to as ‘genes’ and ‘differentially expressed genes’.

      We have corrected the terminology in the manuscript. ATAC‑seq identifies differentially accessible chromatin regions, which are then annotated to the nearest downstream gene.

      In Figure 7, trained WT and Nlrp3-/- mice display similar levels of bacterial clearance. How should this result be interpreted?

      A portion of this phenotype is via the normalization of phagocytosis to ‘fold killing’. Presentation of the raw CFU data shows a trend towards reduced bacterial clearance in Nlrp3-/- mice.

      Reviewer #2 (Public review):

      Sample numbers for experiments 1, 2, and 6 are not provided.

      We have added explicit n values for all experiments and verified their accuracy against the original records.

      The Discussion would benefit from a clear summary of study caveats.

      We agree and have added a dedicated paragraph outlining key Caveats as described.

      Specific identities of DEGs are not provided; only pathway enrichment is shown.

      We have now included a supplementary table listing the differentially expressed genes identified in our analysis.

      Controls for subcellular fractionation and dye microscopy should be included.

      Controls for subcellular fractionation have added in figure 3B and controls for dye microscopy have now been added in supplementary figure 1.

      The text states that protease inhibitors diminish ATP‑induced training effects, but the figure does not show significance.

      We have re‑examined the data and updated the figure to include statistical testing where appropriate.

    1. eLife Assessment

      This manuscript reports valuable results on the role of MDC1 and Treacle in DSB repair in rDNA repeats. It has been previously established that MDC1 is replaced by Treacle as the main adaptor in the nucleolar DNA damage response. This work provides convincing evidence that MDC1 is required for the recruitment of RAD51 and BRCA1 to DSBs in rDNA. The work involves multiple MDC1 knockout models and establishes that RFN8-RNF168 act downstream of MDC1in parallel with the RAP80-ABRAXAS pathway in the recruitment of the HR machinery to nucleolar DSBs.

    2. Reviewer #1 (Public review):

      [Editors' note: this version has been assessed by the Reviewing Editor without further input from the original reviewers. The revised version of the manuscript addresses the previous concerns. Importantly, a major role of the RAP80-ABRAXAS pathway is now demonstrated in the recruitment of BRCA1, PALB2, and RAD51 to nucleolar DSBs.]

      This study elucidates the molecular linkage between the mobilization of damaged rDNA from the nucleolus to its periphery and the subsequent repair process by HDR. The authors demonstrate that the nucleolar adaptor protein Treacle mediates rDNA mobilization, and the MDC1-RNF8-RNF168 pathway coordinates the recruitment of the BRCA1-PALB2-BRCA2 complex and RAD51 loading. This stepwise regulation appears to prevent aberrant recombination events between rDNA repeats. This work provides compelling evidence for the recruitment of the Treacle-TOPBP1-NBS1 complex to rDNA DSBs and demonstrates the critical role of MDC1 in the rDNA damage response. There are some issues with the over-interpretation of results as described subsequently. Some aspects could be strengthened, for example, a potential role of the RAP80-Abraxas axis, the origin of the repair synthesis (HDR vs. NHEJ), and a direct comparison of the RNF8 and RNF168 recruitment in the absence or presence of MDC1.

    3. Reviewer #2 (Public review):

      Summary:

      DNA double-strand breaks (DSB) in repeated DNA pose a challenge for repair by homologous recombination (HR) due to the potential of generating chromosomal aberrations, especially involving repeats on different chromosomes. This conceptual caveat led to a long-held notion that HR is not active in repeated DNA, which was disproven in groundbreaking work by Chiolo showing in Drosophila that DSBs in pericentromeric repeats are mobilized to the nuclear periphery for repair by HR. A similar mechanism operates in mouse cells, as shown by the Gautier laboratory, but the mobilization goes to the nucleolar periphery, called nucleolar caps. In this manuscript, the authors reexamine the role of MDC1 in the mobilization of DSBs in rDNA in human cells. Previous work has shown that MDC1 is replaced by Treacle, the gene associated with Treacher Collins syndrome 1, in its role as the main adaptor of the DNA damage response, and these results are confirmed here. The novelty of this contribution lies in the discovery that MDC1 is required downstream in the recruitment of BRCA1 and RAD51 to nucleolar DSBs that were mobilized to the nucleolar cap. Using multiple MCD knockout models and DSBs induced by the nuclease PpoI, which cleaves at nuclear sites as well as in the 28S rDNA, convincingly documents this role of MDC1 and shows that it acts upstream of the RNF8-RNF168 ubiquitylation axis. Using a proxy assay of co-localization of EdU incorporation at DSBs (gammaH2AX), evidence is provided that MDC1 is required for HR in rDNA. MDC1 was not required for RAD51 recruitment to IR-induced foci, but it is unclear whether this is related to the different DSB chemistry (enzymatic versus IR) or to the localization of the DSB (rDNA versus unique sequence genome).

      Strengths:

      (1) The manuscript is well-written, and the experimental evidence is nicely presented.

      (2) Multiple MDC1 knockout models are used to validate the results.

      (3) Convincing back-complementation data clarify the relationship between MDC1 and RNF8.

    4. Author response:

      The following is the authors’ response to the original reviews.

      Reviewer #1 (Public review):

      This study elucidates the molecular linkage between the mobilization of damaged rDNA from the nucleolus to its periphery and the subsequent repair process by HDR. The authors demonstrate that the nucleolar adaptor protein Treacle mediates rDNA mobilization, and the MDC1-RNF8-RNF168 pathway coordinates the recruitment of the BRCA1-PALB2-BRCA2 complex and RAD51 loading. This stepwise regulation appears to prevent aberrant recombination events between rDNA repeats. This work provides compelling evidence for the recruitment of the Treacle-TOPBP1-NBS1 complex to rDNA DSBs and demonstrates the critical role of MDC1 in the rDNA damage response. There are some issues with the over-interpretation of results as described subsequently. Some aspects could be strengthened, for example, a potential role of the RAP80-Abraxas axis, the origin of the repair synthesis (HDR vs. NHEJ), and a direct comparison of the RNF8 and RNF168 recruitment in the absence or presence of MDC1.

      We thank the reviewer for the positive assessment of our work and for the constructive suggestions. We agree that certain aspects of the manuscript required clarification, in particular the potential contribution of the RAP80–ABRAXAS pathway, the interpretation of the repair synthesis assay, and the role of RNF168 recruitment. We have addressed these points experimentally where feasible and have revised the manuscript accordingly to avoid overinterpretation.

      Reviewer #1 (Recommendations for the authors):

      Major comments

      (1) In Figures 4C, 4D, and S4B-D, BRCA1 and RAD51, recruitment to nucleolar caps is partially reduced upon RNF168 depletion. Despite this, the authors broadly conclude that recruitment mainly depends on the MDC1-RNF8-RNF168 pathway. Since the RAP80-Abraxas pathway may also contribute, as briefly mentioned in the Discussion, siRNA knockdown of Abraxas would help clarify the relative roles of these two pathways.

      We thank the reviewer for this important suggestion. To directly address the potential contribution of the RAP80–ABRAXAS pathway, we depleted RAP80 by siRNA and analysed BRCA1 and RAD51 recruitment to nucleolar caps following I-PpoI-induced rDNA damage.

      Strikingly, RAP80 depletion strongly impaired the formation of both BRCA1 and RAD51 nucleolar caps in two independent cell lines (U2OS and RPE1) (new Figure 5). These findings demonstrate that the RAP80–ABRAXAS pathway plays a critical role in BRCA1 recruitment at nucleolar caps.

      Together with our observation that RNF168 depletion only partially reduces BRCA1 and RAD51 recruitment, these results indicate that both RNF168-dependent and RAP80–ABRAXAS-dependent pathways contribute to HDR factor recruitment downstream of RNF8-mediated chromatin ubiquitylation.

      We have revised the model Figure (Figure 9) and the Results and Discussion sections accordingly to reflect this dual-pathway model and to avoid overemphasising the contribution of RNF168.

      (2) In Figure 7C, the EdU-γH2AX PLA assay detects DNA synthesis at rDNA breaks, but it remains unclear whether this signal reflects HDR- or NHEJ-mediated repair. Since MDC1 functions upstream of the DSB repair pathway choice, the observed reduction in PLA signal upon MDC1 depletion does not necessarily reflect impaired HDR alone. Synchronizing cells in G2 or using cell cycle markers would help clarify the repair context and strengthen the interpretation.

      We thank the reviewer for raising this important point. We agree that the EdU–gH2AX PLA assay does not exclusively report on HDR-mediated DNA synthesis and may also capture other forms of repair-associated DNA synthesis.

      In the revised manuscript, we have therefore tempered our interpretation and now describe this assay more cautiously as a readout of DNA synthesis at sites of rDNA damage, rather than as a direct measure of HDR activity.

      Importantly, our conclusion that MDC1 promotes HDR factor recruitment at nucleolar caps is based primarily on the reduced accumulation of BRCA1, PALB2, and RAD51, which are well-established markers of HDR. The PLA assay is now presented as supportive evidence for ongoing DNA synthesis at these sites rather than as a definitive indicator of HDR.

      We agree that further experiments, such as cell cycle synchronization or the use of phase-specific markers, would help to more precisely define the repair context, and we have included this point in the Discussion.

      (3) The authors propose that MDC1 is essential for RNF8-RNF168 recruitment, specifically at nucleolar rDNA breaks. A side-by-side comparison of RNF8 or RNF168 localization in the presence and absence of MDC1, with IR-treated conditions, would provide important validation of this model. Including representative images in Figure S2C would further support the claim.

      We agree with the reviewer that a direct analysis of RNF8 and RNF168 recruitment in the presence and absence of MDC1 would provide valuable mechanistic insight. We therefore attempted to address this experimentally.

      However, despite testing multiple antibodies, we were unable to obtain specific and reproducible signals for RNF8 and RNF168 at nucleolar caps, precluding a reliable analysis of its recruitment under these conditions.

      Given this technical constraint, we have revised the manuscript to avoid overinterpretation regarding direct RNF168 recruitment and instead focus on functional readouts of downstream ubiquitylation-dependent signalling, such as BRCA1 and RAD51 accumulation.

      We note that the requirement for MDC1 in BRCA1 and RAD51 recruitment at nucleolar caps is consistent with a role of MDC1 upstream of RNF8-dependent chromatin ubiquitylation, in line with its established function at IR-induced DSBs.

      Minor comments:

      (1) The legend for Figure 8 should more clearly explain the proposed mechanism and include concise titles or descriptions for each sub-panel.

      We agree with the reviewer that the model should be described in the Figure legend. We have thus updated the model to accommodate the new data and wrote a legend that concisely explains the proposed model. We do not think that titles for each sub-panel are required. Instead, we separately referred to the sub-panels in the legend.

      (2) Typos:

      (a) Page 8: PRE1 MDC1, as "RPE1 MDC1;

      (b) S3 Figure legend: Dhermacon";

      (c) Page 29: "80.103"-please clarify or correct.

      We thank the reviewer for pointing out these errors. These have been corrected in the revised manuscript

      Reviewer #2 (Public review):

      Summary:

      DNA double-strand breaks (DSB) in repeated DNA pose a challenge for repair by homologous recombination (HR) due to the potential of generating chromosomal aberrations, especially involving repeats on different chromosomes. This conceptual caveat led to a long-held notion that HR is not active in repeated DNA, which was disproven in groundbreaking work by Chiolo showing in Drosophila that DSBs in pericentromeric repeats are mobilized to the nuclear periphery for repair by HR. A similar mechanism operates in mouse cells, as shown by the Gautier laboratory, but the mobilization goes to the nucleolar periphery, called nucleolar caps. In this manuscript, the authors reexamine the role of MDC1 in the mobilization of DSBs in rDNA in human cells. Previous work has shown that MDC1 is replaced by Treacle, the gene associated with Treacher Collins syndrome 1, in its role as the main adaptor of the DNA damage response, and these results are confirmed here. The novelty of this contribution lies in the discovery that MDC1 is required downstream in the recruitment of BRCA1 and RAD51 to nucleolar DSBs that were mobilized to the nucleolar cap. Using multiple MCD knockout models and DSBs induced by the nuclease PpoI, which cleaves at nuclear sites as well as in the 28S rDNA, convincingly documents this role of MDC1 and shows that it acts upstream of the RNF8-RNF168 ubiquitylation axis. Using a proxy assay of co-localization of EdU incorporation at DSBs (gammaH2AX), evidence is provided that MDC1 is required for HR in rDNA. MDC1 was not required for RAD51 recruitment to IR-induced foci, but it is unclear whether this is related to the different DSB chemistry (enzymatic versus IR) or to the localization of the DSB (rDNA versus unique sequence genome).

      Strengths:

      (1) The manuscript is well-written, and the experimental evidence is nicely presented.

      (2) Multiple MDC1 knockout models are used to validate the results.

      (3) Convincing back-complementation data clarify the relationship between MDC1 and RNF8.

      Weaknesses:

      (1) The recruitment of BRCA2 was not directly demonstrated. This caveat could be recognized, as IF for BRCA2 is challenging.

      (2) PpoI also induces DSBs in the non-rDNA genome. These DSBs would be an ideal control to establish nucleolar specificity of the events described and clarify whether the difference between IR and PpoI is the chemical structure of the DSB or the location of the DSB.

      We thank the reviewer for the positive and insightful evaluation of our work. We appreciate the recognition of the conceptual advance and the robustness of our experimental approaches. We have carefully considered the reviewer’s suggestions and have revised the manuscript to clarify interpretation where appropriate, particularly regarding BRCA2 recruitment and the specificity of I-PpoI-induced DNA damage. Where possible, we have also added new analyses to strengthen the conclusions.

      Reviewer #2 (Recommendations for the authors):

      (1) The claim that the BRCA1-PALB2-BRCA2 is recruited (abstract, end of results section, discussion page 15) should be qualified as BRCA2 recruitment was not directly demonstrated.

      We thank the reviewer for this important point. We agree that BRCA2 recruitment was not directly demonstrated in our study, as reliable immunofluorescence detection of BRCA2 remains technically challenging.

      We have therefore revised the manuscript throughout (Abstract, Results, and Discussion) to avoid overstatement and now refer more precisely to the recruitment of BRCA1, PALB2, and RAD51, rather than implying direct recruitment of a BRCA1–PALB2–BRCA2 complex.

      We note that BRCA2 function is supported indirectly by the observed RAD51 loading, which depends on BRCA2 activity. However, we have clarified this point to ensure that our conclusions remain fully supported by the presented data.

      (2) The temporal sequence established in Figure 1, 1hr BRCA1 and 2 hrs PALB2, argues against recruitment of a stable BRCA1-PALB2-(BRCA2) complex. This should be acknowledged.

      We thank the reviewer for this insightful observation. We agree that the temporal separation between BRCA1 accumulation (1 h) and PALB2/RAD51 recruitment (2 h) argues against the recruitment of a pre-assembled, stable BRCA1–PALB2–BRCA2 complex.

      We have revised the manuscript to reflect this interpretation and now describe the recruitment of HDR factors as a sequential process rather than as the assembly of a pre-formed complex. This is consistent with current models in which BRCA1 promotes subsequent PALB2 and BRCA2 recruitment, ultimately leading to RAD51 loading.

      (3) The model predicts that MDC1-KO cells are proficient for transcriptional repression after nucleolar DSB induction. Has this been tested?

      We did not specifically test this in the current work, but previous results published by our group revealed that siRNA-mediated depletion of MDC1 in human cells had a minimal effect on rDNA transcriptional inhibition after DNA damage (Larsen et al., 2024).

      (4) The nuclear PpoI DSBs could be analyzed as a specificity control, and clarify whether the difference between IRIF and PpoI DSBs relates to the DSB chemistry or location.

      We thank the reviewer for this important point. We agree that I-PpoI induces DNA breaks both within rDNA repeats and at additional genomic loci.

      To address this, we have now analysed the formation of gH2AX-positive nucleolar caps and non-nucleolar gH2AX foci over time following I-PpoI expression (new Figure 1–figure supplement 2). We find that nucleolar caps form rapidly and are prominent at early time points, whereas gH2AX foci accumulate more gradually.

      These results indicate that nucleolar caps and non-nucleolar DNA damage responses can be distinguished both spatially and temporally, and support the use of nucleolar caps as a specific readout for rDNA damage in our study.

      In addition, we note that RAD51 recruitment to IR-induced foci is not affected by MDC1 loss, suggesting that the requirement for MDC1 in RAD51 loading is specific to nucleolar rDNA breaks rather than reflecting differences in DSB chemistry alone. We have clarified this point in the Discussion.

      Additional points:

      (5) Page 4 top: Shieldin.

      Corrected.

      (6) The general reader will be interested to learn about the connection of the Treacle function with Treacher Collins syndrome. Maybe a paragraph could be added to discuss this?

      We thank the reviewer for this suggestion. We agree that the relationship between Treacle and Treacher Collins syndrome may be of interest to a broad readership. Since the developmental pathology of Treacher Collins syndrome is currently thought to arise primarily from impaired ribosome biogenesis and nucleolar dysfunction rather than defective nucleolar DNA damage signalling, we felt that an extensive discussion would be beyond the scope of the present study. We have, however, added a brief statement introducing Treacle as the product of the TCOF1 gene mutated in Treacher Collins syndrome and noting that whether its DNA damage response function contributes to disease pathology remains an open question.

      (7) Figure 7: A short explanation could be added as to why hypoxia conditions were chosen for the p53-deficient cell lines.

      We thank the reviewer for pointing this out. We have added a brief explanation in the figure legend to clarify that hypoxia conditions were used to stabilise replication stress and enhance detection of DNA repair intermediates in p53-deficient cells.

      (8) A short statement on whether the repair of nuclear DBS is affected by Treacle could be added.

      We thank the reviewer for this interesting point. While our study focuses on nucleolar DNA damage, we did not observe evidence that Treacle is required for the repair of non-nucleolar DSBs. We have added a brief statement in the Discussion to clarify that Treacle appears to function specifically in the nucleolar DNA damage response.

    1. eLife Assessment

      This study offers valuable insights into the role of post-translational modifiers, specifically SUMO2ylation at K81 in p66Shc, and its impact on endothelial function through reactive oxygen species. A series of compelling experiments demonstrated that lysine 81 of p66Shc is the site of SUMO2 conjugation, which is crucial for mitochondrial localization and essential for S36 phosphorylation, leading to specific pathological effects. The combination of cell overexpression and animal studies provides solid data supporting this mechanistic link.

    2. Reviewer #2 (Public review):

      Summary:

      The manuscript titled "p66Shc Mediates SUMO2-induced Endothelial Dysfunction" by Kumar et al. builds upon established literature demonstrating that both p66Shc and SUMOylation are essential players in nitric oxide (NO)-mediated endothelial vascular homeostasis and development (PMID: 10580504, 28760777, and 35187108).

      In this study, the authors uncover a novel mechanism showing how the SUMO2ylation of p66Shc drives reactive oxygen species (ROS) production in endothelial cells. Specifically, they identify Lysine 81 (K81) as the critical residue on p66Shc conjugated to SUMO2, proving it is essential for the protein's mitochondrial localization.

      The authors convincingly demonstrate that:

      p66Shc is actively SUMO2ylated at the K81 site in cellular models.

      Phosphorylation at Serine 36 (S36) is significantly reduced upon the loss of this critical SUMOylation site.

      Conclusion:

      Overall, this study provides strong evidence for a novel regulatory axis in endothelial cells. It successfully opens the door for further dissection of the complex mechanistic crosstalk between three key post-translational modifications on p66Shc: S36 phosphorylation, K81 SUMO2ylation, and acetylation.