10,000 Matching Annotations
  1. Sep 2026
    1. Reviewer #3 (Public review):

      Summary:

      Recently, the off-target activity of antibiotics on human mitoribosome has been paid more attention in the mitochondrial field. Hafner et al applied mitoribosome profiling to study the effect of antibiotics on protein translation in mitochondria as there are similarities between bacterial ribosome and mitoribosome. The authors conclude that some antibiotics act on mitochondrial translation initiation by the same mechanism as in bacteria. On the other hand, the authors showed that chloramphenicol, linezolid and telithromycin trap mitochondrial translation in a context-dependent manner. More interesting, during deep analysis of 5' end of ORF, the authors reported the alternative start codon for ND1 and ND5 proteins instead of previously known one. This is a novel finding in the field and it also provide another application of the technique to further study on mitochondrial translation.

      Strengths:

      This is the first study which applied mitoribosome profiling method to analyze multiple antibiotics treatment cells. The mitoribosome profiling method had been optimized carefully and has been suggested to be a novel method to study translation events in mitochondria. The manuscript is constructive and well-written.

      Comments on revisions:

      The authors added a discussion to the revised manuscript, and also carefully investigate structural data from others. I have no more comment. Congratulations to the team for a good manuscript!

    2. Author response:

      The following is the authors’ response to the previous reviews

      Reviewer #1 (Public review):

      Summary:

      This study aimed to determine whether bacterial translation inhibitors affect mitochondria through the same mechanisms. Using mitoribosome profiling, the authors found that most antibiotics, except telithromycin, act similarly in both systems. These insights could help in the development of antibiotics with reduced mitochondrial toxicity.

      They also identified potential novel mitochondrial translation events, proposing new initiation sites for MT-ND1 and MT-ND5. These insights not only challenge existing annotations but also open new avenues for research on mitochondrial function.

      Strengths:

      Ribosome profiling is a state-of-the-art method for monitoring the translatome at very high resolution. Using mitoribosome profiling, the authors convincingly demonstrate that most of the analyzed antibiotics act in the same way on both bacterial and mitochondrial ribosomes, except for telithromycin. Additionally, the authors report possible alternative translation events, raising new questions about the mechanisms behind mitochondrial initiation and start codon recognition in mammals.

      Weaknesses:

      All the weaknesses I previously highlighted were adequately addressed.

      We thank the reviewer for carefully considering our revision.

      Reviewer #3 (Public review):

      Summary:

      Recently, the off-target activity of antibiotics on human mitoribosome has been paid more attention in the mitochondrial field. Hafner et al applied mitoribosome profilling to study the effect of antibiotics on protein translation in mitochondria as there are similarities between bacterial ribosome and mitoribosome. The authors conclude that some antibiotics act on mitochondrial translation initiation by the same mechanism as in bacteria. On the other hand, the authors showed that chloramphenicol, linezolid and telithromycin trap mitochondrial translation in a context-dependent manner. More interesting, during deep analysis of 5' end of ORF, the authors reported the alternative start codon for ND1 and ND5 proteins instead of previously known one. This is a novel finding in the field and it also provide another application of the technique to further study on mitochondrial translation.

      Strengths:

      This is the first study which applied mitoribosome profiling method to analyze mutiple antibiotics treatment cells. The mitoribosome profiling method had been optimized carefully and has been suggested to be a novel method to study translation events in mitochondria. The manuscript is constructive and well-written.

      Weaknesses:

      This is a novel and interesting study, however, most of conclusion comes from mitoribosome profiling analysis, as the result, the manuscript lacks the cellular biochemical data to provide more evidence and support the findings.

      Comments on revisions:

      The authors addressed most of my concerns and comments, although there is still no biochemical assay which should be performed to support mitoribsome profiling data.

      The author also carefully investigated the structure of complex I, however, I am surprised that the author chose to analyse a low-resolution structure (3.7 A). Recently, there are more high-resolution structures of mammalian complex I published (7R41, 7V2C, 7QSM, 9I4I). Furthermore, the authors should not only respond to the reviewers but also (somehow) discuss these points in the manuscript.

      We thank the reviewer for suggesting additional structural analyses. Of the suggested additional structures to look at, only 9I4I from Nguyen et al. 2026 was of human-derived complex I. Other structures were derived from species that utilized alternative codons at the 3’ end of these mRNAs, preventing comparison. However, for 9I4I, the authors were able to fit density for every amino acid of all mitochondrially-encoded proteins in complex I, except the first two residues of ND1 and ND5, which, as our manuscript suggest, are not translated.

      In addition to structures, we also attempted to identify additional publicly available mass spectrometry data for complex I. However, data which we identified either derived their proteins from bovine tissue, which utilizes different codons at the initiation sites, or did not capture the N-terminal region of the proteins. Therefore, we did not include this analysis.

    1. eLife Assessment

      This study provides important insights into the crosstalk between ATG2A and components of the early secretory pathway. The authors employ an elegant proximity-labelling approach to identify two ER-Golgi intermediate compartment (ERGIC)-localized proteins. They further identify ARFGAP1 and Rab1A as components of early autophagic membranes, which accumulate at the periphery of pre-autophagosomal structures induced by loss of ATG2. Overall, the study is well executed, and the evidence supporting the claims is convincing.

    2. Reviewer #1 (Public review):

      [Editors' note: this version has been assessed by the Reviewing Editor without further input from the original reviewers. The authors have addressed the comments raised in the previous round of review.]

      Summary:

      D. Fuller et al. set out to study the molecular partners that cooperate with ATG2A, a lipid transfer protein essential for phagophore elongation, during the process of autophagy. Through a series of experiments combining microscopy and biochemistry, the authors identify ARFGAP1 and Rab1A as components of early autophagic membranes, which accumulate at the periphery of aberrant pre-autophagosomal structures induced by loss of ATG2. While ARFGAP1 has no apparent function in autophagy, the authors show that RAB1A is implicated in autophagy, although the precise mechanisms are not explored in the manuscript.

      Strengths:

      The work presented by Fuller et al. provides new insights into the composition of early autophagic membranes. The authors provide a series of MS experiments identifying proteins in close proximity to ATG2A, which is a valuable dataset for the field. Furthermore, they show for the first time the interaction between ATG2A and RAB1A both in fed and starved conditions, which extends the characterisation of the pre-autophagosomal structures observed in ATG2 DKO cells.

    3. Reviewer #2 (Public review):

      The mechanisms governing autophagic membrane expansion remain incompletely understood. ATG2 is known to function as a lipid transfer protein critical for this process; however, how ATG2 is coordinated with the broader autophagic machinery and endomembrane systems has remained elusive. In this study, the authors employ an elegant proximity labeling approach and identify two ER-Golgi intermediate compartment (ERGIC)-localized proteins-Rab1 and ARFGAP1-as novel regulators of ATG2 during autophagic membrane expansion.

      Their findings support a model in which autophagosome formation occurs within a specialized subdomain of the ER that is enriched in both ER exit sites (ERES) and ERGIC, providing valuable mechanistic insight. The overall study is well executed and offers an important contribution to our understanding of autophagy. I support its publication in eLife and offer the following minor comments for clarification and improvement.

    4. Reviewer #3 (Public review):

      The manuscript by Fuller et al describes a crosstalk between ARTG2A with components of the early secretory pathway, namely RAB1A and ARFGAP1. They show that ATG2A is recruited to membranes positive for RAB1A, which they also show to interact with ATG2A. In agreement with earlier findings by other groups, silencing RAB1A negatively affects autophagy. While ARFGAP1 was also found on ATG2A positive membranes, silencing ARFGAP1 had no impact autophagy. Notably, these ARFGAP1 positive membranes are not Golgi membranes.

      The findings are interesting and the data are in general of good quality.

      Comments on the previous version:

      The revisions carried out by the authors are fine. The new data on ArfGAP1 and about the indirectness of the ATG2A and Rab1A interaction improve both clarity and strength of the manuscript. I have no further comments.

    5. Author response:

      The following is the authors’ response to the previous reviews

      (1) Interpretation of LC3-II accumulation and phenocopying

      Therefore, LC3-II accumulation alone is insufficient to support phenocopying in my view.

      We agree with this assessment. Upon reconsideration, we concluded that LC3-II accumulation alone does not justify the use of the term "phenocopying." We have therefore removed this language from the manuscript.

      (2) Strength of conclusions regarding autophagosome biogenesis

      As presented, the findings support a correlative relationship rather than a defined role in autophagosome biogenesis.

      We agree that our original wording overstated the strength of the conclusions. To better reflect the data, we revised the text to state that our findings "expound upon" rather than "elucidate" the role of these membranes in autophagosome biogenesis.

      (3) Title wording

      The title states that ATG2A ‘engages’ Rab1A- and ARFGAP1-positive membranes during autophagosome formation... A more descriptive term, such as ‘associates,’ would more accurately reflect the data.

      We appreciate this suggestion. We revised the title to avoid implying a causal dependency. The title now states that ATG2A interacts with Rab1A- and ARFGAP1-positive membranes, emphasizing the membrane association observed in our study rather than a direct interaction with the proteins themselves.

      (4) ARFGAP1 knockdown phenotype

      The authors claim: ‘siRNA against ARFGAP1 had very little effect’ but the quantification and blots show actually no effect.

      We agree that the original wording was imprecise. The sentence has been revised to state:

      “siRNA against ARFGAP1 had no effect on flux.”

      This more accurately reflects both the data and our original interpretation.

      (5) Interpretation of ARFGAP knockdown experiments

      Conclusions drawn from KD experiments in Fig. S2 should be interpreted with caution, as knockdown efficiency is very low, particularly for ARFGAP1/3 in the triple knockdown.

      We agree with this caution and have revised the text accordingly. The manuscript now states:

      “Knockdown of ARFGAP2, ARFGAP3 and ARFGAP1-3 marginally increased autophagic flux (Fig. S2B,C), suggesting either no role or a minor role as negative regulators of autophagy.”

      This wording more appropriately reflects the limitations of the experiment.

      (6) Discussion of ERGIC/ERES remodeling literature

      It would strengthen the manuscript to discuss previous studies reporting ERES and ERGIC remodeling and formation of ERES-ERGIC contact sites (PMID: 34561617; PMID: 28754694).

      We appreciate the suggestion. These studies were already cited and discussed in the original submission, and therefore no additional changes were required.

      (7) Figure readability

      The font size in Figure 1A and Supplementary Figure S1G is too small for comfortable reading.

      We agree and have enlarged the labels in both figures to improve readability.

      (8) Clarification of starvation conditions in figure legends

      In Figures 2A-C and Figure 4, it is unclear how the cells were treated. Were they starved in EBSS?

      We have updated the corresponding figure legends to explicitly state the starvation conditions used in these experiments.

      (9) Interpretation of ARFGAP1 knockdown and LC3 lipidation

      In Figure 2A, ARFGAP1 knockdown appears to reduce LC3 lipidation without affecting Halo-LC3 cleavage.

      We do not observe a reproducible reduction in LC3 lipidation following ARFGAP1 knockdown and therefore do not believe this conclusion is supported by the data. No changes were made in response to this comment.

      (10) Clarification of protein-protein interaction statement

      The phrase ‘but protein-protein interactions appear to be limited to RAB1’ would benefit from clarification.

      We agree and have adopted the suggested wording. The manuscript now states:

      “but stable protein-protein interactions appear to be limited to RAB1.”

      We thank the reviewers again for their constructive feedback and for helping us improve the clarity and accuracy of the manuscript.

    1. eLife Assessment

      These results add to a growing body of literature demonstrating a role for IL-21 in T follicular helper cell expansion and differentiation, as well as in germinal center formation and maintenance. The authors use an IL21R ENU-induced mutant, together with complementary in vitro and in vivo experiments, to investigate the roles of STAT1 and STAT3 signaling downstream of IL-21R. The experimental approach is valuable, but the main conclusion is incompletely supported by the evidence, as the mutant exhibits a broad reduction in STAT signaling rather than a selective impairment of STAT1 signaling.

    2. Reviewer #1 (Public review):

      Summary:

      King and colleagues generated a mouse with a point mutation in IL21R and investigate the influence on IL-21-mediated T and B cell activation and differentiation. They find that mutant mice show a reduced T and B cell response with CD4 T cell differentiation into T follicular helper cells being primarily affected.

      Strengths:

      The authors combine in vitro and in vivo analysis, including bone-marrow chimeric mice.

      Weaknesses:

      The effect of the IL21R EINS mutant does not specifically affect STAT1, as clearly shown in Figure 1 H, I. Particularly at lower doses of IL21 which may be more relevant in vivo, the effects are very similar. A second key weakness is the very small Tfh response, a not very clear PD-1 and CXCR5 staining to identify Tfh and a lack of a steady-state (prior to immunisation) comparison of Tfh numbers in the different mouse strains. The latter makes it impossible to know what fraction of the response is antigen-specific.

      Comments on revised version.

      Thank you for responding to some of my suggestions/comments.

      Unfortunately, these responses do not address the concerns raised by the other reviewer and me (see 'weaknesses'). The argument that statistical analysis shows that there is no difference in p-STAT3/5 does not make any sense - it simply reflects the absence of sufficient statistical power.

      No further experiments have been performed, and the authors didn't appropriately adjust the conclusions to reflect the limitations of the study.

      To provide an example, the authors decided to keep the title 'An IL-21R hypomorph circumvents functional redundancy to define STAT1 signalling in germinal center responses' although both reviewers clearly highlighted that no conclusion about STAT1 can be made and the authors responded that 'We agree that further experiments are needed to definitively show that the effect is attributable to reduced STAT1 activation alone. Rescue experiments will be a focus of future experiments.'

      These rescue experiments (and additional analysis of the Tfh response) are essential to allow firms conclusions to be made.

      In its current form, the manuscript is unfortunately misleading.

    3. Reviewer #2 (Public review):

      Summary:

      In the manuscript, "An IL-21R hypomorph circumvents functional redundancy to define STAT1 signaling in germinal center responses," Cecile King and colleagues identify a cytoplasmic site of the IL-21 receptor that differentially regulates STAT1 and STAT3 activation upon IL-21 stimulation. They further examine the immunological consequences of this site-specific alteration on Tfh differentiation and Tfh-dependent humoral immunity, raising important questions about how gene-knockout models may obscure nuanced functional roles of signaling molecules.

      Strengths:

      The study convincingly highlights a non-redundant role for STAT1 downstream of IL-21-IL-21R signaling in the Tfh differentiation pathway. This conclusion is supported by in vitro analyses of STAT1 and STAT3 activation in CD4 T cells stimulated with IL-21 or IL-6; by in vivo assessments of Tfh and germinal center B cell responses in WT and IL21R-EINS mutant mice, including bone-marrow chimera systems; and by investigating the expression of Tfh-related molecules in WT versus IL21R-EINS CD4 T cells.

      Weaknesses:

      Although the experiments were carefully executed with appropriate controls, a key question remains unresolved: whether the Tfh differentiation defect in IL21R-EINS mice is directly attributable to reduced STAT1 activation. Rescue experiments that restore STAT1 signaling in IL21R-EINS TCR-transgenic CD4 T cells would provide strong evidence linking the mutation to impaired STAT1 activation and, consequently, defective Tfh differentiation. Without such evidence, it remains formally possible that additional, uncharacterized mutations introduced during ENU mutagenesis contribute to the phenotypes observed, particularly given the discrepancies between IL21R knockout and IL21R-EINS mutant mice.

      Comments on revised version.

      The revised manuscript failed to address the key question, whether the Tfh differentiation defect in IL21R-EINS mice results from the reduced STAT1 activation in CD4 T cells.

    4. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      King and colleagues generated a mouse with a point mutation in IL21R and investigated the influence on IL-21-mediated T and B cell activation and differentiation. They found that mutant mice show a reduced T and B cell response, with CD4 T cell differentiation into T follicular helper cells being primarily affected.

      Strengths:

      The authors combined in vitro and in vivo analysis, including bone-marrow chimeric mice.

      Weaknesses:

      The effect of the IL21R EINS mutant does not specifically affect STAT1, as clearly shown in Figure 1 H, I. Particularly at lower doses of IL21, which may be more relevant in vivo, the effects are very similar. A second key weakness is the very small Tfh response, a not very clear PD-1 and CXCR5 staining to identify Tfh, and a lack of a steady-state (prior to immunisation) comparison of Tfh numbers in the different mouse strains. The latter makes it impossible to know what fraction of the response is antigen-specific.

      Reviewer #2 (Public review):

      Summary:

      In the manuscript, "An IL-21R hypomorph circumvents functional redundancy to define STAT1 signaling in germinal center responses," Cecile King and colleagues identify a cytoplasmic site of the IL-21 receptor that differentially regulates STAT1 and STAT3 activation upon IL-21 stimulation. They further examine the immunological consequences of this site-specific alteration on Tfh differentiation and Tfh-dependent humoral immunity, raising important questions about how geneknockout models may obscure nuanced functional roles of signaling molecules.

      Strengths:

      The study convincingly highlights a non-redundant role for STAT1 downstream of IL-21-IL-21R signaling in the Tfh differentiation pathway. This conclusion is supported by in vitro analyses of STAT1 and STAT3 activation in CD4 T cells stimulated with IL-21 or IL-6; by in vivo assessments of Tfh and germinal center B cell responses in WT and IL21R-EINS mutant mice, including bonemarrow chimera systems; and by investigating the expression of Tfh-related molecules in WT versus IL21R-EINS CD4 T cells.

      Weaknesses:

      Although the experiments were carefully executed with appropriate controls, a key question remains unresolved: whether the Tfh differentiation defect in IL21R-EINS mice is directly attributable to reduced STAT1 activation. Rescue experiments that restore STAT1 signaling in IL21R-EINS TCR-transgenic CD4 T cells would provide strong evidence linking the mutation to impaired STAT1 activation and, consequently, defective Tfh differentiation. Without such evidence, it remains formally possible that additional, uncharacterized mutations introduced during ENU mutagenesis contribute to the phenotypes observed, particularly given the discrepancies between IL21R knockout and IL21R-EINS mutant mice.

      We agree that further experiments are needed to definitively show that the effect is attributable to reduced STAT1 activation alone. Rescue experiments will be a focus of future experiments.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) Figure1

      I would recommend changing the conclusion in Line 141 to 'potentially less affected' rather than unaffected as there is a clear and very consistent effect on pSTAT3 in every IL21 dose tested. Also, it seems that much more IL21 is required to induce STAT1 phosphorylation, which may explain the increased effect of the EINS mutant on this signalling pathway. Similarly, for the STAT5 data in Figure S1, there is very little phosphorylation beyond baseline phosphorylation (unstimulated), but there is a clear and consistent reduction in pSTAT5 in the EINS mutants.

      Neither pSTAT3 nor pSTAT5 were significantly different between WT and IL21rEINS cells in either the percentage or MFI. We have edited line 141 to state “the levels of phosphorylated STAT3 in CD4+ T cells were significantly less affected by the Il21r<sup>EINS</sup> mutation (Fig. 1I).”

      A minor point: Why is the MFI in IL21R-/- mice at 300 in panel C and at 200 in panel E? How representative is the reduced baseline pSTAT5 in IL21r-/- mice? 

      This is likely due to machine voltage during acquisition in a different experiment.

      Collectively, I would suggest concluding from that data that the EINS mutation affects IL21R signaling, which results in reduced STAT1, 3, and 5 phosphorylation, with pSTAT1 being most strongly reduced, particularly at high IL21 concentration. 

      Please also see response to above comment. Since neither pSTAT3 nor pSTAT5 were significantly different between WT and IL21rEINS cells, the data does not support that conclusion.

      All subsequent data therefore do not investigate the effect of the EINS mutation on STAT1, but on overall reduced IL21R signalling. This needs to be considered when interpreting the data. For example, the text in lines 171, 172, and 174 should be adjusted as the effect is neither only dependent on STAT1 nor is STAT3 signalling intact.

      Please also see responses above. It is possible that a different method for detection of phosphorylated STAT3 and STAT5 could have looked more closely into the effect at very low concentrations of IL-21. However, our findings using Westen Blot and flow cytometry only observed a significant difference in STAT1 activation. We have edited our sentence on line 333 to state “response in the presence of an IL-21 receptor mutant that predominantly affects IL-21 activation of STAT1”.

      A key signalling pathway downstream of IL21R is AKT and S6 phosphorylation. It would be important to also investigate the effect of the EINS mutation on these pathways.

      We agree and this will be a focus of future experiments.

      (2) Figure 2

      PNA or BCL-6 staining would be preferable to identify GC in Figure 1A, but the flow cytometry data in Figure 3 are convincing, so this is not absolutely necessary.

      Also, no conclusions can be made here about STAT1 specifically, and there could be other reasons why Tfh are slightly and temporarily reduced in EINS mice.

      (3) Figure 3

      Line 200: Please explain what is meant by 'despite an expansion of the IgG1 FAs B cell population on day5, the percentages of EINS IgG1* GC B cells were significantly lower (Fig. 3F). I cannot see any expansion of IgG1 FAS B cells, nor can I see a specific effect on day 7. Both total GC B cells and IgG1 GC B cells are similarly affected throughout the response. Some data points may not reach statistical significance, but the trend is very clear.

      Figure 3E shows the percentage of GC B cells increasing from day 3 to day 7 in WT and from day 3 to day 5 in IL21rEINS, with significant differences between WT and IL21rEINS on days 5 and 7. In Figure 3F he percentage of IgG1+ GC B cells increase from day 3 to day 14 in both genotypes, with a significant difference between IL21rEINS on day 7. We have edited to manuscript to state “Despite an increase in the IgG1<sup>+</sup> FAS<sup>+</sup> B cell population from day 3 in response to immunization, the percentages of Il21r<sup>EINS</sup> IgG1+ GC B cells were significantly lower relative to WT cells 7 days after SRBC immunisation (Fig. 3F).”

      (4) Figure 4

      How many days after SRBC immunization was the analysis done?

      The data show an intrinsic role of IL21R signalling to Tfh development, which may include a role for STAT1. The absence of any effect on GC B cells is somewhat surprising. Chimeric and irradiated mice sometimes mount poor immune responses, and GC B cell numbers are very low. What is the frequency of GC B cells in non-immunized mice? This would be important to know if the mice responded at all, and if they did, if the 'baseline' of GC activity differs.

      As stated in the figure legend for Figure 4 – on day 7 “Mixed BM chimaeras were reconstituted with equal ratios of WT CD45.1+ BM cells and Il21rEINS CD45.2+ BM cells. 8 weeks after transfer, the mice were immunized with SRBC and analysed 7days later.

      (5) Figure 7

      Please highlight that while only IL21R-/- mice showed a significant difference in the frequency of Tfh, a similar trend was observed in WT and IL21Reins mice. The data spread is smaller in the IL21R-deficient mice, facilitating statistical significance. As throughout the manuscript, this is not a STAT1 IL-21R mutant; it is a mutant with reduced IL21R signalling. In fact, the finding that IL-6 does not compensate for the EINS mutation may suggest that STAT1 plays a minor role in the biological effects observed.

      We can only report on the statistical significance of the data we have.

      (6) Other comments

      Figure 1 F/G. I think the Y axis should read pSTAT1 and pSTAT3, respectively.

      Thank you, we have corrected the graph accordingly.

      Figure S1A. Please change the order of WT, IL21EINS, and IL21R-/- to match the main figures (IL21R-/- last). Currently, A, B, and D have a different order, but C is like the main figures.

      The figure panels are aligned to show media, then either IL-2 or IL-6 and then IL21.

      Please provide complete flow cytometry gating strategies for all figures.

      Flow cytometry dating for P-STAT1 and p-STAT3 is shown in figure 1, for Tfh cells and Tfr cells in Figure 2, 4 and 5. Please also see supplementary figures for T cell gating and methods for detailed description of antibodies and dilutions used for immunostaining.

      Reviewer #2 (Recommendations for the authors):

      Line 332 requires revision.

      We have edited the final sentence to state” Taken together these findings demonstrate that, despite the strong ability of IL-6 to activate STAT1, IL-6 is ineffective at fully compensating for the germinal centre response in the presence of an IL-21 receptor mutant that predominantly affects IL-21 activation of STAT1.”

    1. eLife Assessment

      This important study reports that Sox17 is key to the formation and function of the Sertoli valve, a transition region between the rete testis and seminiferous tubules that remains an understudied domain of testicular biology. The supporting data are solid and convincing. This work will be of interest to developmental and reproductive biologists, as well as andrologists who work on male fertility and men's health.

    2. Reviewer #1 (Public review):

      Summary:

      This manuscript is an excellent follow-up to your 2022 study, in which Sox17 expression was localized to the rete testis and shown to be required for proper formation of the Sertoli cell valve (transition region). By using Nr5a1-Cre to drive conditional deletion of Sox17 specifically in rete testis cells, you demonstrate that testis weights remain normal at 2 weeks of age but become significantly reduced by 8 weeks in Sox17-cKO males. At the later time point, the seminiferous epithelium is severely disrupted, with apparent arrest of spermiogenesis: the epididymal lumen is essentially devoid of sperm, and most tubules lack elongated spermatids.

      Strengths:

      Clearly shows the role of Sox17 in Sertoli cells being important to the SV function. The SV (transition region) between the rete testis and seminiferous tubules remains an understudied domain of testicular biology. The present work, together with your prior study, highlights intriguing mechanisms operating in this specialized niche.

      Weaknesses:

      The available data do not fully explain either the developmental assembly of the Sertoli valve or the precise consequences of its functional disruption. These studies are nonetheless valuable precisely because they raise more questions than they answer; the conceptual implications are thought-provoking.

    3. Reviewer #2 (Public review):

      This manuscript investigates the role of SOX17 in the formation and function of the Sertoli valve (SV) at the interface between seminiferous tubules and the rete testis (RT). Building on previous work showing that rete testis-specific deletion of Sox17 disrupts SV formation, leading to defective spermiogenesis and male infertility, the authors explore how SOX17 overexpression in Sertoli cells regulate SV of rodent testes.

      Using transgenic mouse models with ectopic Sox17 expression in Sertoli cells, the study demonstrates that SOX17 is not only required but can also modulate SV formation. Ectopic expression in Sertoli cells induces expansion of the SV structure and partially rescues SV defects and spermatogenesis in RT-specific Sox17 conditional knockout animals. The data support a model in which SOX17 acts through paracrine signaling to regulate SV formation, although the precise mechanisms remain to be clarified.

      Overall, this is a well-executed study with novel and significant findings. The ability to experimentally manipulate SV size is particularly compelling and provides a valuable framework to study fluid dynamics and epithelial interactions in the testis. This work will be of broad interest to the reproductive biology and developmental biology communities.

    4. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This manuscript is an excellent follow-up to your 2022 study, in which Sox17 expression was localized to the rete testis and shown to be required for proper formation of the Sertoli cell valve (transition region). By using Nr5a1-Cre to drive conditional deletion of Sox17 specifically in rete testis cells, you demonstrate that testis weights remain normal at 2 weeks of age but become significantly reduced by 8 weeks in Sox17-cKO males. At the later time point, the seminiferous epithelium is severely disrupted, with apparent arrest of spermiogenesis: the epididymal lumen is essentially devoid of sperm, and most tubules lack elongated spermatids.

      Strengths:

      The study clearly shows the role of Sox17 in Sertoli cells as being important to SV function. The SV (transition region) between the rete testis and seminiferous tubules remains an understudied domain of testicular biology. The present work, together with the authors' prior study, highlights intriguing mechanisms operating in this specialized niche.

      Weaknesses:

      At the same time, the available data do not yet fully explain either the developmental assembly of the Sertoli valve or the precise consequences of its functional disruption. These studies are nonetheless valuable precisely because they raise more questions than they answer; the conceptual implications are thought-provoking.

      Reviewer #2 (Public review):

      This manuscript investigates the role of SOX17 in the formation and function of the Sertoli valve (SV) at the interface between seminiferous tubules and the rete testis (RT). Building on previous work showing that rete testis-specific deletion of Sox17 disrupts SV formation, leading to defective spermiogenesis and male infertility, the authors explore how SOX17 overexpression in Sertoli cells regulates the SV of rodent testes.

      Using transgenic mouse models with ectopic Sox17 expression in Sertoli cells, the study demonstrates that SOX17 is not only required but can also modulate SV formation. Ectopic expression in Sertoli cells induces expansion of the SV structure and partially rescues SV defects and spermatogenesis in RT-specific Sox17 conditional knockout animals. The data support a model in which SOX17 acts through paracrine signaling to regulate SV formation, although the precise mechanisms remain to be clarified.

      Overall, this is a well-executed study with novel and significant findings. The ability to experimentally manipulate SV size is particularly compelling and provides a valuable framework to study fluid dynamics and epithelial interactions in the testis. This work will be of broad interest to the reproductive biology and developmental biology communities.

      Reviewer #3 (Public review):

      Summary:

      These studies are based on previously published work that showed that deletion of expression of the Sox17 gene in the testis essentially deleted the formation of the Sertoli valve in the Rete testis. The authors extended this work by constructing a vector that resulted in increased Sox17 expression by Sertoli cells and enhanced formation of the Sertoli valve in both wild type and Sox17 knockout mice. The work provides strong evidence supporting the requirement for Sox17 expression to allow formation of the Sertoli valve.

      Strengths:

      The general approach was to express Sox17 from a Tg mouse that expressed Sox17 from Sertoli cells. This Tg mouse was bred into both the WT and the Sox17 KO mouse. The Sertoli valve was enhanced in both the WT/Tg mouse and KO/Tg mouse, showing that ectopic Sox17 could compensate in the Sox17 Ko and act in a concentration-dependent manner in the WT mouse. The results are strong and support the conclusions from the authors. The results were as expected from the original paper describing the KO of Sox 17. These results strengthen these conclusions and provide ideas for additional conclusions. These studies were technically challenging, and the authors provided a very solid manuscript.

      Weaknesses:

      The authors refer several times to high or low expression, but it all appears to be based on immunohistochemistry, and there is no real quantification using PCR, for example. The process used for cell quantification lacks a rationale for why certain numbers were assigned.

      We sincerely thank the reviewers for their careful evaluation of our manuscript and for their constructive and encouraging comments. We are grateful for the recognition of the significance of the Sertoli valve as an understudied transition region between the rete testis and seminiferous tubules, as well as for the positive assessment of our genetic approach and the evidence that ectopic SOX17 expression can modulate SV formation. We have carefully considered all points raised in the assessment and have revised the manuscript accordingly. The major revisions include:

      (1) Clarification of the scope and limitations of the study (Reviewers #1 and #2):

      In response to the comments that the developmental assembly of the Sertoli valve and the precise consequences of its functional disruption remain incompletely understood, we clarified the scope and limitations of the present study at the end of 7th paragraph in the Discussion. Although our findings support a model in which SOX17 regulates SV formation through paracrine signaling, the downstream effectors and precise molecular mechanisms remain to be identified. We therefore revised the Discussion to avoid overinterpretation of the molecular mechanisms and to emphasize that comprehensive mechanistic analyses, including transcriptomic analyses using the Tg mouse model, represent an important direction for future research. We also added histological analyses of the earliest detectable lesions at 4 weeks of age and low-magnification images of adult Sox17 cKO testes (new Figure S1), revealing selective sloughing of round spermatids despite preserved Sertoli cell architecture and subsequent mosaic spermatogenic defects among individual seminiferous tubules. These observations provide additional insights into the altered luminal microenvironment and suggest that spermatogenic defects may progress in a tubule-by-tubule manner.

      (2) Clarification of quantitative analysis and methodology (Reviewer #3):

      In response to concerns regarding the basis and methodology of cell quantification, we revised the Methods to provide detailed information on tissue preparation, fixation, orientation of the rete testis–Sertoli valve region, and the criteria used for quantitative analysis of SV-associated Sertoli cells (new Figure S4). We clarified that Sertoli cells were counted within the SV region extending approximately 100 μm from the RT boundary, including Sertoli cells protruding into the RT lumen, based on previously established criteria (Aiyama et al., 2015).

      (3) Clarification of the limitations of expression-level assessment (Reviewer #3):

      In response to concerns regarding the quantitative assessment of SOX17 and other SV-associated molecules, we clarified the technical limitations of selectively isolating the very small SV region and obtaining sufficient material for quantitative molecular analyses such as qPCR at the end of 7th paragraph in the Discussion. We therefore clarified that expression of SV-associated molecules in the present study was primarily evaluated using histological and immunohistochemical approaches and added the relevant text to acknowledge these limitations.

      We also made additional revisions to clarify each mouse Tg line, phenotypic descriptions, standardize gene nomenclature, improve methodological descriptions, and refine the relevant Discussion where appropriate.

      We sincerely appreciate the reviewers’ thoughtful and constructive comments. Their feedback has helped us clarify the scope of our conclusions, strengthen the methodological descriptions, and improve the overall presentation of the study. All changes have been incorporated into the revised manuscript.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (i) Although the current paper is not responsible for interpreting the 2022 findings, both datasets show reduced spermatid production accompanied by multinucleated giant germ-cell syncytia. This phenotype has been attributed to backflow of tubular fluid and consequent microenvironmental perturbation. While this is a reasonable hypothesis, it is not entirely consistent with earlier experimental observations. Complete ligation of the efferent ductules reliably produces giant cells, whereas estrogen-receptor knockout, which also causes massive luminal fluid accumulation, does not. In addition, ligation of the testicular artery itself can induce giant-cell formation. Although this may have already been answered in the papers, can you be sure that a direct or indirect effect on the vasculature can be excluded in the Sox17-cKO model?

      We thank the reviewer for this important comment. In our models, SOX17 expression was manipulated specifically in the Sertoli cell lineage, either by SF1-Cre-mediated Sox17 deletion or by ectopic SOX17 expression under the hAMH-promoter. SOX17-expressing vascular endothelial cells were not targeted in either model, making a direct effect of Sox17 manipulation on the testicular vasculature unlikely. Moreover, the partial rescue of the Sox17 cKO phenotype by hAMH-Sox17 supports the interpretation that the phenotype primarily results from altered SOX17 function in Sertoli cells and RT epithelia.

      However, indirect effects on the vascular or interstitial environment by aberrant luminal flow cannot be completely excluded, particularly with the substantial accumulation of sloughed round spermatids (giant cells) within the rete testis. Addressing the potential for an initial luminal flow defect, we newly added histological images of 4-week-old testes (Figure S1), where selective post-meiotic germ cell sloughing occurs despite preserved Sertoli cell process architecture, suggesting an altered adluminal microenvironment that impairs Sertoli–spermatid adhesion. Furthermore, low-magnification images of adult mature Sox17 cKO testes (Figure S1B) display a mosaic pattern of spermatogenic defects across individual tubules. While 3D reconstruction was not conducted, this structural pattern supports the view that spermatogenic failure progresses on a tubule-by-tubule basis, potentially linked to the structural integrity of individual Sertoli valves.

      (ii) A related and important unresolved issue is the total number of Sertoli cells per testis in cKO males. The number of Sertoli cells per tubule cross-section is reported to be equivalent to controls; however, the substantial reduction in testis weight implies a corresponding reduction in tubule length. Under these conditions, maintenance of a normal per-cross-section count would still be compatible with an overall decrease in total Sertoli-cell number. Although it is generally accepted that murine Sertoli cells exit the cell cycle around postnatal day 15, continued growth of the testis may still occur in the Sertoli valve region, where Sertoli cells retain proliferative capacity. Your discussion of possible heterogeneity in the embryonic origin of Sertoli cells near the rete testis is therefore particularly intriguing and commendable. Should this hypothesis be substantiated, it would raise the possibility that Sertoli cells derived from the valve region, especially those that migrate into the seminiferous tubules, are intrinsically less competent to support full spermatogenesis than those of classic gonadal-ridge origin.

      To help readers appreciate the overall severity and topographic distribution of the spermatogenic defect (particularly in tubule segments distant from the rete), inclusion of a low-magnification photomicrograph of a well-fixed (Bouin's) testicular cross-section would be very useful.

      We thank the reviewer for this important comment. We agree that the maintenance of Sertoli cell numbers per seminiferous tubule cross-section does not necessarily indicate preservation of the total Sertoli cell number per testis, particularly given the substantial reduction in testis size and potential reduction in overall seminiferous tubule length. Although total Sertoli cell numbers can theoretically be estimated using stereological approaches, such analyses are technically demanding and beyond the scope of the present study.

      We also appreciate the reviewer’s insightful suggestion regarding potential heterogeneity among Sertoli cell populations. Sertoli cells associated with the Sertoli valve region may have distinct developmental origins or functional properties compared with classical gonadal ridge-derived Sertoli cells, which could potentially influence their capacity to support complete spermatogenesis. Although this hypothesis was not directly tested in this study, we have expanded the Discussion to highlight the developmental and functional heterogeneity of Sertoli cell populations associated with the Sertoli valve as an important topic for future investigation.

      In addition, as requested, we have added a low-magnification image of well-preserved testicular cross-sections in Supplementary Figure S1B to better illustrate the overall severity and topographic distribution of spermatogenic defects throughout the testis.

      Specific Comments:

      (1) Figure 3A and associated fertility/histology data. The results state that epididymal spermatozoa were detected in only 2 of 7 cKO;Tg males at 8 weeks of age, yet Materials and Methods indicate that spermatogenesis was evaluated in only 5 males. a) Were the remaining two males also examined histologically? b) It would be interesting to determine if the severity of pathological changes was the same in regions more distant from the rete testis, or possibly different tubules. See: Nakata H, Wakayama T, Sonomura T, Honma S, Hatta T and Iseki S (2015). "Three-dimensional structure of seminiferous tubules in the adult mouse." J Anat 227(5): 686-694. c) In addition, mating trials were performed with four independent cKO;Tg males, two of which sired offspring. It is unclear whether the testes of these four mating males were included among the five (or seven) animals evaluated for histology, and whether the two fertile males correspond exactly to the two individuals that retained epididymal sperm. Please clarify these relationships explicitly so that readers can correctly interpret the link between histological findings and fertility.

      We thank the reviewer for this important comment. We apologize that the relationship among the groups of animals used for histological analysis, epididymal sperm detection, and fertility assessment was not sufficiently clear in the original manuscript. Because this study focused specifically on the anatomically minute RT–SV region, our sampling strategy had to prioritize the maximal utilization of this limited tissue. In this study, the RT–SV region, the remaining testicular tissue, and the epididymis were processed separately as three tissue blocks for each animal (Figure S4) and were independently evaluated for distinct analysis sets. Briefly, the proximal quarter containing the rete testis and Sertoli valve region was used for SV analysis, whereas the remaining three-quarters of the testis were used for evaluation of spermatogenesis, and the epididymis was analyzed separately for the presence of spermatozoa. Therefore, due to these technical requirements, tissue allocation, and independent analytical evaluation, the numbers of animals used for RT–SV analysis, testicular histology, epididymal sperm detection, and fertility testing were not identical.

      For quantitative histological analyses, we also used virgin males to minimize potential variation associated with mating experience and to allow comparison with age-matched littermate controls. Therefore, these animals were not used for fertility testing. Fertility assessment was performed using an independent cohort of cKO; Tg males that were subjected to long-term mating trials with wild-type females. Thus, fertility outcomes and histological findings were not designed to be directly matched at the individual level.

      In response to the reviewer’s suggestion, we have clarified the selection of experimental animals and the relationship among fertility assessment and histological analyses in the Materials and Methods and added a schematic illustration of the sampling strategy in Figure S4. We also corrected the citation for Nakata H et al., 2015 in the revised manuscript.

      (2) Page 8, line 301 (Sertoli-cell quantification). The description of the counting method-"counted in each ... (~100 μm from the edge of the RT; Fig. 4C)"-is ambiguous.

      (a) Does this mean that cells were counted beginning at the rete boundary and extending radially outward for approximately 100 μm, or is a circumferential sampling area intended? (b) Figure 4C shows a large standard deviation, indicating substantial variability with the current approach. An alternative strategy (for example, counting Sertoli cells within standardized areas or per tubule specifically within the valve region) might reduce variability and improve reproducibility. Regardless of the method ultimately chosen, a more precise, step-by-step description of the quantification protocol is required so that it can be reliably replicated by other laboratories.

      We thank the reviewer for pointing out that the description of the Sertoli cell quantification method was not sufficiently clear. The Sertoli cell quantification was performed using the same criteria as previously described (Aiyama et al., 2015; Uchida et al., 2022), in which SOX9-positive Sertoli cell nuclei within the SV-associated region were counted.

      In the revised manuscript, we have clarified that the SV region was operationally defined as comprising (i) the terminal 100 μm segment of the seminiferous tubule immediately adjacent to the rete testis (RT) and (ii) the protruded SV extending into the RT lumen. Based on the distribution of spermatogonial stem cells, the ~100 μm region extending from the RT boundary along the seminiferous tubule toward the ST side was defined as the SV region (Aiyama et al., 2015). Only sagittal sections showing a continuous RT–SV–ST axis and sectioning the SV approximately through its mid-sagittal plane were included for quantitative analysis.

      Furthermore, to improve reproducibility, we have added a more detailed description of the tissue preparation and quantification procedures in the Materials and Methods and provided a schematic illustration of the quantification strategy in the new Figure S4.

      Reviewer #2 (Recommendations for the authors):

      (1) Phenotypic differences between transgenic lines: the phenotypic differences between the tg26 and tg27 lines are intriguing and warrant further clarification. While tg27 mice exhibit infertility and defective spermatogenesis, tg26 animals remain fertile with SV expansion. Could the authors elaborate on the underlying causes of these differences? In particular, is infertility in tg27 mice due to excessive SOX17 expression impairing Sertoli cell function? A comparison of Sox17 expression levels between tg26 and tg27 lines would be informative. In addition, it would be useful to assess whether acetylated tubulin (Ac-Tub) expression is present in the Sertoli cells of the tg27 mouse testis.

      We thank the reviewer for this highly constructive and insightful comment. We clarified in the revised manuscript that the analysis of the Tg27 mouse was performed using the F0 founder male and added an explanation that only the Tg26 line could be established as its heterogenous SOX17 expression in Sertoli cells did not impair overall fertility. We agree that the phenotypic differences between the Tg26 line and the Tg27 mouse provide important clues regarding the dosage-dependent effects of SOX17 in Sertoli cells. Unfortunately, we were unable to establish a stable, multi-generational transgenic line from this Tg27 founder (F0) male. Consequently, we could not perform detailed molecular or immunohistochemical analyses on this line beyond the initial histological evaluation of the F0 generation presented in Figure 1. For this reason, we cannot provide a quantitative comparison of Sox17 expression levels or evaluate acetylated tubulin (Ac-Tub) expression in Tg27 Sertoli cells.

      To address the reviewer's concern without overstepping the available data, we removed direct quantitative comparisons of Sox17 expression levels between the two lines from the text. Instead, we added a clear description of their contrasting cellular expression patterns - specifically, the mosaic, heterogeneous SOX17 expression in Tg26 Sertoli cells versus the ectopic, uniform SOX17 expression in the infertile #27 F0 male - in the 'Animals' section of Materials and Methods. This mosaic pattern in Tg26 testes suggests the presence of Sertoli cells with low or undetectable SOX17 levels, which may be associated with sustaining overall fertility.

      (2) Mechanism of SOX17 action: although SOX17 is a transcription factor, the author's studies indicate it regulates SV formation via paracrine and/or autocrine signaling. The underlying mechanisms remain unclear. Which downstream factors mediate this effect? The observed upregulation of RSPO1 and WNT4 is suggestive, but more direct evidence would strengthen this conclusion. For example, does SV expansion in tg26 mice depend on the activation of RSPO1/WNT signaling? Additional molecular analyses, such as bulk RNA-seq comparing control and transgenic testes, could help identify pathways regulated by SOX17 and clarify its mode of action.

      We thank the reviewer for this important and insightful suggestion. At present, comprehensive analyses, including scRNA-seq of Sox17 cKO and littermate control testes, have not identified definitive downstream targets of SOX17 (Uchida et al., 2022). As the reviewer rightly points out, the Tg26 mouse model generated in this study represents a valuable tool for investigating SOX17-dependent molecular pathways. To this end, we are currently conducting transcriptomic analyses of Tg26 seminiferous tubules to identify genes altered in SOX17+ Sertoli cells. However, determining whether these candidate genes represent direct transcriptional targets of SOX17 and whether they function specifically in the rete testis-associated region during Sertoli valve formation will require extensive functional and expression studies. Therefore, we feel it would be premature to draw definitive conclusions regarding the underlying molecular mechanisms, including the precise involvement of the RSPO1/WNT signaling pathway, in the present manuscript. Accordingly, rather than overinterpreting the available data, we have revised the Discussion to clarify this limitation (at the end of 7th paragraph in the Discussion). Furthermore, incorporating initial insights from our ongoing Tg26 transcriptomic analyses, we have added a brief discussion, supported by relevant literature, on the possibility that SOX17 may regulate Sertoli valve formation by modulating cell adhesion and extracellular matrix (ECM) organization and altering the responsiveness of SOX17-positive Sertoli cells to morphogenetic signals originating from the rete testis (new 5th paragraph in Discussion).

      (3) Minor comment: Gene nomenclature should be standardized: e.g. line 245, Sox17 and hAMH should be italicized.

      We thank the reviewer for pointing this out. All gene names have been italicized throughout the manuscript.

      Reviewer #3 (Recommendations for the authors):

      No suggestions except to quantify some of the changes in concentration of agents by PCR rather than eyeball levels with immunocytochemistry. Verify the cell quantification procedure used.

      We thank the reviewer for this comment. The Sertoli valve (SV) is an extremely small transitional structure, with only approximately 20 sites per mouse testis. As a result, selective isolation of the SV region to collect sufficient material for molecular analyses, such as quantitative PCR, remains technically challenging. We have therefore added this limitation to the Discussion.

      Regarding the cell quantification procedure, we have clarified the methodology in the revised Materials and Methods and added a schematic illustration in Figure S4. Specifically, the Sertoli cell number in the SV region was quantified by counting SOX9-positive Sertoli cell nuclei within a standardized SV-associated region in RT–SV–ST sagittal sections.

    1. eLife Assessment

      The study investigates how CD1c-restricted T cells respond to Mtb-infected APCs, leading to increased cytokine production and cytotoxic activity that may help control Mtb infection. The work is important and will interest researchers in the field. The supporting evidence is solid.

    2. Reviewer #1 (Public review):

      Summary:

      T cells that recognize lipids-CD1c are frequent in circulation; however, their role in infection is unclear. This study aims to understand how Mtb infection can shape the responses of CD1c-specific T cells. CD1c is expressed in MTB granuloma, but in lower amount than in nearby inflamed tissue. Mtb infection downregulates the expression of CD1c on monocyte derived DCs. Single cell RNA sequencing revealed the cytotoxic program inherent to the lipids-CD1c-specific T cells. Using an in vitro APC system where CD1c expression remains intact upon Mtb infection, the authors suggest that these T cells react better to Mtb-infected than uninfected Cd1c-expressing APC and reduce Mtb burden in infected cells. Therefore, Cd1c downregulation could be an immune evasion strategy used by Mtb.

      Strengths:

      This study asks an important question. The single cell transcription analysis suggests the inherent cytotoxic program of lipid-CD1c cells and provides insights into their phenotypic and potential functional profiles. Function experiments suggest that these autoreactive T cells can react to Mtb infection, adding to the paradigm of infection control by these non-conventional T cell population. In summary, this study reports potential role of lipid-CD1c specific T cells in Mtb infection., and the evidence overall supports the conclusions.

      Weaknesses:

      (1) The study engineered THP-1 cells that lacked endogenous beta2 microglobulin (beta2m) and stably expressed chimeric CD1c-beta2m (CD1c-THP-1 cells). T cells isolated from healthy donors showed stimulation when mixed with CD1c-THP-1, suggesting the presence of CD1c reactive T cells in healthy individuals (Fig. 1). It appears that antibody staining profiles for different markers are not from the same experiment (Fig. 1). It is better to show all the representative antibody staining from a single experiment. And a quantification of MFI values from at least three independent experiments should be shown with statistics. Further, in each of the representative antibody profile, unstained/isotype control should be incorporated. The characterization of beta2m KO cells - using DNA sequencing and western blotting - should be shown in supplementary.

      (2) Next, using immunohistochemistry, the study suggests expression of CD1c in lung biopsy from TB patients, and that CD1c expression is reduced on primary MoDCs upon Mtb infection (Fig. 2).

      (3) To assess CD1c-reactive T cell cytotoxicity in Mtb infection, the authors generated CD1c-T cell lines (from T cells from healthy individuals) using two different approaches. In one, CD1c-specific T cells were expanded upon mixing with CD1c-THP-1 cells, were sorted using CD1c-endo tetramer, and expanded. In another approach, CD1c-specifc T cells were sorted (using CD1c-endo dextramer) directly from blood-derived T cells and were expanded. The line obtained using one of the approaches to assess cytotoxicity towards CD1c-THP-1 cells, uninfected or infected with Mtb. It is seen that CD1c-T cells exert cytotoxicity towards infected, and not towards uninfected, CD1c-THP-1 cells (Figs. 3,4).<br /> The authors should clearly indicate the T line, obtained from which one of the two approaches, was used for the cytotoxicity assay, and what was the result with the line obtained from other approach. Fig. S4 shows the sorting strategy and final profile of the T line obtained after CD1c-endo tetramer-based sorting and expansion. This figure is mistakenly referred as supplementary fig. 3 in the caption of Fig. 3; should be corrected. Further, to complete the Fig. S4, the intermediate step of sorting scheme, which shows negative and positive fraction as defined by streptamer binding, should be included. And Y-axis of left-most lower-most panel (no tetramer) should be labelled. The Profile of the T line shows three distinct population: tetramer-low, tetramer-intermediate, and tetramer-high. The authors should comment on this, particularly on the tetramerlow population that can be CD1c-negative cells that have non-specific weak binding to the tetramer.<br /> The profile of the T line obtained from the other approach should also be shown, and the experiments that used those should be clearly indicated.

      (4) To assess whether the elevated response of CD1c-T cells is mediated by the T cell receptor (TCR) that they harbour, the authors performed single T cell sequencing for TCR. They could identify 11 αβ pairs, two of them (named as EM1 and EM2) constituted 10 out of the 11 αβ pairs. It is surprising that such a shallow sampling (despite of the filtering mentioned in Fig. S9) such a high degree of clonal expansion and yielded two context-relevant TCR sequences. These TCRs were indeed DC1c specific, as indicated by their ability to upregulate CD69 upon engagement with CD1c (Fig. 5C). However, they respond very weakly to CD1c-THP-1 cells - the difference (THP-1 KO versus CD1c-THP-1 cells) is small for EM1, and the upregulation of CD69 per se is very weak (within the noise range) for EM2 (Fig. 5D). Nonetheless, the difference shown is statistically significant for both EM1 and EM2.

      (5) The study further indicate that the CD1c-specific T cells are enriched for the marker associated with their cytotoxic functions, and they perform slightly better in controlling Mtb growth in culture.

      The authors are suggested to carefully go through the manuscript a couple of times so that supplementary figures are referred appropriately and accurately.

    3. Reviewer #2 (Public review):

      Summary:

      The study by Milton et al titled "Human CD1c-autoreactive T cells recognise Mycobacterium tuberculosis-infected antigen-presenting cells and display cytotoxic effector programmes" characterises CD1c-restricted autoreactive T cells and their potential role in controlling Mtb infection. The authors develop a well-controlled system to assay for the functioning/activation of autoreactive T cells. They report the presence of CD1c-restricted autoreactive T cells in the circulating blood of healthy donors. They show that these T cells respond to CD1c and get activated even in the absence of any exogenous antigen. They next show that CD1c, along with CD1a and b, are typically downregulated on APCs during Mtb infection. These autoreactive T cells are cytotoxic, indicating they respond to Mtb treatment and/or to changes in the T cell ratio. The autoreactive T cells could effectively lyse Mtb-infected or PAMP-stimulated CD1c+APCs. Next, using TCR sequencing, they show that T cell responses were mediated by specific TCR clones with common sequence features. They show that these autoreactive T cells could curtail Mtb growth as measured by luminescence. Finally, using scRNAseq, they selectively identify the CD1c-reactive T cell pool and detect enrichment of typical effector memory CD4 and CD8 cells expressing cytolytic markers such as Granzyme, granulolysin, etc. The lung biopsy staining, along with the other data presented here, suggests that while CD1c-restricted T cells could have potential anti-bacterial roles, Mtb downregulation effectively shuts down this mechanism for TB control.

      Strengths:

      The study is designed well and has developed many exciting tools to generate specific information.

      Weaknesses:

      The revised manuscript addresses many concerns, but one section remains weak. The efficiency of these CD1c-restricted T cells in controlling TB remains very limited. The only result that addresses the bacterial control through this mechanism is Fig. 6C, which shows a very modest impact. Even THP1-KO cells show a decline in CFU when cultured with autoreactive CD1c-autoreactive T cells, and the further dip in THP1 CD1c cells is very minimal.

      Another issue left unaddressed is the cytolytic response on Mtb-infected cells. How efficient are lytic responses in controlling Mtb infection? Usually, bacteria can emerge from lysed cells and divide extracellularly. How would one show this mechanism in vivo?

    4. Reviewer #3 (Public review):

      The authors have addressed most concerns from the initial review, significantly enhancing the manuscript. The expanded characterisation of the engineered THP1-CD1c system provides strong evidence that the observed T-cell responses are unlikely to result from residual conventional MHC recognition. The specificity of the CD1c-autoreactive T-cell lines is further confirmed by multimer staining, their inactivity against THP1-KO cells, and TCR-transfer experiments.

      The revised data support the main conclusion that human CD1c-autoreactive T-cells recognise Mtb-infected cells that express Cd1c and exhibit cytotoxic effector functions. The difference between TCR-dependent recognition, shown by EM1 and EM2, and cytotoxicity, demonstrated by the original T-cell lines, is now more clearly presented.

      Several limitations remain. The specific CD1c-associated lipid signalling molecule that enhances recognition of Mtb-infected cells has yet to be identified. Additionally, the bacterial luminescence assay lacks validation against CFU counts and cannot differentiate between intracellular and extracellular bacteria. The single-cell RNA sequencing was conducted with only two donors, and the lung immunohistochemistry remains qualitative without comparison to healthy or non-TB inflammatory tissues. These constraints limit detailed mechanistic insights and broader applicability, but they are appropriately reflected in the manuscript.

      Overall, this is an important study that advances understanding of human CD1c-autoreactive T-cells in the context of mycobacterial infection. The evidence supporting the principal conclusions is solid, provided that altered CD1c-associated lipid presentation remains as a mechanistic hypothesis and the reduction in Mtb luminescence is not taken as direct evidence of selective intracellular bacterial killing.

    5. Author response:

      The following is the authors’ response to the original reviews.

      We have carefully considered all comments and have revised the manuscript to address the key points raised. We have also updated the author list to include Kinga Niedobecka, who performed the additional flow cytometric validation of the engineered THP1 cell lines included in the revised manuscript. In particular, we have strengthened the validation of the THP1-CD1c system, clarified and better signposted the characterisation of CD1c-autoreactive T-cells using existing data, and refined the explanation of the mechanisms underlying enhanced responses to Mtb-infected cells. Some of the suggestions represent significant additional experimental work beyond the scope of this manuscript, and in these instances we have amended the text to clarify interpretation and limitations.

      eLife Assessment

      The study investigates how CD1c-restricted T cells respond to Mtb-infected APCs, leading to increased cytokine production and cytotoxic activity that may help control Mtb infection. While the work is important and will interest researchers in the field, the supporting evidence is incomplete and could be strengthened by additional experiments. Experiments would: (i) evaluate THP1-CD1c cells to determine whether MHC surface expression is reduced or entirely abolished, (ii) enhance confidence in the purity of the CD1c-specific T cell population isolated from blood, and (iii) suggest what additional signal THP1-CD1c cells treated with Mtb express that is absent from the untreated cells.

      (i) evaluate THP1-CD1c cells to determine whether MHC surface expression is reduced or entirely abolished

      We thank the Editor for highlighting this important point. We agree that it is essential to establish whether conventional MHC-mediated antigen presentation could contribute to the observed T-cell responses. To address this directly, we repeated and extended our flow cytometric validation of the engineered THP1 system. These data are now presented in an expanded Fig. 1A and include assessment of CD1c, classical MHC class I, MHC class II, β2m, CD1b and HLA-E across WT THP1, THP1-KO and THP1-CD1c cells. Our THP1-KO system is based on CRISPR-mediated knockout of both β2microglobulin (β2m) and the Class II transactivator (CIITA). Loss of β2m removes surface expression of β2m-dependent molecules, including classical MHC class I and endogenous CD1 proteins, while CIITA knockout prevents MHC class II expression. In this new analysis, WT THP1 cells expressed β2m and classical MHC class I, with low detectable MHC class II and HLA-E. In contrast, THP1-KO cells lacked detectable β2m, MHC class I, MHC class II, HLA-E, CD1b and CD1c. Importantly, THP1-CD1c cells retained robust CD1c expression through the CD1c-β2m fusion construct, while MHC class I, MHC class II, CD1b and HLA-E remained undetectable by flow cytometry.

      These extended validation data support the conclusion that residual MHC expression does not account for the observed T-cell responses, which are instead dependent on CD1c expression. We have revised the relevant section of the Results to incorporate these data and to clarify that the engineered THP1-CD1c APC system provides robust CD1c expression in the absence of detectable surface MHC-I or MHC-II (revised manuscript, page 5-6, lines 111-122; Fig. 1A and Fig. 1 legend).

      (ii) enhance confidence in the purity of the CD1c-specific T-cell population isolated from blood

      We agree that confidence in the specificity and purity of the CD1cautoreactive T-cell populations is essential. The relevant data were included in the original manuscript, but we recognise that they were not signposted clearly enough. We have therefore revised the Results to describe the enrichment, sorting, post-expansion validation and functional specificity of the T-cell lines more explicitly on page 8, lines 176-191.

      CD1c-autoreactive T-cell lines were generated from two independent donors using two complementary strategies. One line was generated by expansion with THP1-CD1c APCs followed by CD1c-endo tetramer-guided sorting and expansion. A second line was generated by direct enrichment using CD1c-endo streptamers, followed by CD1c-endo dextramer sorting and expansion. The gating strategy and post-sort validation are shown in Fig. S4. Importantly, after expansion, the enriched cells stained strongly with CD1c-endo tetramers, whereas unstained and irrelevant tetramer controls showed no detectable staining. In the main figure, both donor-derived lines are shown to be strongly CD1c-endo tetramer-positive, with post-expansion tetramer positivity of 97100% (Fig. 3A and 3C). Both lines were αβTCR+CD4+ and lacked detectable γδTCR or CD8 expression (Fig. 3B and 3D).

      We also highlight the functional validation of specificity. Both T-cell lines were activated by THP1-CD1c APCs but not THP1-KO APCs, as assessed by CD69 and CD25 upregulation (Fig. 3E). Importantly, the specificity of these cells was further supported by TCR transfer experiments. TCRs cloned from one of the CD1c-endo tetramer-positive T-cell lines were expressed in Jurkat reporter cells and conferred CD1c-endo tetramer binding, activation in response to plate-bound CD1c-endo protein, and enhanced activation in response to Mtb-infected THP1-CD1c APCs (Fig. 5B-D, page 10, lines 225244). This provides independent confirmation that the enriched T-cell line contained CD1c-reactive TCRs capable of mediating CD1c-dependent recognition.

      Together, these data support that the T-cell populations used in the functional assays are highly enriched CD1c-specific T-cell lines rather than mixed or nonspecific populations.

      (iii) suggest what additional signal THP1-CD1c cells treated with Mtb express that is absent from the untreated cells.

      We agree that identifying the additional signal provided by Mtb-treated THP1-CD1c cells is an important mechanistic question. We have now revised the Discussion to clarify our interpretation and to more explicitly outline the likely mechanisms (revised manuscript, page 16-17, lines 380-400).

      Our data suggest that the enhanced response to Mtb-infected THP1-CD1c cells is unlikely to be explained simply by increased CD1c expression, generic APC activation, or soluble cytokine release. CD1c expression was maintained but not increased on THP1-CD1c cells after Mtb infection, and stimulation with TLR2 or TLR4 agonists did not reproduce the enhanced cytotoxicity observed after Mtb infection. In addition, Mtb-treated THP1-CD1c cells alone produced IL-8 and RANTES, but not the broader cytokine profile observed in T-cell co-cultures. Together, these data suggest that Mtb exposure provides an additional CD1c-dependent activating signal.

      We now discuss that this signal is most likely an altered CD1c-presented lipid repertoire on Mtb-exposed APCs. Possible mechanisms include presentation of Mtb-derived lipids, infection-induced accumulation of host-derived stimulatory “stress lipids”, presentation of bacterial and mammalian shared lipids, or altered lipid processing and trafficking during infection. These possibilities are consistent with prior studies showing enhanced responses of autoreactive CD1-restricted T-cells to microbial stimulation and our TCR transfer experiments seemingly support a CD1c-TCR-dependent recognition mechanism. However, because we have not directly identified the lipid ligands presented by CD1c on Mtb-infected APCs, we now state this as a mechanistic hypothesis rather than a conclusion, and a key outstanding question.

      We have revised the Discussion (page 16-17, lines 380-410) to make this limitation explicit. Future studies will require isolation of CD1c molecules from Mtb-infected cells and then lipidomic analysis and mass spectrometry to define the CD1c-associated lipid species.

      Reviewer #1 (Public review):

      Strengths:

      (1) This study asks an important question. The single-cell transcription analysis suggests the inherent cytotoxic program of lipid-CD1c cells and provides insights into their phenotypic and potential functional profiles. Function experiments suggest that these autoreactive T-cells can react to Mtb infection, adding to the paradigm of infection control by these non-conventional T-cell populations.

      We thank the reviewer for this positive assessment of the importance of the study and for recognising the value of the single-cell transcriptional analysis and functional experiments. We are pleased that the reviewer agrees that our findings provide insight into the cytotoxic effector programme of CD1c-autoreactive T-cells and their potential contribution to immune responses during Mtb infection.

      Weaknesses:

      (2) The study lacks sufficient rigor; conclusions may be strengthened with the incorporation of more controls, and some deeper characterization of the THP1 system and the CD1c-specific T-cells isolated from blood. Crucial conclusions are drawn from the cell mixing experiments involving the engineered THP-1 system and CD1c-lipidspecific T-cells from blood. These cells need more in-depth characterization. The expression of MHC-I/II is clearly reduced in THP1-CD1c cells. However, it is important to ensure that it is completely abolished, since a residual expression can skew the result with activation of conventional T-cells in the blood or low levels of conventional T-cells that may be present in the CD1c-tetra/multimer sorted T-cells

      We agree that this is an important point and have addressed it by adding new experimental controls and by clarifying the validation of the CD1c-autoreactive T-cell lines.

      First, we repeated flow cytometric validation of the existing markers and extended the panel to assess additional surface molecules across WT THP1, THP1-KO and THP1CD1c cells. The revised Fig. 1A therefore includes repeat staining for β2m, classical MHC class I, MHC class II and CD1c, together with newly added staining for HLA-E and CD1b.

      The THP1-KO system is based on CRISPR-mediated knockout of both β2-microglobulin (β2m) and the Class II transactivator (CIITA). Loss of β2m removes surface expression of β2m-dependent molecules, including classical MHC class I and endogenous CD1 proteins, while CIITA knockout prevents MHC class II expression. The repeated analyses confirmed the original staining pattern, while the additional HLA-E and CD1b stains further extended validation of the system. WT THP1 cells expressed β2m and classical MHC class I, with low detectable MHC class II and HLA-E. In contrast, THP1-KO cells lacked detectable β2m, MHC class I, MHC class II, HLA-E, CD1b and CD1c.

      Importantly, THP1-CD1c cells retained robust CD1c expression through the CD1c-β2m fusion construct, while MHC class I, MHC class II, HLA-E and CD1b remained undetectable by flow cytometry. We have revised the Results to describe these new validation experiments more clearly (revised manuscript, page 5-6, lines 111-122, Fig. 1A and Fig. 1 legend).

      Second, we have strengthened the description of the purity and specificity of the CD1c-autoreactive T-cell lines. These lines were generated using two complementary approaches, namely expansion with THP1-CD1c APCs followed by CD1c-endo tetramer-guided sorting, and direct enrichment using CD1c-endo streptamers followed by CD1c-endo dextramer sorting and expansion. The gating strategy and post-sort validation are shown in Fig. S4. After expansion, the enriched cells stained strongly with CD1c-endo tetramers, whereas unstained and irrelevant tetramer controls showed no detectable staining. Both donor-derived lines were strongly CD1c-endo tetramer positive, with post-expansion tetramer positivity of 97 to 100%, and both were αβTCR+CD4+ with no detectable γδTCR or CD8 expression (Fig. 3A-D). Functionally, both lines responded to THP1-CD1c APCs but not parental THP1-KO APCs, as assessed by CD69 and CD25 upregulation (Fig. 3E). We have revised the Results to signpost these data more clearly (revised manuscript, page 8, lines 176–191).

      Finally, TCR transfer experiments provide independent confirmation of CD1c-specific recognition. TCRs cloned from one of the CD1c-endo tetramer-positive T-cell lines conferred CD1c-endo tetramer binding and CD1c-dependent activation when expressed in Jurkat reporter cells (Fig. 5B-D, revised manuscript, page 10, lines 225244). Together, the absence of detectable MHC-I/MHC-II expression in the engineered APC system, the high CD1c-endo tetramer enrichment of the T-cell lines, the lack of activation against THP1-KO cells, and the TCR transfer experiments support the conclusion that the observed responses are driven by CD1c-dependent recognition rather than residual conventional MHC-mediated activation.

      (3) Figure 2: The immunohistochemistry appears to be shown only for one biopsy; it may be worth quantifying the immunohistochemistry of all five.

      We thank the reviewer for this helpful suggestion. We agree that quantitative analysis of CD1c immunohistochemistry across all biopsies would be valuable. We examined CD1c staining across all five TB lung biopsies and observed a consistent spatial pattern, with CD1c staining generally low or infrequent in central granulomatous regions and more apparent in distal inflammatory tissue and lymphoid/B-cell follicle-rich areas.

      However, because these were diagnostic human biopsy samples with substantial variation in tissue size, architecture, granuloma representation and inflammatory composition, we do not think that simple bulk quantification of CD1c-positive area across biopsies would be robust or biologically interpretable. In particular, quantification would be strongly affected by whether a section captured granuloma centre, peripheral inflammatory regions, lymphoid aggregates, or uninvolved lung tissue. We have therefore retained the IHC as representative spatial evidence of CD1c expression in TB lung tissue, rather than presenting it as a quantitative comparison across anatomical compartments.

      We have revised the Results to make this clearer, stating that CD1c expression was observed across the biopsies analysed but was spatially heterogeneous, with staining most apparent away from the granuloma centre and in lymphoid/inflammatory regions (revised manuscript, page 7, lines 144-151). We have also tempered the interpretation in the Discussion to avoid overstatement and now highlight systematic quantitative spatial analysis of larger tissue cohorts as an important future direction (revised manuscript, page 18, lines 427-430).

      (4) The expression of CD1 molecules goes up during the differentiation of MoDC, and Mtb infection prevents or dampens the upregulation. Does Mtb infection downregulate the CD1 expression of mature DCs? Can the effect of Mtb on the expression of CD1a,b,c molecules be investigated using CD1c-expressing DCs from blood? What could be the reason THP-1 cells do not downregulate CD1 molecules upon Mtb infection, and how about the expression of CD1a and b?

      We agree that the distinction between impaired CD1 upregulation during MoDC differentiation and active downregulation of CD1 expression on already differentiated CD1-expressing DCs is important.

      In the revised manuscript, we have clarified that our primary cell data assess the effect of Mtb infection on differentiated MoDCs that already express CD1 molecules, rather than only examining failure of CD1 induction during differentiation. Specifically, we analysed a published RNA-sequencing dataset from differentiated human MoDCs infected with live Mtb and observed reduced expression of group 1 CD1 genes, including CD1A, CD1B and CD1C, at 48 hours after infection. We then validated this experimentally at the protein level by flow cytometry, showing reduced CD1c expression on primary MoDCs after live Mtb infection. These data support the conclusion that Mtb infection can reduce CD1 expression on CD1c-expressing primary DCs.

      We agree that analysis of freshly isolated blood CD1c+ DCs would be valuable. However, these cells are rare in peripheral blood and are technically challenging to isolate in sufficient numbers for live Mtb infection assays and downstream flow cytometric or functional analysis. For this reason, we used MoDCs as a tractable primary human DC model to assess infection-induced changes in CD1 expression. We now acknowledge in the revised Discussion that validation in primary blood-derived CD1c+ DCs would be an important future direction.

      We have also clarified why CD1c expression is not downregulated in the engineered THP1-CD1c system. In primary DCs, Mtb-mediated suppression of CD1c has been linked to host regulatory mechanisms, including post-transcriptional regulation by miRNAs such as miR-381-3p, which targets the 3′ UTR of endogenous CD1c transcripts. In contrast, CD1c expression in our THP1-CD1c cells is driven by a lentiviral CD1c-β2m fusion construct under a heterologous promoter and expressed from a cDNA lacking the native untranslated regions. Therefore, CD1c in this system is not expected to be regulated in the same way as endogenous CD1c in primary DCs.

      This is a deliberate feature of the model. It allows us to assess CD1c-dependent T-cell responses to Mtb-infected APCs without the confounding effect of infection-induced CD1c loss. THP1-CD1c cells do not express endogenous CD1a or CD1b because the parental THP1-KO cells lack β2m-dependent endogenous CD1 surface expression, and only CD1c is reintroduced through the CD1c-β2m fusion construct. We have revised the Results and Discussion to clarify these points (revised manuscript, pages 7- 8, lines 165174 and pages 17- 18, lines 411-430).

      (5) Figure 3: (F) What does the X-axis read for the no infection group? The value for MOI = 0 should be incorporated for the infected T-cell group.

      We agree that the original presentation could be clearer. The uninfected condition corresponds to MOI = 0, whereas the remaining points represent THP1-CD1c APCs exposed to increasing amounts of UV-killed Mtb. We have retained the figure layout but revised the figure legend to clarify that the MOI values on the x-axis apply only to the Mtb-treated conditions, and that the no-infection/no-treatment control represents MOI = 0 (Fig. 3 legend).

      (6) Figure 4: In the lysis assay, THP1-CD1c cells (uninfected and infected) incubated alone should be incorporated.

      We agree that APC-only controls are essential for interpreting the lysis assay, and we apologise that this was not sufficiently clear in the original manuscript. THP1-CD1c cells cultured alone, both uninfected and Mtb-infected, were included in all assays and used to define baseline target-cell viability for each matched condition.

      The data in Fig. 4 are presented as specific lysis to isolate the effect of T-cells on target cell viability. Specifically, THP1 viability in APC-only wells was used as the baseline and subtracted from the corresponding T-cell co-culture condition within the same experiment. Thus, lysis of uninfected THP1-CD1c cells was calculated relative to uninfected THP1-CD1c cells cultured alone, and lysis of Mtb-infected THP1-CD1c cells was calculated relative to Mtb-infected THP1-CD1c cells cultured alone. This presentation allows the T-cell-mediated effect to be visualised while accounting for baseline viability differences.

      We have revised the Methods and Fig. 4 legend to make this calculation more explicit (revised manuscript, page 25, lines 608-614; Fig. 4 legend).

      (7) A quantitative brief on the single cell TCR sequencing - including how many T-cells were sequenced and the frequency of different clone including EM1 and EM2 - should be shown.

      We agree that the single-cell TCR sequencing data required clearer quantitative description. We have expanded the Results and Fig. 5 legend to include the number of single cells analysed and the frequency of the dominant clonotypes. Single CD1c-endo tetramer-positive T-cells were sorted into individual wells for targeted TCR sequencing. After filtering and manual curation, 11 single cells yielded productive paired αβ TCR sequences. The repertoire was oligoclonal, with two dominant productive clonotypes accounting for 10 of 11 paired TCRs. EM1 was detected in 6 of 11 cells and EM2 was detected in 4 of 11 cells. These data support the selection of EM1 and EM2 for TCR-transfer experiments and clarify that they were dominant clonotypes within the CD1c-endo tetramer-positive T-cell line rather than arbitrarily selected TCRs. We have revised the Results and Fig. 5 legend accordingly (revised manuscript, page 10, lines 225-235, Fig. 5 legend).

      Reviewer #1 (Recommendations for the authors):

      (8) Perform an experiment to assess activation of T-cells expressing EM1 or EM2, upon mixing with CD1c-expressing dendritic cells isolated from human blood, with and without Mtb infection.

      We agree that testing EM1 and EM2 TCRs against primary dendritic cells is an important question. However, in the specific context of Mtb infection, the proposed experiment is difficult to interpret because Mtb downregulates CD1c expression on primary dendritic cells. This is supported by previous studies showing that Mtb and BCG suppress CD1c expression on DCs [1,2], and by our own data showing reduced CD1 group 1 transcript expression in Mtb-infected MoDCs and reduced CD1c protein expression on primary MoDCs following Mtb infection (Fig. 2B and 2C). Therefore, mixing EM1- or EM2-expressing Jurkat T-cells with Mtb-infected primary CD1c-expressing DCs would introduce a major confounder: reduced T-cell activation could reflect loss of CD1c expression rather than absence of a CD1c-dependent Mtb-induced activating signal. This is precisely why we used the engineered THP1-CD1c system, in which CD1c expression is preserved during Mtb infection (Fig. 2D). This model allowed us to test whether Mtb infection enhances CD1c-TCR-dependent activation without the confounding effect of infection-induced CD1c loss.

      Using this controlled system, we show that EM1 and EM2 TCRs confer CD1c-endo tetramer binding, activation in response to plate-bound CD1c-endo protein, activation in response to THP1-CD1c but not THP1-KO APCs, and enhanced activation in response to Mtb-infected THP1-CD1c APCs (Fig. 5B-D). These data support the conclusion that the enhanced response to Mtb-infected APCs is mediated through CD1c recognition by the TCR.

      We have revised the Results and Discussion to clarify this rationale and to explain why the engineered THP1-CD1c system was necessary for these experiments (revised manuscript, page 10, lines 242-244; page 17-18, lines 411-430).

      (9) Conduct an experiment to assess T-cell cytotoxicity expressing EM1 or EM2, in the presence and absence of Mtb infection.

      We agree this would be a valuable experiment. As outlined in our response to 8, EM1 and EM2 were cloned into Jurkat T-cells to test TCR-dependent CD1c recognition and activation, not cytotoxic effector function. Jurkat T-cells are not cytotoxic effector cells<sup>3</sup>, so they are not suitable for target-cell killing assays.

      Cytotoxicity was instead assessed using the original CD1c-autoreactive T-cell lines (Figs. 3F-G and 4D-E). Testing EM1- or EM2-mediated killing would require engineering and validating primary human T-cells expressing these TCRs, which is a substantial additional workflow. We have clarified in the revised manuscript that the EM1/EM2 experiments demonstrate TCR-dependent recognition, while cytotoxicity was assessed using the CD1c-autoreactive T-cell lines (revised manuscript, page 10, lines 234–244).

      (10) A list of primers used for TCR sequencing should be provided.

      We have now provided the primer sequences used for targeted single-cell TCR sequencing in a new supplementary table (Table S1). We have also revised the Methods to provide additional detail on the single-cell TCR sequencing workflow, including CD1c-endo tetramer-guided single-cell sorting, oligo-dT reverse transcription, universal cDNA amplification, targeted amplification of TCR variable regions using TRAC-, TRBC-, TRGC- and TRDC-specific primers, well-specific 8-bp barcoding, size selection, library preparation and MiSeq sequencing. In addition, we now cite the SMART-seq2 protocol on which the approach was based (Picelli et al., 2014) (revised manuscript, page 20, lines 491-503; new Table S1).

      Reviewer #2 (Public review):

      Strengths:

      (1) The study is designed well and has developed many exciting tools to generate specific information.

      We thank the reviewer for this positive assessment of the study design and for recognising the value of the experimental tools developed in this work.

      Weaknesses:

      (2) The study has weaknesses in two important parameters - novelty and relevance in controlling TB. Further, the results could be better presented and discussed to allow easy understanding of the experimental design

      We accept that the novelty and relevance to TB control could be made clearer in the manuscript. However, we believe the study makes several important and previously unreported contributions, and we have revised the Introduction, Results and Discussion to improve the clarity of the experimental design and to state the conceptual advance more explicitly.

      First, to our knowledge, this is the first study to demonstrate that human CD1c-autoreactive T-cells respond more strongly to Mtb-infected CD1c+ APCs than to uninfected CD1c+ APCs. Previous work has shown that CD1c-autoreactive T-cells exist in human blood and can respond to CD1c-expressing cells in the absence of exogenous antigen. However, their role during infection has remained unclear. Our findings extend the field beyond the established steady-state, autoimmune and tumour contexts of CD1c autoreactivity by identifying Mtb-infected APCs as a biologically relevant setting in which these cells acquire enhanced effector activity. We propose that CD1cautoreactive T cells may not simply represent autoreactive bystanders that become pathogenic in disease, but instead form an evolutionarily conserved arm of lipid immune surveillance that can detect infection-associated changes in antigen presentation. Given the long-standing selective pressure imposed by microbial infection throughout human evolution, it is plausible that protection against infection represents a central physiological function of these cells, with their roles in autoimmunity and cancer reflecting the same capacity to sense altered self-lipid landscapes in other settings. Our data provide initial functional evidence supporting this model.

      Second, the study provides functional evidence that these cells are not simply activated by infected APCs, but can mediate effector functions relevant to antimicrobial immunity. CD1c-autoreactive T-cells showed enhanced activation, cytokine production and cytotoxicity in response to Mtb-infected APCs, and led to reduced Mtb burden under in vitro conditions. These findings are directly relevant to TB immunity because cytotoxic T-cell pathways and antimicrobial molecules such as granulysin have been implicated in control of intracellular Mtb.

      Third, the study links these functional observations to the ex vivo biology of human CD1c-autoreactive T-cells. Single-cell transcriptomic profiling demonstrates that these cells are enriched for cytotoxic effector-memory programmes and express molecules associated with target-cell killing and antimicrobial activity. This provides an independent, unbiased cellular basis for the functional assays and strengthens the conclusion that CD1c-autoreactive T-cells represent a plausible effector population in anti-mycobacterial immunity.

      We agree that the experimental design needed clearer presentation. In the revised manuscript, we have improved signposting of the stepwise logic of the study: (1) defining and validating the THP1-CD1c APC system, (2) demonstrating CD1cautoreactive T-cell enrichment and specificity, (3) testing responses to UV-killed and live Mtb, (4) confirming TCR-dependent CD1c recognition using EM1 and EM2 TCR transfer, and (5) integrating these functional data with single-cell transcriptomic profiling of ex vivo CD1c-autoreactive T-cells. These revisions aim to make the experimental design easier to follow and to clarify how each section supports the overall conclusion.

      We have revised the Introduction and Discussion accordingly to more clearly state the novelty and TB relevance of the work (revised manuscript, page 4-5, lines 83-104; page 14, lines 323-329).

      (3) At several places, UV-killed or live Mtb were used. What is the rationale behind that?

      We have now added a Methods statement explaining that UV-killed Mtb was used for controlled exposure to defined amounts of Mtb-derived antigen, particularly in dose-response cytotoxicity and cytokine-release assays, whereas live Mtb was used to assess T-cell activation, target-cell lysis and relative bacterial burden during APC infection with proliferating Mtb. We have also ensured that the figure legends clearly specify whether UV-killed or live Mtb was used in each experiment (revised manuscript, page 22, lines 534-540).

      (4) Why use irradiated THP1-CD1c cells for activating T-cells?

      Irradiated THP1-CD1c cells were used only during the T-cell expansion phase to provide sustained CD1c-mediated stimulation while preventing proliferation of the THP1 APCs. This was necessary because the expansion cultures lasted up to 12 days, during which non-irradiated THP1 cells would continue to divide and could overgrow the T-cell culture. Irradiation therefore allowed THP1-CD1c cells to function as APCs while maintaining controlled culture conditions and enabling selective expansion of CD1c-reactive T-cells. We have clarified this rationale in the Methods (revised manuscript, page 19-20, lines 472-474).

      (5) While functional assays identified only CD4+ cells as CD1c-restricted, scRNAseq shows that both CD4+ and CD8+ cells exhibit this phenotype

      We agree with the reviewer’s observation and have clarified this point in the revised Discussion. The functional assays were performed using CD1c-autoreactive Tcell lines generated from two donors. Both lines were CD4+αβTCR+, reflecting the outcome of the enrichment, sorting and expansion process used to generate sufficient T-cells for functional assays. These lines therefore provide mechanistic evidence that CD4+ CD1c-autoreactive T-cells can recognise CD1c+ APCs and respond more strongly to Mtb-infected APCs, but they are not intended to represent the full diversity of the CD1c-autoreactive T-cell compartment.

      By contrast, the single-cell RNA-seq analysis was designed to provide a broader ex vivo assessment of CD1c-endo-binding T-cells without relying on prolonged in vitro expansion. This revealed that CD1c-autoreactive T-cells include both CD4+ and CD8+ populations, with enrichment of cytotoxic effector-memory programmes. We therefore interpret the functional and single-cell datasets as complementary: the functional assays provide mechanistic validation using tractable CD1c-reactive T-cell lines, while the single-cell data demonstrate that the broader ex vivo CD1c-autoreactive compartment is phenotypically diverse and includes both CD4+ and CD8+ cytotoxic populations.

      We have revised the Discussion to clarify this point and to emphasise that combining in vitro functional assays with ex vivo single-cell profiling allowed us to capture both mechanistic activity and broader cellular diversity (revised manuscript, page 14-15, lines 341-349).

      (6) Identifying the specific lipid antigen presented by CD1c could add greater value to the study.

      We agree that identifying the specific CD1c-presented lipid antigen(s) would add important mechanistic insight. As noted in the response to the Editor, our data suggest that the enhanced response to Mtb-infected THP1-CD1c APCs is most likely due to altered CD1c-associated lipid presentation. However, defining these lipid species would require isolation of CD1c from infected APCs followed by specialised mass spectrometry-based lipidomics, which is a substantial additional workflow. We have revised the Discussion to state this limitation clearly and to highlight lipid identification as an important next step (revised manuscript, page 16-17, lines 380-410).

      (7) Since autoreactivity was independent of exogenous antigen, the cytotoxic activity should also be independent of exogeneous antigens? What additional signal a THP1-CD1c cells treated with UV-killed Mtb express that is absent from the untreated cells?

      CD1c-autoreactive T-cells likely recognise self-lipids presented by CD1c, but our data show that this response is enhanced after Mtb exposure. We interpret this as evidence that Mtb alters the quality or abundance of CD1c-associated lipid ligands, potentially through Mtb-derived lipids or infection-induced changes in host lipid metabolism. We have revised the Discussion to clarify that the precise lipid ligand(s) remain unidentified and will require future CD1c-lipidomic analysis (revised manuscript, page 16-17, lines 380-410).

      (8) The relative Mtb growth assay is confusing. CD1c cells with Mtb infection triggers massive lytic response, as shown in Figure 4. Under similar conditions, in Figure 6, the authors report a significant decline in Mtb growth in these cells. The problem is that with the kind of lytic response observed, a lot more Mtb could be present extracellularly and would evade killing. How do we reconcile the two observations?

      In the Mtb lux assay, extracellular bacteria were removed by washing after the initial infection step, so the starting bacterial population measured in the co-culture assay is expected to be predominantly cell-associated. We also recognise that the luminescence readout measures total viable lux-expressing Mtb under the assay conditions and does not distinguish intracellular from extracellular bacteria at later time points.

      The cytotoxicity observed in Fig. 4 reflects enhanced but incomplete lysis of infected APCs. Therefore, although T-cell-mediated lysis could release some bacteria from infected target T-cells, the reduced luminescence observed in Fig. 6C indicates a lower net viable Mtb burden under these co-culture conditions. We interpret this as the combined outcome of CD1c-autoreactive T-cell effector activity, including cytotoxicity and antimicrobial mediators such as granulysin and cytokines, rather than as direct evidence of selective intracellular bacterial killing.

      We have revised the Results and Methods to clarify that the Mtb lux assay measures relative viable bacterial burden/luminescence under in vitro co-culture conditions. We have changed the discussion to avoid over-interpreting this assay as distinguishing intracellular from extracellular Mtb killing (revised manuscript, page 11-12, lines 266274, page 26, lines 633-637).

      Reviewer #2 (Recommendations for the authors):

      (9) Nearly 40-50% of the samples did not respond to THP1-CD1 stimulation. What contributes to this diversity?

      We agree that there is clear donor-to-donor variability in the response to THP1-CD1c stimulation [4]. Approximately one-third of donors did not show detectable expansion under these assay conditions. This likely reflects differences in the precursor frequency and TCR repertoire composition of CD1c-autoreactive T-cells between donors, together with variation in activation state and responsiveness during short-term in vitro expansion. Apparent non-response may also reflect low-frequency CD1c-reactive populations that are present but fall below the detection threshold. We have revised the Discussion to acknowledge donor heterogeneity as an expected feature of primary human CD1c-autoreactive T-cell responses (revised manuscript, page 14-15, lines 341-349).

      (10) For lung biopsy staining, how is CD1 expression in healthy tissue or some unrelated inflammatory condition?

      The purpose of the lung biopsy staining was to determine whether CD1c-expressing cells are present in human TB lung tissue and to assess their spatial relationship to granulomatous inflammation, rather than to perform a formal comparison between healthy, non-TB inflammatory and TB lung tissue.

      Across the TB biopsies analysed, CD1c staining was spatially heterogeneous. CD1c expression was generally low or infrequent in central granuloma regions and in tissue regions remote from granulomatous inflammation, whereas staining was more apparent in distal inflammatory tissue and lymphoid/B-cell follicle-rich regions. We have revised the Results to clarify that these data are presented as representative spatial observations within TB lung tissue, rather than as a quantitative comparison with healthy or unrelated inflammatory tissue.

      We agree that comparison with healthy lung and non-TB inflammatory lung tissue would provide useful additional context, particularly for distinguishing TB-associated changes from more general inflammatory induction of CD1c. We now acknowledge this as an important future direction (revised manuscript, page 18, lines 425-430). We have also revised the Results to clarify the spatial pattern of CD1c staining within TB lung tissue (revised manuscript, page 7, lines 144-151).

      (11) What was the rationale for using UV-killed or live Mtb for different experiments?

      This point is addressed in our response to 3 above.

      Reviewer #3 (Public review):

      Strengths:

      (1) The manuscript is well written, and the novelty, impact, and limitations of this study are precisely highlighted by the authors.

      We thank the reviewer for this positive assessment of the manuscript, particularly their recognition of the study’s novelty, impact and balanced discussion of its limitations.

      (2) Lipid antigen identification and direct lipid identification via lipidomics/MS of CD1c-bound lipids from Mtb-infected APCs would clarify whether the enhancement arises from altered self-lipids or subtle Mtb lipids

      We agree that direct identification of CD1c-bound lipids from Mtb-infected APCs would provide important mechanistic insight and help determine whether enhanced activation reflects altered self-lipids, Mtb-derived lipids, or shared lipid species. As noted above, this would require isolation of CD1c from infected APCs followed by specialised mass spectrometry-based lipidomic analysis, which represents a substantial additional workflow. We have revised the Discussion to state this limitation clearly and to highlight CD1c-lipidomic analysis as an important next step (revised manuscript, page 16-17, lines 380-410).

      Reviewer #3 (Recommendations for the authors):

      (3) Figure 2Ai-vi, lines 134-136. The authors should include the data from central granuloma staining to solidify their claim of the presence of CD1c expression remote from the centre of TB granulomas.

      Central granuloma regions are included in Fig. 2A, including panels showing staining within granulomatous tissue where CD1c expression is low or infrequent compared with distal inflammatory and lymphoid/B-cell follicle-rich regions. We agree that this spatial distinction was not sufficiently clear in the original text and figure legend.

      We have therefore revised the Results and Fig. 2 legend to more explicitly guide the reader through the central versus distal regions shown in Fig. 2A. The revised text now states that CD1c expression was observed across lung biopsies from all five TB patients, but was spatially heterogeneous, with staining most apparent in distal inflammatory tissue and lymphoid/B-cell follicle-rich areas, and generally low or infrequent in central granuloma regions (revised manuscript, page 7, lines 144-151; Fig. 2 legend).

      (4) Figure 2D, lines 149-151. The authors should clarify whether the CD1c resistance to downregulation is model-specific to THP1-CD1c-APCs or an overexpression artefact

      As described in our response to Reviewer 1, 5, we agree that this point required clearer explanation. We have clarified in the revised Results and Discussion that the preservation of CD1c expression in THP1-CD1c APCs likely reflects the engineered nature of this system, rather than a general feature of endogenous CD1c regulation during Mtb infection.

      Specifically, in primary MoDCs, Mtb infection reduces CD1c expression at both transcript and protein levels (Fig. 2B and 2C). In contrast, CD1c in THP1-CD1c APCs is expressed from a heterologous CD1c-β2m fusion construct rather than from the endogenous CD1C locus. Therefore, its resistance to downregulation is likely model-specific and related to the expression system. We now state this in the Results and Discussion, and explain that this feature allows CD1c-dependent T-cell responses to Mtb-infected APCs to be assessed without the confounding effect of infection-induced CD1c loss (revised manuscript, page 7-8, lines 165-174; page 17-18, lines 411-430).

      (5) Figure 6C. The relative Mtb burden is measured through luminescence. While this correlates closely with CFUs, confirmation with plating is better evidence.

      We have previously shown close correlation between luminescence in our system and CFUs<sup>5</sup> (Bielecka mBio 2017, PMID: 28174307). Perhaps controversially, we propose that luminescence is a better readout of total Mtb load. Luminescence captures all metabolically active Mtb, whilst CFUs may be confounded by clumping of bacteria, for example, giving an underestimate. However, we agree with the reviewer that luminescence is an indirect measure of bacterial burden and that CFU plating would provide additional confirmatory evidence. We have therefore revised the Results, Discussion and Methods to describe the assay more cautiously as a measure of relative viable Mtb burden/luminescence, and we now acknowledge the absence of CFU confirmation as a limitation of the study (revised manuscript, page 11-12, lines 266-274; page 17, lines 400-403; page 26, lines 633-637).

      (6) Figure 7. The authors show that CD1c-autoreactive T-cells exhibit cytotoxic effector memory phenotype. While the sc-RNAseq subsampling is robust, the number of sample donors being 2 might create a potential bias.

      We agree that the use of two donors for the single-cell RNA-seq analysis is a limitation and could introduce donor-specific bias. We have now stated this more explicitly in the Discussion. Importantly, the scRNA-seq data are not used alone to define function, but rather provide an ex vivo phenotypic framework that complements the functional assays showing CD1c-dependent activation, cytokine production, cytotoxicity and reduced relative Mtb burden.

      The subsampling analysis supports the robustness of the transcriptional patterns within this dataset, but we agree that larger donor cohorts will be required to determine how consistently these cytotoxic effector-memory programmes are represented across the broader human CD1c-autoreactive T-cell compartment. We have revised the Discussion accordingly (revised manuscript, page 17, lines 403-410).

      (7) The authors should on how it might compare with non-autoreactive CD1c-restricted T-cells.

      We agree that it is important to place these findings in the context of non-autoreactive, antigen-specific CD1c-restricted T-cells. Previous studies have shown that CD1c can present microbial lipid antigens, such as mycobacterial lipids, to T-cells with defined antigen specificity. In contrast, the CD1c-autoreactive T-cells studied here recognise endogenous ligands and appear to respond to infection through changes in the CD1c-presented lipid repertoire rather than through recognition of a single defined foreign antigen.

      Our findings suggest that autoreactive CD1c-restricted T-cells may provide a complementary mode of immune surveillance, capable of sensing infection-induced changes in lipid presentation, whereas non-autoreactive CD1c-restricted T-cells may respond more directly to specific microbial lipid antigens. A head-to-head comparison would clearly be very interesting, but an extensive new piece of work beyond the scope to the current study. We have expanded the Discussion to more clearly highlight this distinction, and that direct comparison is required (revised manuscript, page 16-17, lines 380-410).

      Concluding remarks

      In summary, we have addressed the key concerns raised by the reviewers by strengthening validation of the experimental system, improving characterisation of T cell populations, and clarifying mechanistic interpretation. We believe these revisions significantly improve the clarity and rigour of the manuscript. We accept that identification of the CD1-presented lipids is an important next step that will give significant mechanistic insight, but is beyond the scope of the current work.

      References:

      (1) Wen, Q. et al. MiR-381-3p Regulates the Antigen-Presenting Capability of Dendritic Cells and Represses Antituberculosis Cellular Immune Responses by Targeting CD1c. J Immunol 197, 580-589 (2016). https://doi.org/10.4049/jimmunol.1500481

      (2) Gagliardi, M. C. et al. Bacillus Calmette-Guerin shares with virulent Mycobacterium tuberculosis the capacity to subvert monocyte differentiation into dendritic cell: implications for its efficacy as a vaccine preventing tuberculosis. Vaccine 22, 3848-3857 (2004). https://doi.org/10.1016/j.vaccine.2004.07.009

      (3) Grailer, J. et al. A Novel Cell-based Luciferase Reporter Platform for the Development and Characterization of T-Cell Redirecting Therapies and Vaccine Development. J Immunother 46, 96-106 (2023). https://doi.org/10.1097/CJI.0000000000000453

      (4) Guo, T. et al. A Subset of Human Autoreactive CD1c-Restricted T Cells Preferentially Expresses TRBV4-1(+) TCRs. J Immunol 200, 500-511 (2018). https://doi.org/10.4049/jimmunol.1700677

      (5) Bielecka, M. K. et al. A Bioengineered Three-Dimensional Cell Culture Platform Integrated with Microfluidics To Address Antimicrobial Resistance in Tuberculosis. mBio 8 (2017). https://doi.org/10.1128/mBio.02073-16

    1. eLife Assessment

      By using a combination of patch clamp recordings, calcium imaging and computer modeling, the authors analyze the spatial distribution of voltage gated calcium channels at glutamatergic synapses formed between layer 5 pyramidal neurons (L5PNs) and between layer 2/3 and L5PNs in the prefrontal cortex (PFC) and primary somatosensory cortex (S1); they conclude that the calcium channel-vesicle coupling is looser in the PFC compared to S1, and future ultrastructural studies may directly confirm the proposed differences in the spatial organization of voltage-gated calcium channels and vesicular release sites. Overall, these findings are important because they have implications for shaping synaptic plasticity and neural circuit function across brain regions. They are convincing because they are based on the use of a multi-pronged approach, although the presentation would benefit from stronger integration of the current findings with the existing literature and a more explicit discussion of potential limitations and confounding factors for data interpretation.

    2. Reviewer #1 (Public review):

      Summary:

      This study asks whether synapses formed by the same broad neuronal class (excitatory pyramidal neurons, PN) adapt their presynaptic organization in a cortex-specific manner, comparing prefrontal cortex (PFC) with primary somatosensory cortex (S1). The authors combine sophisticated electrophysiology (paired recordings and extracellular minimal stimulation), pharmacological perturbations of presynaptic Ca²⁺-secretion coupling, bouton Ca²⁺ imaging, and mechanistic modeling. Across two prominent excitatory connections (Layer 5 (L5) PN-L5PN and L2/3-L5PN), they provide convergent evidence that mature PFC synapses operate with looser Ca²⁺ channel-release sensor coupling than their S1 counterparts.

      Overall, the study provides an appealing mechanistic link between synaptic nano/micro-architecture and cortical-area specialization. The idea that PFC synapses retain a more "plasticity-favoring" presynaptic state, while primary sensory cortex emphasizes reliability and timing precision, is potentially impactful for how we think about circuit computation and plasticity across cortical hierarchies.

      Strengths:

      A major strength is the multi-pronged experimental strategy. The paper first establishes robust, area-dependent differences in synaptic efficacy, reliability, timing, and short-term plasticity (facilitation prevailing in PFC versus depression in S1), using both paired recordings and minimal extracellular stimulation paradigms. The coupling interpretation is then directly supported by differential sensitivity to EGTA (and appropriate positive-control effects of fast chelators). Finally, volume averaged calcium signals are reported to be similar across areas, arguing against trivial explanations based on gross differences in calcium influx, and the modeling provides a quantitative framework for interpreting the observed chelator effects.

      Weaknesses:

      Limitations are minor and concern interpretation/clarity rather than core results. Some key inferences rely on indirect readouts (chelator sensitivity, fluctuation analysis-derived parameters, bouton-averaged calcium signals), each of which carries assumptions and potential confounds that should be discussed more explicitly. In particular, the, repatching paradigm for the paired-recording EGTA experiment, though very impressive, and the limited number of extracellular calcium conditions used for fluctuation analysis (three concentrations) can influence quantitative estimates and the confidence intervals around them.

    3. Reviewer #2 (Public review):

      Schwarze et al. investigated whether synaptic efficacy is brain-region specific. To this end, they compared synaptic connections established by layer 5 (L5) neocortical pyramidal cells and between L5 and L2/3 pyramidal cells. In order to identify the mechanism of this brain region specificity, the authors employed several experimental approaches, including paired electrophysiological recordings, extracellular stimulation, low- and high-affinity intracellular calcium chelators (EGTA and BAPTA), multiple probability fluctuation analysis (MPFA), and intracellular measurements of calcium transients as well as computational modelling. The findings of the present study indicate that synaptic connections in the primary somatosensory cortex (S1) are significantly stronger and more reliable than those in the prefrontal cortex (PFC).

      The study is timely and the topic is of significant interest to the neuroscience community. Despite the extensive research that has been carried out on the neuroanatomy and receptor distribution of different brain regions, comparatively little attention has been paid to differences in synaptic physiology. The authors' approach is characterised by its elegance and comprehensive nature, and the conclusions drawn are compelling.

      Comments on revised manuscript:

      I have no further issues with the present version of the manuscript. All my concerns and/or recommendations were satisfactorily addressed.

    4. Reviewer #3 (Public review):

      Summary:

      In this manuscript, Max Schwarze and colleagues examined the coupling distance between presynaptic Ca²⁺ channels and the vesicular release sensor at neocortical synapses in mouse. They propose that Ca²⁺ channel-release sensor coupling differs across cortical areas, with relatively loose (microdomain) coupling in prefrontal cortex (PFC) and tighter (nanodomain) coupling in primary somatosensory cortex (S1) for comparable pyramidal-neuron synapse types. To test this, they combine paired recordings and minimal stimulation with chelator manipulations (EGTA/BAPTA), mean-variance/MPFA-style analyses, presynaptic Ca²⁺ imaging, and computational modeling. They conclude that presynaptic coupling organization is area-specific in the mature cortex and contributes to regional differences in synaptic timing, reliability, and short-term plasticity.

      Strengths:

      This study tackles an important question and is strengthened by a cohesive body of evidence assembled from multiple complementary approaches. A major asset is the inclusion of high-value datasets, particularly the paired recordings between L5 pyramidal neurons and the systematic assessment of EGTA sensitivity, which provide a solid functional foundation for the authors' central claims. The work is further distinguished by its genuinely multimodal design: combining electrophysiology with presynaptic calcium imaging (and integrating these observations with quantitative analyses and modeling) offers a more mechanistic view of neurotransmitter release than any single method could provide. Overall, the direct, within-framework comparison of presynaptic release-control mechanisms across cortical areas for comparable synapse types is compelling and gives the conclusions a level of robustness and interpretability that is often difficult to achieve in studies of cortical synaptic diversity.

      Weaknesses:

      The principal limitation is incomplete cellular and synaptic specificity in parts of the study. The L2/3-L5PN experiments rely on minimal extracellular stimulation and therefore do not unambiguously identify the presynaptic neuron, its subtype or the number of recruited axons. Similarly, calcium imaging was performed at boutons on L5PN axon collaterals without identifying their postsynaptic targets. The imaging measurements could therefore combine boutons contacting pyramidal neurons and interneurons, potentially obscuring target-dependent differences in presynaptic calcium regulation. Recent connectomic studies demonstrate that local L5 pyramidal-cell axons can distribute substantial fractions of their output to inhibitory neurons, although the exact proportions depend strongly on pyramidal-cell subtype and distance along the axon.

      The quantitative coupling-distance estimate is also model-dependent. The approximately 50-nm estimate for PFC synapses follows from a particular ring-like VGCC geometry and release-sensor model. The simulations demonstrate that this configuration is compatible with the data, but they do not uniquely identify the underlying molecular architecture.

      Overall, the experiments support the narrower conclusion that the examined PFC and S1 synapses differ in functional Ca²⁺-channel-release-sensor coupling. The associated differences in synaptic timing, efficacy and plasticity are compelling, although coupling distance is not isolated causally from other regional differences in release-site number and quantal properties. The proposal that loose coupling is a general correlate of higher-order cortical function remains an interesting but currently speculative interpretation. Further comparisons across additional cortical regions and genetically or projection-defined synapse types will be particularly helpful in establishing the broader generality of this concept

      Comments on revised version.

      The authors have addressed most of my comments in the revised manuscript. I have only one remaining, relatively minor suggestion concerning point 5. I appreciate that the authors now acknowledge the possibility that the imaged boutons may contact different postsynaptic targets. However, the argument that interneuron-targeting boutons are likely to make only a minor contribution, based on the overall proportions of excitatory neurons or inhibitory synapses in the cortex, may not fully resolve this concern. Excitatory pyramidal neurons can distribute their outputs non-randomly across excitatory and inhibitory targets, and this distribution may depend on pyramidal-cell subtype and axonal distance. For example, a recent MICrONS/Allen Institute connectomic analysis of L5 extratelencephalic neurons in mouse visual cortex found that approximately two-thirds of their proximal synaptic outputs contacted inhibitory neurons. The proportion was close to 80% near the soma and decreased progressively with distance along the axon.

      These findings concern a specific L5 pyramidal-cell subtype in visual cortex and therefore cannot be transferred directly to the PFC and S1 preparations examined here. Nevertheless, they illustrate that the postsynaptic target distribution of L5 pyramidal-neuron boutons cannot necessarily be inferred from the overall abundance of excitatory and inhibitory neurons or synapses.

    5. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This study asks whether synapses formed by the same broad neuronal class (excitatory pyramidal neurons, PN) adapt their presynaptic organization in a cortex-specific manner, comparing the prefrontal cortex (PFC) with the primary somatosensory cortex (S1). The authors combine sophisticated electrophysiology (paired recordings and extracellular minimal stimulation), pharmacological perturbations of presynaptic Ca<sup>2+</sup>-secretion coupling, bouton Ca<sup>2+</sup> imaging, and mechanistic modeling. Across two prominent excitatory connections (Layer 5 (L5) PN-L5PN and L2/3-L5PN), they provide convergent evidence that mature PFC synapses operate with looser Ca<sup>2+</sup> channel-release sensor coupling than their S1 counterparts.

      Overall, the study provides an appealing mechanistic link between synaptic nano/micro-architecture and cortical-area specialization. The idea that PFC synapses retain a more "plasticity-favoring" presynaptic state, while the primary sensory cortex emphasizes reliability and timing precision, is potentially impactful for how we think about circuit computation and plasticity across cortical hierarchies.

      Strengths:

      A major strength is the multi-pronged experimental strategy. The paper first establishes robust, area-dependent differences in synaptic efficacy, reliability, timing, and short-term plasticity (facilitation prevailing in PFC versus depression in S1), using both paired recordings and minimal extracellular stimulation paradigms. The coupling interpretation is then directly supported by differential sensitivity to EGTA (and appropriate positive-control effects of fast chelators). Finally, volume-averaged calcium signals are reported to be similar across areas, arguing against trivial explanations based on gross differences in calcium influx, and the modeling provides a quantitative framework for interpreting the observed chelator effects.

      Weaknesses:

      Limitations are minor and concern interpretation/clarity rather than core results. Some key inferences rely on indirect readouts (chelator sensitivity, fluctuation analysis-derived parameters, bouton-averaged calcium signals), each of which carries assumptions and potential confounds that should be discussed more explicitly. In particular, the repatching paradigm for the paired-recording EGTA experiment, though very impressive, and the limited number of extracellular calcium conditions used for fluctuation analysis (three concentrations), can influence quantitative estimates and the confidence intervals around them.

      We would like to thank the reviewer for his/her overall positive assessment of our manuscript and the constructive advice, which helped us to improve our manuscript. We discussed the limitations, assumptions and potential confounding factors in more detail. We addressed them pointwise in the recommendations for the authors.

      Reviewer #2 (Public review):

      Schwarze et al. investigated whether synaptic efficacy is brain-region specific. To this end, they compared synaptic connections established by layer 5 (L5) neocortical pyramidal cells and between L5 and L2/3 pyramidal cells. In order to identify the mechanism of this brain region specificity, the authors employed several experimental approaches, including paired electrophysiological recordings, extracellular stimulation, low- and high-affinity intracellular calcium chelators (EGTA and BAPTA), multiple probability fluctuation analysis (MPFA), and intracellular measurements of calcium transients as well as computational modelling. The findings of the present study indicate that synaptic connections in the primary somatosensory cortex (S1) are significantly stronger and more reliable than those in the prefrontal cortex (PFC).

      The study is timely, and the topic is of significant interest to the neuroscience community. Despite the extensive research that has been carried out on the neuroanatomy and receptor distribution of different brain regions, comparatively little attention has been paid to differences in synaptic physiology. The authors' approach is characterised by its elegance and comprehensive nature, and the conclusions drawn are compelling. Nevertheless, there are a number of unresolved issues.

      First, we would like to thank the reviewer for his/her detailed survey of our work, which was very helpful in improving our manuscript. We are happy about the overall positive evaluation and the constructive comments. To fully clarify all points, we performed new experiments and analyses, in particular we determined EGTA sensitivity in PFC and S1 from the same animal and we performed MPFA with an additional extracellular Ca<sup>2+</sup> concentration. We extended the discussion on the examined cell types. Overall, we carefully revised the manuscript to address all points. Please see below our point-wise response.

      Major points:

      (1) The authors state that data from the S1 cortex were obtained in a previous study. In the context of an explicitly comparative study (PFC vs. S1cortex), it would have been advantageous for the authors to perform a subset of experiments in which both cortices were obtained from a single animal. This is a feasible undertaking, given the spatial separation of the PFC and S1 cortex.

      This is only true for the paired recordings from L5PN-L5PN connections in S1, which were obtained in a previous study and partially reanalyzed. All recordings from L2/3-L5PN connections in S1 and PFC as well as the paired recordings on L5PN-L5PN synapses in PFC were obtained in the present study. To make this clearer, we have added Table 1. This lists which data and associated figures are from this study and which are from previous studies (Bornschein et al., Cell Rep. 2019; Bornschein et al., Front. Syn. Neurosci. 2019).

      Our experiments are lengthy and therefore it is challenging to achieve two successful recordings within the lifetime of acute brain slices. For this reason, the previous version of the manuscript did not include recordings from PFC and S1 of the same animal. We have now measured EGTA effects in L2/3-L5PNs from S1 and PFC of the same animal. Two new recordings were added to the EGTA-AM plots in Figure 2C-F and in the results section. Example recordings are shown in Figure S3A and B, as is the comparison of EGTA-AM effects in L2/3-L5PNs from PFC and S1 in Figure S3C, with data points from the same animal marked.

      “On the other hand, EGTA significantly reduced EPSC amplitudes only in PFC (0.54, 0.47-0.67; 55% of control) but not in S1 (0.85, 0.81-1.04; 100% of control). For a better comparison of EGTA effects some recordings were performed in PFC and S1 derived from the same animal to rule out interindividual effects (see example recording in Figure S3A-C).”

      (2) Figure 1A is somewhat misleading because it could suggest that the authors have performed dual recordings in identified PFC pyramidal cells.

      We thank the reviewier for this helpful note. We added “L2/3 or L5” to the stimulation panel of Figure 1A to illustrate that we stimulated either extracellularly in L2/3 or L5PNs directly via the patch pipette.

      (3) PFC and S1 cortex in rodents differ markedly in their morphological organisation. For example, in all sensory cortices, layer 4 is very pronounced; however, in the PFC of rodent,s no clear layer 4 can be found. On the other hand, PFC shows a clear separation of layers 2 and 3, which is not visible inthe S1 cortex. Furthermore, PFC pyramidal cells in layers 2, 3, and 5 exhibit significant heterogeneity, diverging considerably from those found in layers 5a and 5b of S1 cortex. Thus, there is no clear correlation between L5 pyramidal cells in the PFC and the S1 cortex. In order to achieve a meaningful comparison of the data obtained in PFC and S1 cortex, it is necessary for the authors to determine whether the record is from similar pyramidal cell populations.

      (3) In addition, PFC pyramidal cells in layer 2, 3 and 5 are highly heterogeneous and differ markedly from those in layer 5a and 5b of S1 cortex. To achieve a meaningful comparison of the data obtained in the PFC and the S1 cortex, the authors need to determine whether the record from similar pyramidal cell populations.

      We apologize for having not been precise about the specific location and type of pyramidal neurons in the original manuscript. Extracellular stimulation in PFC and S1 was always performed in layer 2, where the first large cell bodies, relative to the pia mater, are located within a cortical column. Therefore, we assume that the same cell populations were stimulated in both brain regions (van Aerde & Feldmeyer, Cereb. Cortex 2015; Oberlaender et al., Cereb. Cortex 2012; Lefort et al, Neuron 2009). To stick with the standard terminology, we refer to it in the manuscript as upper layer 2/3. We kept stimulation intensity as low as possible to ensure that only a few presynaptic cells within the target region were activated.

      In S1, paired recordings were obtained from pyramidal neurons in layer 5A following the procedures and criteria described in detail in our previous work (Bornschein et al., Cell Rep. 2019; Bornschein et al., Front. Syn. Neurosci. 2019; Bornschein et al., Science 2025). Briefly, these criteria are as follows: close proximity to layer 4 and the barrels as well as the PPR of 0.78 (0.69-0.90), which is consistent with depression dominating in L5A (Frick et al., Cereb. Cortex 2008; Bornschein et al., Front. Syn. Neurosci. 2019) and different from L5B with PPR ≥ 1 (Lefort & Petersen, Cereb. Cortex 2017). Within layer 5A we did not attempt to further differentiate between pyramidal neuron types. For recordings in S1 with extracellular stimulation we focused on the same locations as for the paired recordings. We extended the corresponding section in Materials and Methods of the revised manuscript.

      We agree that there is no clear layer 4 in PFC, making the distinction between layer 2/3 and layer 5 less clear. Layer 2/3 and layer 5 have approximately the same diameter (van Aerde & Feldmeyer, Cereb. Cortex 2015). Based on this, we performed recordings in upper layer 5 of PFC. We did neither morphologically nor electrophysiologically differentiate between pyramidal neuron cell types. It should be noted that within a cortical area (S1 or PFC), we did not find a difference between glutamatergic synapses from L2/3 onto L5PNs and L5PN-to-L5PN synapses, neither with regard to the EGTA-sensitivity of release nor with regard to the release probability. In particular, we found homogeneous results and similar variability in both, the examined connections in the PFC and in S1, with no discernible clustering in the data that would suggest stimulation of different cell populations. These findings suggest that excitatory inputs to L5PNs exhibit similar properties (PPR, p<sub>N</sub>, CD) irrespective of whether they originate in L2/3 or in neighbouring PNs in L5A. However, we do see significant differences between synapses in the different cortical areas S1 and PFC. Thus, intra-area specific differences in morphology and spiking patterns among pyramidal neurons appear to not be reflected on the level of their synapses.

      The Reviewer probably refers to such differences and heterogeneity in morphology and spiking patterns of pyramidal neurons. If he/she has more specific differences in mind, it would be helpful if references for the significant heterogeneity could be given.

      Please also note that the type of experiments we perform with paired recordings and long-lasting patch-clamp measurements is not suitable for analyzing population differences among pyramidal neuron types.

      We refer to the problem of pyramidal neuron heterogeneity in the revised manuscript in the discussion.

      “Patch-clamp recordings from L5PNs located in the upper layer 5 (L5A in S1) were established according to the criteria described in detail in our previous work on this connection in S1 (Bornschein et al., 2019b; Bornschein et al., 2025). Presynaptic neurons were stimulated extracellularly in upper layer 2/3 (L2/3-L5PN connections) straight above the patched L5PN or in on-cell mode in L5A right next to the postsynaptic cell (L5PN-L5PN connections; Figure 1).“

      “We did neither morphologically nor based on spiking patterns differentiate further between PN subtypes within a given layer. However, within a cortical area (S1 or PFC) we did not find a difference between glutamatergic synapses from L2/3 onto L5PNs and L5PN to L5PN synapses, neither with regard to the EGTA sensitivity of release nor with regard to p<sub>N</sub>. In particular, we found homogeneous results and similar variability in both, the examined connections in the PFC and in S1, with no discernible clustering in the data that would indicate stimulation of different cell populations. These findings suggest that excitatory inputs to L5PNs exhibit similar properties (PPR, p<sub>N</sub>, CD) irrespective of whether they originate in L2/3 or in neighboring PNs in L5A. However, we do see significant differences between synapses in the different cortical areas S1 and PFC. Thus, intra-area specific differences in morphology and spiking patterns among PNs appear to be not reflected on the level of their synapses.“

      (4) For the S1 cortex, in rats it has been found that L5 synaptic connection between pairs of L5a pyramidal cells and pairs of L5b pyramidal cells differ markedly with respect to mean EPSP amplitude, latency and coefficient of variation (cv, a surrogate measure for the synaptic release probability) (cf. Markram et al., 1997; Frick et al., 2008). It is therefore likely that PFC and S1 pre- and postsynaptic pyramidal cells are not only morphologically and electrophysiological distinct but also with respect to their synaptic properties. At least, the authors need to discuss these confounding issues and preferentially address them experimentally. For example, it would be helpful to demonstrate that paired recordings were made from the same pyramidal cell types, perhaps by documenting their morphology and/or firing patterns. In addition, they should discuss the marked difference in EPSP amplitude and putative release probability between their data and the earlier studies.

      We agree that Markram et al. (J. Physiol. 1997) and Frick et al. (Cereb. Cortex 2008) provided highly valuable insights into synaptic transmission between pyramidal neurons in S1. We referred to their work in detail in our previous work on developmental changes in the presynaptic organization of transmitter release in L5APN synapses in S1 (Bornschein et al., Cell Rep. 2019). Both studies were performed in young rats and EPSPs were measured, whereas we worked in mice and recorded EPSCs. This impedes a direct comparison of amplitudes.

      Markram et al. (J. Physiol. 1997) recorded in 2-week-old rats from thick tufted PNs, corresponding to L5BPNs. Given the longer lifespan and slower development of rats compared to mice, this likely reflects a maturation state that corresponds better to our previous measurements in 8 to 10-day-old mice. Markram et al. found small failure rate (median 7%), which is similar to what we found in our previous study for young L5APN synapses (low failure rates and high p<sub>N</sub>; Bornschein et al., Cell Rep. 2019).

      The study by Frick et al. (Cereb. Cortex 2008) is closer to our present study and to the mature age window in our previous study, although they also recorded from rats but from L5APN-L5APN pairs in almost 3-week-old animals in S1. Again, EPSPs rather than EPSCs were recorded, impeding a direct comparison of amplitudes.

      Both studies concluded, based on the synaptic failure rate and CV analysis of EPSP amplitude, that the synapses they investigated operate with high release probability. This is fully in line with our findings. Of note, we found no significant difference between L2/3-L5PN and L5PN-L5PN synapses within a given area, indicating that varibality on the synaptic level between PNs of a given area is not pronounced.

      In order to further substantiate this, we determined the relative variability in median EPSC amplitudes to test whether there is a higher variability of recorded cell types in PFC compared to S1. The relative MAD (median absolute deviation) of EPSC amplitudes was 0.46 in PFC and 0.50 in S1. The similarity in these values argues against higher cell-type variability in PFC compared to S1. We have discussed the results of these studies in relation to our own findings.

      “Two other previous studies on L5APN (Frick et al., 2008) and L5BPN (Markram et al., 1997) connections concluded that these synapses operate with high release probability, which nicely agrees with our previous (Bornschein et al., 2019b) and current results. It is remarkable that we did not even detect any differences between the L2/3-L5PN and L5PN-L5PN synapses within a given cortical area. Overall these results from different studies (Markram et al., 1997; Reyes and Sakmann, 1999; Frick et al., 2008; Bornschein et al., 2019b; Bornschein et al., 2019a) may indicate that variability on the synaptic level between PNs of a given area is not pronounced. In order to further substantiate this, we determined the relative variability in median EPSC amplitudes to test whether there is a higher variability of recorded cell types in PFC compared to S1. The relative MAD (median absolute deviation) of EPSC amplitudes was 0.46 in PFC and 0.50 in S1. The similarity in these values argues against higher cell-type variability in PFC compared to S1.

      (5) In order to perform multiple probability fluctuation analysis (MPFA), a parabolic fit with a mere three points is inadequate, particularly because 2 mM and 5 mM Ca<sup>2+</sup> are close to the peak of the variance-to-mean parabola, and only 1 mM Ca<sup>2+</sup> is on its initial linear part. A more meaningful result would have been obtained with an additional Ca<sup>2+</sup> concentration between 1.0 and 2.0 mM, as these are closer to the physiological range. In this context, the authors should have quoted the more recent and more detailed paper by the Silver group (Saviane and Silver, 2006; Lanore and Silver, 2016) and not just the Clements and Silver review paper.

      We used only three Ca<sup>2+</sup> concentrations for MPFA as these resulted in a low (<0.5), a medium (~0.5) and a large (>0.5) p<sub>N</sub> condition, thereby clearly determining a parabola. Also Saviane and Silver (Nature 2006) performed MPFA with three extracellular Ca<sup>2+</sup> concentrations (1, 2, and 8 mM). We have now cited this paper, as well as the more recent work by Lanore and Silver (Neuromethods 2016), in relation to the MPFA method. To verify the reliability of MPFA, the determined parameters were compared with values estimated based on the EPSC amplitudes (EPSC = N p<sub>N</sub> q; PFC, 6 pA; S1, 48 pA) and failure rates (F = (1-p<sub>N</sub>)^N; PFC, 0.25, S1, 0.0001). The estimated values did in fact match those of the MPFA (EPSCs in PFC: 8 pA, 5-15 pA, and S1: 29 pA, 18-53 pA; failure rates in PFC: 0.16, 0.08-0.28, and S1: 0, 0-0.03; see original manuscript.

      To further support this, we have now conducted additional experiments using four Ca<sup>2+</sup> concentrations. The results are consistent with those from the experiments using three concentrations. The additional Ca<sup>2+</sup> concentration of 1.5 mM did not improve the parabolic fit, as it yielded p<sub>N</sub> values very close to those determined with 2 mM Ca<sup>2+</sup>. Therefore, the additional experiments are shown in Author response image 1. The novel p<sub>N</sub> data are included in the summary of p<sub>N</sub> values (now n=6) in the results section and in Figure 3F.

      Author response image 1.

      MPFA with four different extracellular Ca<sup>2+</sup> concentrations. (A) MPFA of EPSC amplitudes recorded at the indicated [Ca<sup>2+</sup>]<sub>e</sub> from L2/3-L5PNs in PFC. Top: Individual EPSCs (grey, average in black) recorded from L5PNs after extracellular stimulation in L2/3. Middle: Plot of EPSC amplitudes over time. Bottom: Corresponding mean-variance plot fitted with a parabola estimating the quantal parameters of release. p<sub>N</sub> is for 2 mM [Ca<sup>2+</sup>]<sub>e</sub>. Recordings were made in the presence of 10 µM Bicuculline, 0.25 mM Kynurenic acid and 50 µM 2-Amino-5-phosphonovaleriansäure. (B) As in (A), but for a L2/3-L5PN connection in S1. (C) Summary of determined p<sub>N</sub> values in PFC and S1 (dots represent individual experiments; P=0.485, Mann-Whitney U rank-sum test).

      Additionally, we performed bootstrap analyses with 10,000 replicates which were generated with replacement from the original data sample. Distributions of bootstrap 25% trimmed means showed a clear separation for N and q between PFC and S1 but not for p<sub>N</sub>. The bootstrap results are shown in Author response image 2.

      Author response image 2.

      Bootstrap analysis (A-C) Distribution of bootstrap 25% trimmed means of the quantal parameters p<sub>N</sub> (A), N (B) and q (C) in PFC (orange), S1 (blue) and S1 with gDGG (light blue). 10,000 bootstrap replicates were generated with replacement from the original data sample obtained by MPFA in L5PN-L5PN connections (cf. Figure 3D). (D) Same as in (A) but for MPFA in L2/3-L5PN connections.

      (6) Methods: The authors should clarify whether their paired recordings from L5 pyramidal cells involved whole-cell recordings from both pre- and postsynaptic neurons. From Figure 1B, it appears as if the presynaptic neurons were not recorded in whole cell mode but rather stimulated in cell-attached mode. This is also reflected in the artefact visible in the current trace recorded in the postsynaptic neuron. The authors should explicitly state their methodological approach and mention how reliable the timing of the presynaptic action potential was under these circumstances. The same holds true for the extracellular stimulation protocol. A significantly more detailed description of the experimental protocol is necessary here.

      In the paired recordings, presynaptic cells were stimulated in the cell-attached mode. For presynaptic EGTA application the whole-cell configuration was established after re-patching to allow buffer perfusion of the presynaptic L5PN. This is described in the methods section of the original version of the manuscript. We extended this description as follows:

      “In paired recordings, presynaptic L5PNs were stimulated in on-cell configuration (200-500 mV, 1-2 ms). In the chelator wash-in experiments, presynaptic neurons were repatched with a pipette solution supplemented with 10 mM EGTA (K-gluconate concentration was reduced to 135 mM to adjust osmolarity) and whole-cell configuration was established to allow EGTA perfusion of the presynaptic neuron.”

      The amplitudes were determined by fitting a product of two exponential functions to the baseline-subtracted currents, which allows for independent adjustment of the time constants of the rising and falling phases and minimizes noise effects (cf. Bornschein et al., J. Physiol. 2013). Synaptic delays were determined from the onset of stimulation to the fitted EPSC onset. We added this more detailed explanation to the methods section. in the timing of the presynaptic action potential was similar in recordings from PFC and S1. This applies to the paired-recordings with on-cell stimulation of presynaptic neurons as well as to the extracellular stimulation experiments.

      “Synaptic responses were determined by fitting a product of two exponential functions to the baseline-subtracted currents, which allows for independent adjustment of the time constants of the rising and falling phases and minimizes noise effects (Bornschein et al., 2013). Synaptic delays were determined from the onset of stimulation to the fitted onset of the EPSC. PPRs were calculated by dividing the second amplitude of two consecutive EPSCs by the first.”

      (7) Methods: The authors use Student's t-test for data comparison. The authors should verify that the data distribution was indeed normal, e.g. by using a Shapiro-Wilk test. If this is not the case, non-parametric tests should be used.

      We typically used non-parametric tests as stated in the figure legends of the corresponding figures. We have now explained the abbreviations for the Mann-Whitney U test (MWU) and the Wilcoxon signed-rank test (WSR) in the figure legends. A paired t-test was used only in Figure 5F after testing for normal distribution with the Shapiro-Wilk test. This is described in the methods section of the original version of the manuscript. Additionally, results of the Shapiro-Wilk test were now included in the figure legends.

      “Normality was tested using the Shapiro-Wilk test. Normally distributed data were compared with the t-test (two groups) or a one-way ANOVA (more than two groups). Non-normally distributed or small samples of data were compared with the Mann-Whitney U rank-sum test (MWU; two groups) or a Kruskal-Wallis ANOVA on ranks (more than two groups). (…). To compare pre- and post-treatment data the paired t-test or the Wilcoxon signed-rank test (WSR) was used, depending on the distribution of the data.”

      Reviewer #3 (Public review):

      Summary:

      In this manuscript, Max Schwarze and colleagues examined the coupling distance between presynaptic Ca<sup>2+</sup> channels and the vesicular release sensor at neocortical synapses in mice. They propose that Ca<sup>2+</sup> channel-release sensor coupling differs across cortical areas, with relatively loose (microdomain) coupling in prefrontal cortex (PFC) and tighter (nanodomain) coupling in primary somatosensory cortex (S1) for comparable pyramidal-neuron synapse types. To test this, they combine paired recordings and minimal stimulation with chelator manipulations (EGTA/BAPTA), mean-variance/MPFA-style analyses, presynaptic Ca<sup>2+</sup> imaging, and computational modeling. They conclude that presynaptic coupling organization is area-specific in the mature cortex and contributes to regional differences in synaptic timing, reliability, and short-term plasticity.

      Strengths:

      This study tackles an important question and is strengthened by a cohesive body of evidence assembled from multiple complementary approaches. A major asset is the inclusion of high-value datasets, particularly the paired recordings between L5 pyramidal neurons and the systematic assessment of EGTA sensitivity, which provide a solid functional foundation for the authors' central claims. The work is further distinguished by its genuinely multimodal design: combining electrophysiology with presynaptic calcium imaging (and integrating these observations with quantitative analyses and modeling) offers a more mechanistic view of neurotransmitter release than any single method could provide. Overall, the direct, within-framework comparison of presynaptic release-control mechanisms across cortical areas for comparable synapse types is compelling and gives the conclusions a level of robustness and interpretability that is often difficult to achieve in studies of cortical synaptic diversity.

      Weaknesses:

      Several aspects would benefit from clearer explanation, stronger integration with the existing literature, and a more explicit discussion of limitations and potential confounds. Without these additions, some conclusions remain speculative. Throughout the manuscript, the authors also often imply that different measurements reflect the same underlying synapse population. This is unlikely to be strictly true across all experiments and makes it difficult to integrate results from the various approaches into a single, unified set of functional synaptic properties. In addition, some statements-particularly those linking coupling mode to "higher-order neocortical functions"-appear broader than what is directly supported by the experiments and should be tempered or more precisely scoped.

      Below, I list several topics that could help better frame the main findings of the present study and clarify how it relates to previously published work.

      We would like to thank the reviewer for the comprehensive and detailed assessment of our manuscript and his/her overall positive evaluation. We have addressed all of the reviewer's points. We expanded the model description and discussion, and slightly toned down our conclusion.

      (1) The authors use EGTA sensitivity of EPSCs (together with additional metrics) to argue that S1 and PFC synapses differ in Ca<sup>2+</sup> channel-release sensor coupling. While this is a plausible interpretation, EGTA effects are not uniquely determined by coupling distance and can also reflect differences in Ca<sup>2+</sup> entry kinetics, action potential waveform, endogenous buffering/extrusion, or release-sensor/vesicle state. The authors use a constrained modeling approach, but the rationale for the different constraint sets is not fully clear from the current description. It would be helpful to expand and clarify the Methods section to explain how these constraints were defined, justified, and applied (and how alternative constraint choices would affect the results). In this context, the Abstract's broader claim that the study "reveals microdomain coupling as a presynaptic structure-function correlate of higher-order neocortical functions" appears overstated. Given the well-known diversity of cortical synapses even within a single region (e.g., synapses onto different interneuron subclasses or different PN cell types, extracortical sources like thalamus), the authors should clarify the intended scope: is the conclusion meant to apply broadly across synapse classes in S1 and PFC, or only to the specific connection type(s) examined here?

      We would like to thank the reviewer from pointing out that our description fell a bit short, in particular with respect to the interpretation of the EGTA effects. We addressed the points as follows in the revised manuscript: We discussed the interpretation of EGTA effects in more detail. We toned down the concluding statement in the last sentence of the Abstract.

      “Differences in the sensitivity of release to low to moderate concentrations of EGTA (≤ 30 mM) are a standard indicator of differences in the coupling distance (e.g. Adler et al., 1991; Bucurenciu et al., 2008; reviewed in Eggermann et al., 2012; Vyleta and Jonas, 2014; Kusch et al., 2018; Bornschein et al., 2019b). p<sub>N</sub> is determined by the size of the Ca<sup>2+</sup> signal at the release sensor and the binding kinetics and affinity of the sensor. The former in turn is determined by the details of the Ca<sup>2+</sup> influx and the diffusional coupling distance between the VGCCs and the sensor. The similarity of Ca<sup>2+</sup> signals between synapses in PFC and S1 (Figure 4) indicates that Ca<sup>2+</sup> influx is similar between boutons, although more subtle differences in the influx kinetics may have remained undetected in these volume-averaged signals. Regarding sensor affinity, results in a previous study indicate that differences in EGTA sensitivity show differences in coupling rather than sensor affinity even if k<sub>on</sub> of the sensor and its affinity should differ as much as ten-fold, which appears to be an unlikely scenario given that even the two major isoforms of Synaptotagmin that trigger synchronous release differ by less than a factor of three to four in their affinity (Bollmann et al., 2000; Schneggenburger and Neher, 2000; Bornschein et al., 2025). Finally, the increase in the PPR induced by the application of Cd<sup>2+</sup> further supports our conclusion of microdomain coupling in the PFC synapses (Scimemi and Diamond, 2012).”

      “They suggest that microdomain coupling in pyramidal neuron synapses could be a presynaptic structure-function correlate of higher order neocortical functions.”

      (2) The chelator logic is sound in principle, but the Discussion should more explicitly acknowledge standard caveats and alternative explanations. The authors partly address this by including presynaptic Ca<sup>2+</sup> imaging and modeling, yet it would help to explain more clearly how the combination of (i) chelator sensitivity, (ii) presynaptic Ca<sup>2+</sup> signals, and (iii) model constraints rules out-or substantially reduces the likelihood of-changes in AP waveform, Ca<sup>2+</sup> influx kinetics, buffering/extrusion, or sensor/vesicle state as the primary drivers. In addition, recent hypotheses emphasizing vesicle priming and/or release-site occupancy as contributors to apparent EGTA sensitivity should be discussed as a complementary or alternative interpretation.

      Please see above the first part of the discussion to point one.

      (3) A substantial portion of the S1 comparison appears to rely on previously published datasets. This should be made unambiguous in the Results and Methods, and it would be helpful to summarize this clearly (e.g., in a table indicating which figures/analyses use new data versus reanalysis of published data). If this information is already present, it should be highlighted more prominently.

      Please excuse us for not having made it clearer which data had already been published. Only the paired recordings from L5PN-L5PN connections in S1 were obtained in previous studies and partially reanalyzed. Paired recordings on the same synapses in PFC as well as all recordings from L2/3-L5PN connections in PFC and S1 were obtained in the present study. At your suggestion, we have added Table 1 highlighting which data and associated figures are from this study and which were acquired in previous studies (Bornschein et al., Cell Rep. 2019; Bornschein et al., Front. Syn. Neurosci. 2019).

      (4) The modeling is informative, but the choice of a specific VGCC-release-site geometry and channel arrangement is not sufficiently justified. The manuscript adopts a particular spatial configuration, yet the rationale for selecting this geometry, rather than other plausible architectures discussed in the literature, is not clearly explained, nor is it meaningfully revisited in the Discussion. The authors should justify why the same organization is assumed across two distinct cortical areas and, ideally, include (or at a minimum discuss) a sensitivity analysis showing how key inferences (e.g., coupling distance and channel number) depend on the assumed geometry.

      We extended the discussion of why a ring-like structure of VGCCs was assumed in the model.

      “The microdomain was assumed to be formed by a ring-like structure of VGCCs around a vesicle (Figure 5D). This topography was chosen because such a microdomain was found to best predict the experimental data of transmitter release from PNs in young S1 (Bornschein et al., 2019b). Other previously described distributions of VGCCs suitable to reproduce release data cover random distributions of VGCCs (Scimemi and Diamond, 2012), VGCC clusters (Meinrenken et al., 2002; Nakamura et al., 2015), and exclusion zones (Keller et al., 2015). In the early S1, all of these models predicted a higher EGTA sensitivity of the microdomain, however, these models provided a poorer fit to the full set of the experimental data than the ring-like structure (Bornschein et al., 2019b). Since the experimental data from PNs in the mature PFC were similar to those in young S1, these other microdomain models were not tested explicitly here.”

      (5) The calcium imaging data are valuable, but given the diversity of synapses within each cortical layer, it is not clear that imaged boutons can be confidently assigned to the specific connection types being interrogated electrophysiologically. A substantial fraction of boutons likely corresponds to different postsynaptic targets (including interneurons and distinct pyramidal-cell classes), and this heterogeneity could complicate interpretation. This limitation should be discussed explicitly

      Excitatory pyramidal cells make up 80-85% of cortical neurons, with the highest density in layer 5 (Keller et al., Front. Neuroanat. 2018). In the somatosensory cortex, inhibitory synapses account for only about 10% (Santuy et al., Brain Struct. Funct. 2018). We imaged a large number of presynaptic boutons within layer 5 (about 10 boutons per cell, in total 85 boutons in PFC and 100 boutons in S1, numbers of boutons were now included in Figure 4). In this respect, the impact of inhibitory synapses is minor. Since connectivity between neighboring PNs in layer 5A is high (Feldmeyer, Front. Neuroanat. 2012), we assume that a large proportion of the imaged boutons target neighbouring L5PNs. We added a sentence on potential postsynaptic targets in the results section.

      “The imaged presynaptic boutons most likely connect to neighboring pyramidal cells, as connectivity between L5PNs in layer 5A is high (Feldmeyer, 2012). Nevertheless, a small proportion of other postsynaptic targets, such as interneurons, cannot be ruled out.”

      (6) In unitary connections, the authors assess EGTA effects alongside other functional parameters (strength, delay, short-term plasticity), which is a major strength. However, for L2/3 to L5 connections, it appears that EGTA sensitivity was tested primarily using extracellular stimulation. Given anatomical and circuit differences between PFC and S1, extracellular stimulation may recruit different synapse populations across regions, potentially confounding regional comparisons of EGTA sensitivity. This limitation should be acknowledged explicitly. While I am not requesting technically demanding L2/3↔L5 paired recordings in S1, the possibility that different synapse identities are being sampled should be treated as a meaningful source of uncertainty. The Discussion would also benefit from placing the magnitude of EGTA effects in the context of prior "loose coupling" literature, where comparatively large EGTA effects have been reported in some systems. In addition, the reported difference between adult PFC EGTA effects and S1 inhibition appears small (on the order of <10%) and should be interpreted cautiously, especially given that PFC and S1 mature on different timelines and P21-P26 is unlikely to reflect a mature PFC circuit state. The adult cohort (P90-P100) is therefore important, but the age mismatch complicates PFC-S1 comparisons; ideally, S1 should be assessed at matched ages, or this limitation should be discussed explicitly. Finally, for statistical robustness, in panel D of Figure 2, were the comparisons corrected for multiple testing to control Type I error?

      EGTA sensitivity was examined in PFC and in S1, for two connections in each region - using paired recordings for L5PN-L5PN connections and using extracellular stimulation for L2/3-L5PN connections. The L5PN-L5PN data from S1 were collected in an earlier study (Bornschein et al., Cell Rep. 2019; Bornschein et al., Front. Syn. Neurosci. 2019), have now been reanalyzed for the test period between 20 and 30 min, and included in Figure 2B for the sake of consistency (see Table 1). To emphasize this point, despite stimulating different input synapses with different stimulation methods, we obtained similar results in the respective brain regions. This suggests that the EGTA sensitivity observed in the investigated PFC connections is not a solely synapse-specific property.

      It is difficult to compare the absolute EGTA sensitivities from different synapses from different publications, since EGTA effects do not depend exclusively on the coupling distance, as the reviewer also noted in point 1. They are, among other factors, influenced by the Ca<sup>2+</sup> sensitivity of the release machinery, which differs between our Syt1-expressing cortical synapses and Syt2-expressing synapses in other brain regions (Schneggenburger et al., Nature, 2000; Bollmann et al., Science, 2000; Bornschein et al., Science 2025), such as the calyx of Held or the cerebellar basket to Purkinje cell synapse. Furthermore, direct patching and loading of the presynaptic bouton with EGTA - as feasible at the calyx of Held and other large synapses - results in higher effective EGTA concentrations compared to somatic loading of presynaptic terminals, despite identical pipette concentrations. The buffer-AM method introduces additional uncertainty regarding the effective intra-bouton EGTA concentration, since the loading efficacy has to be estimated. Thus, although differences in EGTA sensitivity primarily show differences in coupling distances, the comparison of absolute values between different publications is difficult. Consistently, data-constrained models are used to estimate the coupling topography and to compare these topographies rather than comparing the absolute EGTA effects (e.g. Buccurenciu et al., Neuron, 2008; Vyleta and Jonas, Science, 2014; Bornschein et al., Cell Rep., 2019; Chen et al., Neuron, 2024; Bornschein et al., Science, 2025).

      We include a note on this in the discussion.

      “Thus, although the absolute EGTA sensitivity is influenced by different factors, which necessitates data-constrained models for quantitative comparisons, the general sensitivity of release to EGTA indicates loose coupling.”

      In mouse neocortex postnatal maturation in S1 and PFC follows the same time course. Kroon et al. (Sci. Rep. 2019) reported that maturation of dendritic morphology and intrinsic properties of pyramidal neurons occurs within the first two weeks after birth, now cited in the discussion. Therefore, it is unlikely that the EGTA effect in PFC is due to a delayed maturation. The difference in EGTA sensitivity between PFC at P90-100 and S1 at P21-26 is indeed small but significant (P=0.009, Mann-Whitney-U rank sum test).

      “Since postnatal development follows the same time-course in mouse PFC and S1 and occurs predominantly within the first two weeks after birth, (…) (Kroon et al., 2019).”

      Thank you for the advice concerning statistical robustness. We replaced the Mann-Whitney-U rank sum test in Figure 2D by a one-way ANOVA and performed a Holm-Sidak post-hoc test correcting for multiple comparisons. Similarly, ANOVA was used to compare more than two groups in Figures 1K and 2F. We changed the corresponding P values and tests in the figure legends and added the performed post-hoc tests in the methods section.

      “For multiple comparisons post-hoc testing was performed with the Holm-Sidak (one-way ANOVA) or Dunn´s method (ANOVA on ranks).”

      (7) Alterations in initial release probability are often associated with changes in short-term plasticity. In the present manuscript, the authors report similar initial release probability at PFC and S1 synapses, yet observe differences in short-term plasticity profiles. The mechanistic basis for this apparent dissociation is not addressed and should be discussed explicitly, including potential explanations.

      Various other factors besides p<sub>N</sub> can influence short-term plasticity, that are the coupling distance, the number of occupied release sites (N<sub>occ</sub>), the replenishment of N<sub>occ</sub> or the recruitment of newly formed N<sub>occ</sub> as well as the expression of endogenous Ca<sup>2+</sup> buffers (Blatow et al., Neuron 2003; Felmy et al., Neuron 2003; Matveev et al., Biophys. J. 2004; Neher, Cell Calcium 1998; Regehr, CSH Perp. Biol. 2012) or fascilitation sensors (Turecek & Regehr, J. Neurosci. 2018; Shin et al., eLife 2025). Traditionally, p<sub>N</sub> had been assumed to have a major impact on short-term plasticity (STP) which is indeed the case at low replenishment rates (e.g. Feldmeyer and Radnikow, J. Physiol. 2009; Zucker and Regehr, Ann. Rev. Physiol. 2002). But at several synapses very fast replenishment rates have been described driving a progressive overfilling of the initial RRP and increasing N<sub>occ</sub> above baseline levels (Brachtendorf et al., Front. Cell. Neurosci. 2015; Doussau et al., eLife 2017; Miki et al., Neuron 2016; Valera et al., J. Neurosci. 2012) making replenishment the stronger determinant of STP.

      Additionally, the size and organization of sub-pools from which vesicle recruitment and release occurs affects the speed and reliability of vesicular release. In our previous study on L5PN-L5PN connections in S1 we found that developmental tightening of CDs was associated with an increase in PPR without altering p<sub>N</sub> (Bornschein et al., Cell Rep. 2019). We could show that the maturation of a replenishment pool during postnatal development increases vesicle recruitment and reliability thereby affecting STP (Bornschein et al., Front. Syn. Neurosci. 2019).

      We discussed this in the revised manuscript.

      “Classically, p<sub>N</sub> was considered as the major determinant of short-term plasticity (e.g. reviewed in Zucker and Regehr, 2002; Feldmeyer and Radnikow, 2009). More recently other factors, including the number of occupied release sites, their replenishment or an increase in their occupancy, or the expression of endogenous Ca<sup>2+</sup> buffers have been considered as more important determinants of short-term plasticity (Rozov et al., 2001; Blatow et al., 2003; Felmy et al., 2003; Matveev et al., 2004; Bornschein et al., 2013; Miki et al., 2016; Doussau et al., 2017; Jackman and Regehr, 2017; Neher and Brose, 2018). (…)”

      Short-term plasticity changes during postnatal development at different cortical PN connections without alterations in p<sub>N</sub> (Reyes and Sakmann, 1999; Bornschein et al., 2019a). For L5PN-to-L5PN connections these differences were found to result from the maturation of an intermediate replenishment vesicle pool (Bornschein et al., 2019).

      (8) There are multiple instances where the text appears to cite non-existent or misnumbered figure panels (e.g., references to "Figure 4G-I / 4J" when the relevant material appears elsewhere). These should be corrected throughout, as they currently reduce readability and confidence.

      We apologize for the misnumbering which originated from a previous version of this manuscript. Figure 4G-J is actually Figure 5A-D. We corrected the references to Figure 5 in the methods section.

      (9) The Methods describe P21-P26 animals, whereas the Results include older cohorts (e.g., P90-P100) and additional regions (e.g., mPFC). The Methods should be updated so that all cohorts and regions analyzed in the Results are fully described.

      Thank you for thoroughly reading the methods. We added the missing cohort (P90-100) and brain region (mPFC) to the method section.

      “C57BL/6J mice at P21-26 and P90-100 of either sex were decapitated under deep Isoflurane (Curamed) inhalation anaesthesia. (…) Coronal neocortical slices (150-250 μm thick) were cut from the lateral PFC, medial PFC (mPFC) or S1 region (Figure 1A) with a vibratome (HM 650 V, Microm).”

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      These are mostly points for discussion; there is no need for additional experiments.

      (1) Discuss potential effects of (re-)patching on presynaptic physiology and the "control" time course in Figure 2A.

      The repatching strategy is a strength, but it also introduces opportunities for physiological drift (dialysis effects, changes in access resistance, altered excitability, and the switch to presynaptic stimulation in whole-cell after repatching). Please discuss (and, if already quantified, briefly report) why EPSC amplitudes appear to increase in the control time course after repatching (as seen in Figure 2A). Even a short explanation (e.g., run-up after whole-cell access, recovery from on-cell stimulation, washout of endogenous buffering, improved spike waveform reliability, etc.) plus reassurance that baseline stationarity criteria were met would strengthen confidence in the repatching-based inference.

      Baseline recordings were usually performed with the presynaptic neuron in cell-attached mode. In this configuration intracellular ion concentration, second messenger systems as well as cell-specific resting membrane potential stay essentially unaffected. When switching to whole-cell mode after re-patching of the presynaptic cell, the pipette solution determines the intracellular environment, which may also affect second messenger systems and mobile endogenous buffers will be washed out. Although layer 5 pyramidal neurons do not express relevant concentrations of these mobile buffers (Helmchen et al., Biophys. J., 1996; Tran and Stricker, Biophys. J., 2018; Bornschein et al., Cell Rep., 2019), this might have contributed to the moderate and temporary run-up after whole-cell access to presynaptic cells.

      In postsynaptic neurons, we also routinely controlled the stability of R<sub>s</sub> and I<sub>leak</sub> during our prolonged measurements. In the controls, we found a median initial increase in relative EPSC amplitudes to 1.17 (1.02-1.38) between 0 and 10 min after the presynaptic whole-cell access was established, which correlated with a temporary decline in R<sub>s</sub> in 3 out of 6 recordings. EPSC amplitudes returned to their baseline values after 20 min at the latest (1.02, 0.75-1.20). We added potential reasons for the temporary amplitude increase in the controls in the results section.

      “In control recordings a temporary initial increase in EPSC amplitudes after whole-cell access to the presynaptic neuron was evident. Although L5PNs do not express large concentrations of mobile buffers (Helmchen et al., 1996; Tran and Stricker, 2018; Bornschein et al., 2019b), their wash-out might have contributed to the temporary run-up. On the other hand, run-up correlated with a temporary decline in R<sub>s</sub> in some recordings (3 out of 6). Run-up effects normalized after 20 min at the latest (1.02, 0.75-1.20).”

      (2) Clarify interpretation/robustness of MPFA-derived "N" given the large range, and discuss uncertainty from using only three [Ca<sup>2+</sup>]e conditions for L2/3-L5PN MPFA.

      The manuscript reports that "N" is markedly smaller in PFC than S1 (median ~2.1 vs 8), but the S1 range is very broad (3-19).

      (a) Briefly discuss whether/why such a wide N range is expected and how it should be interpreted (binomial "N" vs anatomical release sites; sensitivity to CV assumptions; potential dependence on connection geometry, bouton number, dendritic filtering, etc.).

      (b) Add a short statement on uncertainty/identifiability when fitting MPFA with only three conditions for L2/3-L5PN (e.g., whether confidence intervals/bootstraps were examined; how stable q and N are to small changes in the variance estimates). Even a qualitative note would help readers judge how much weight to put on the absolute N estimates versus the overall cross-area trend.

      (a) In our previous study on this connection we determined a wide range for N (8, 3-19) even though we used four extracellular Ca<sup>2+</sup> concentrations in MPFA (Bornschein et al., Cell Rep. 2019). The binomial parameter N can be considered to represent the number of release sites, including empty release sites (see Brachtendorf et al., Front. Cell. Neurosci. 2025). The range of release sites is likely to reflect the variability in the number of anatomical synaptic contacts ranging from 1 to 6 for these synapses (Frick et al., Cereb. Cortex 2008). Since 1 to 3 active zones/release sites per synaptic contact appear to be typical for small cortical synapses (e.g. Xu-Friedman et al., J. Neurosci. 2001), this results in a wide range of 1-18 release sites per connection.

      In our previous study (Bornschein et al., Cell Rep. 2019), we also investigated the effects of different values of CV1 and CV2 by repeating the MPFA fitting procedures for different combinations of CV1 and CV2 ranging from 0.1 to 1 each. We have quantified a deviation of ≤10% in the estimates of vesicular release probability from the typically used CV values of 0.3 across a wide range of CV value combinations (see also Schmidt et al., Curr. Biol. 2013). We have added a note regarding CV sensitivity in the methods section.

      “For CV assumptions that deviate from the standard value of 0.3, deviations in the calculated p<sub>N</sub> values of less than 10% are to be expected (Schmidt et al., 2013; Bornschein et al., 2019b).”

      (b) The three Ca<sup>2+</sup> concentrations we used for MPFA resulted in a low (<0.5), a medium (~0.5) and a large (>0.5) p<sub>N</sub> condition. With this, a parabola is uniquely determined by three parameters. To further ensure the reliability of the parameters determined by MPFA, we compared them to values estimated from EPSC amplitudes (EPSC = N p<sub>N</sub> q; PFC, 6 pA; S1, 48 pA) and failure rates (F = (1-p<sub>N</sub>)^N; PFC, 0.25, S1, 0.0001), which yielded values similar to those from MPFA (EPSCs in PFC: 8 pA, 5-15 pA, and S1: 29 pA, 18-53 pA; failure rates in PFC: 0.16, 0.08-0.28, and S1: 0, 0-0.03; see original manuscript).

      To further support this, we have now conducted additional experiments using four Ca<sup>2+</sup> concentrations. The results are consistent with those from the experiments using three concentrations. The additional experiments are shown for review purposes in Figure R1. The novel p<sub>N</sub> data are included in the summary of p<sub>N</sub> values (now n=6) in the results section and in Figure 3F.

      Additionally, we performed bootstrap analyses with 10,000 replicates which were generated with replacement from the original data sample. Distributions of bootstrap 25% trimmed means showed a clear separation for N and q between PFC and S1 but not for p<sub>N</sub>. The bootstrap results are shown for review purposes in Author response image 2 (cf. point 5 of Reviewer#2).

      (3) Broaden the discussion, e.g. by linking to nanodomain/microdomain coupling as a general strategy for stimulus encoding, including sensory periphery examples.

      The work will resonate beyond the cortex if the authors explicitly connect their findings to broader principles: how the spatial coupling regime shapes the transfer function between Ca<sup>2+</sup> entry and vesicle fusion, thereby tuning reliability, timing, and dynamic range. Requested addition: Please consider adding a short subsection discussing analogous implementations in the sensory periphery, especially ribbon synapses of cochlear inner hair cells and rod photoreceptors, where nanodomain coupling has been discussed as a key determinant of encoding and release dynamics. Also, citing relevant work such as that by Scimemi and Diamond 2012 would further strengthen the paper.

      We agree that the work by Scimemi and Diamond (J.Neurosci. 2012) is important and we cited and discussed their work in several of our previous publications. We now also included the paper in the revised version of the present manuscript. As requested, we also included a discussion on findings from ribbon type synapses and also from the neuromuscular junction.

      “Nanodomain coupling was also found in the peripheral nervous system, in particular at retinal (Singer and Diamond, 2003; Jarsky et al., 2010) and auditory (Moser and Beutner, 2000; Brandt et al., 2005) ribbon-type synapses and at the neuromuscular junction (Harlow et al., 2001; Shahrezaei et al., 2006). These synapses have highly specialized properties and appear to be optimized for very reliable transmission and, in the case of ribbon synapses, also for high-frequency coding of sensory information (reviewed in Matthews and Fuchs, 2010; Eggermann et al., 2012). Thus, it appears that synapses in the sensory pathways, in particular those engaged in reliable high-frequency coding of sensory information, both in the periphery and in the lower processing stages of the CNS, up to primary sensory cortices, operate with nanodomain coupling. In the executing motor pathway, the neuromuscular junction uses nanodomain coupling and, as recent results from our group suggest, also PNs in the primary motor cortex (Yarim et al., in preparation). It is tempting to speculate that complete loops from or to the primary cortices to their peripheral target organs operate with nanodomains. Microdomain coupling, on the other hand, appears to come into play only if integration of information from multiple sources and plasticity are the main focus, as at certain synapses in PFC (this study) or hippocampus (Vyleta and Jonas, 2014).”

      “The microdomain was assumed to be formed by a ring-like structure of VGCCs around a vesicle (Figure 5D). This topography was chosen because such a microdomain was found to best predict the experimental data of transmitter release from PNs in young S1 (Bornschein et al., 2019b). Other previously described distributions of VGCCs suitable to reproduce release data cover random distributions of VGCCs (Scimemi and Diamond, 2012), VGCC clusters (Meinrenken et al., 2002; Nakamura et al., 2015; Rebola et al., 2019), and exclusion zones (Keller et al., 2015; Rebola et al., 2019). In the early S1, all of these models predicted a higher EGTA sensitivity of the microdomain, however, these models provided a poorer fit to the full set of the experimental data than the ring-like structure (Bornschein et al., 2019b). Since the experimental data from PNs in the mature PFC were similar to those in young S1, these other microdomain models were not tested explicitly here.”

      “(…) Finally, the increase in the PPR induced by the application of Cd<sup>2+</sup> further supports our conclusion of microdomain coupling in the PFC synapses (Scimemi and Diamond, 2012).”

      (4) Address limitations of basal/resting Ca<sup>2+</sup> estimates and make explicit that measured Ca<sup>2+</sup> signals are volume-averaged (not microdomain) readouts.

      (a) The reported basal [Ca<sup>2+</sup>]i values are in the ~tens of nM range. Given the stated in vitro KD for Fluo-5F in the authors' pipette solution (439 nM), the resting estimates are far below KD; this does not invalidate the approach, but it does warrant a brief discussion of sensitivity/uncertainty (influence of Rmin estimation, background subtraction, and how errors propagate into basal [Ca<sup>2+</sup>]i). Repeating experiments is not necessary-just clearer framing of limitations.

      (b) Please also emphasize more prominently (ideally in Results and/or Discussion) that the bouton signals are volume averaged and therefore do not directly report calcium microdomains at active zones or nanodomains at release sensors. The Methods already state this point; echoing it in the main text would prevent over-interpretation by readers.

      (a) We agree that Fluo5F is less suitable for determining absolute basal calcium levels. In a previous study (Bornschein et al., Science 2025) we determined the basal Ca<sup>2+</sup> concentration with OGB1 (K<sub>D</sub>=166 nM; basal [Ca<sup>2+</sup>]<sub>i</sub>=44 nM, 24-58 nM, n=43 boutons from 10 cells) and observed no significant difference to basal [Ca<sup>2+</sup>]<sub>i</sub> values determined with Fluo5F despite the K<sub>D</sub> of 439 nM (31 nM, 16-54 nM, 14 boutons from 3 cells; P=0.204, MWU; data not published). We added this limitation to the results section and swapped Figure panels 4E and F for confluence. The calibration curve of Fluo5F as well as the comparison to basal [Ca<sup>2+</sup>]<sub>i</sub> values determined with OGB1have been included in Figure S4.

      “The quantification of absolute basal [Ca<sup>2+</sup>]<sub>i</sub> was limited by the K<sub>D</sub> of Fluo5F (439 nM), which slightly underestimated basal [Ca<sup>2+</sup>]<sub>i</sub> values in comparison to quantification with OGB1 (K<sub>D</sub>=166 nM, Figure S4). Nevertheless, relative comparison of basal [Ca<sup>2+</sup>]<sub>i</sub> yielded no significant differences between PFC (30 nM, 21-34 nM) and S1 (22 nM, 13-38 nM; Figure 4F).”

      (b) In the results section, we have now emphasized that volume-averaged Ca<sup>2+</sup> signals were measured.

      “We performed dual-dye two-photon Ca<sup>2+</sup> imaging (Sabatini et al., 2002) to quantify volume-averaged Ca<sup>2+</sup> signals at presumed presynaptic boutons located on axon collaterals of L5PNs in PFC and in S1.”

      Reviewer #2 (Recommendations for the authors):

      (1) For a meaningful comparison, recordings from the PFC and the S1 cortex of the same animals should be undertaken. Additionally, I suggest performing additional experiments regarding the different cell types of L5 pyramidal cells in layer 5a.

      We performed new experiments to determine EGTA sensitivity in PFC and S1 from the same animal. The results from these experiments agree with the previous results. They are included in the results section , in Figure 2C-F and in Figure S3A-C.

      Additionally, we extended the discussion on the examined cell types. For S1 cortex we refer in more detail to our previous work, where we described in depth where and under consideration of which criteria our recordings were established and that based on these criteria we recorded from pyramidal neurons in layer 5A in S1 (Bornschein et al., Cell Rep. 2019; Bornschein et al. Front. Synapt. Neurosci. 2019; Bornschein et al., Science 2025). Within layer 5A, we did not attempt to further distinguish between types of pyramidal neurons. We include this in the methods section.

      “Patch-clamp recordings from L5PNs located in the upper layer 5 (L5A in S1) were established according to the criteria described in detail in our previous work on this connection in S1 (Bornschein et al., 2019b; Bornschein et al., 2025). Presynaptic neurons were stimulated extracellularly in upper layer 2/3 (L2/3-L5PN connections) straight above the patched L5PN or in on-cell mode in L5A right next to the postsynaptic cell (L5PN-L5PN connections; Figure 1).”

      For the recordings in PFC and heterogeneity in pyramidal neuron types we refer to our detailed response to the point 3 of Reviewer 2. There we also discuss that the heterogeneity in morphology and spiking patterns is probably not reflected on the synaptic level. We would also like to emphasize that the type of experiments we perform with paired recordings and long-lasting patch-clamp measurements is not suitable to differentiate between subpopulations of pyramidal neurons. This would require successful recordings form several tens of different pyramidal neurons, which is not feasible in our type of experiment. We discuss this limitation of the discussion.

      (2) The authors need to comment in depth on their MPFA data, and if feasibl,e perform additional experiments.

      Concerning the robustness of quantification of synaptic parameters by MPFA, we refer to our comments on point 2b of the recommendations for the authors to Reviewer 1. Additionally, we performed new MPFA experiments with four extracellular Ca<sup>2+</sup> concentrations that agree with our results with three Ca<sup>2+</sup> concentrations.

      (3) The statistical analysis should be revised and a test for the normality of data distribution should be implemented.

      A test for normal distribution (Shapiro-Wilk test) has already been described in the methods section in the previous version of this manuscript.

      (3) Figure 1A is somewhat misleading because it could suggest that the authors have performed dual recordings in identified PFC pyramidal cells.

      We added “L2/3 or L5” to the stimulation panel of Figure 1A to illustrate that we stimulated either extracellularly in L2/3 or L5PNs directly via the patch pipette.

      (4) Is the relative variance of the mean EPSC amplitude and latency between connections larger in the PFC connections than in S1 cortex? This could indicate a variability in cell types.

      The relative variance of EPSC amplitudes calculated as median absolute deviation (MAD) was 0.46 in PFC and 0.50 in S1 arguing against differences in the variability in cell types. The larger variability in delays expressed as SD<sub>Delay</sub> is the result of the larger coupling distance in PFC compared to S1 (Bullmann et al., J. Neurosci. 2024). Consequently, also the relative MAD is larger (0.89) in PFC compared to S1 (0.14) and is therefore not able to detect differences in the variability of recorded cell types.

      (5) Reyes and Sakmann (1999) have previously described differences for L2/3-L5b and L5b-L5b synaptic connections in S1 cortex at different developmental stages. This paper needs to be cited as it is highly relevant to this study.

      Reyes and Sakmann (J. Neurosci. 1999) reported layer-specific differences in short-term plasticity in young sensorimotor cortex which disappeared as maturation progressed and short-term plasticity increased. In a previous study (Bornschein et al., Front. Syn. Neurosci. 2019) we also described a developmentally driven increase in short-term plasticity caused by the maturation of vesicle pools. In the present study we used mature animals and would therefore not expect layer-specific differences neither in S1 nor in PFC since the time course of postnatal maturation was described to be comparable in both neocortical circuits (Kroon et al., Sci. Rep. 2019).

      We discussed this paper in the context of developmental changes in short-term plasticity.

      “Short-term plasticity changes during postnatal development at different cortical PN connections without alterations in p<sub>N</sub> (Reyes and Sakmann, 1999; Bornschein et al., 2019a). For L5PN-to-L5PN connections these differences were found to result from the maturation of an intermediate replenishment vesicle pool (Bornschein et al., 2019a). Such pool maturation may also underlie the elimination of layer-specific differences in short-term plasticity between L2/3-L5B and L5B-L5B synaptic connections that were evident in young rats but eliminated during the first weeks of postnatal development (Reyes and Sakmann, 1999). Since postnatal development follows the same time course in mouse PFC and S1 and occurs predominantly within the first two weeks after birth, significant layer-specific differences in PN synapses are unlikely in both areas in our experimental time window (Kroon et al., 2019). Consistently, we found similar PPRs at L2/3-L5PN synapses and L5PN-L5PN synapses in both areas, with facilitation in PFC and depression in S1, irrespective of the presynaptic PN synapse type.

      (6) Regarding the point of loose or tight Ca<sup>2+</sup> channel coupling: Could some of the differences result from differences in the presynaptic Ca<sup>2+</sup> channel complement? Please comment.

      This can be excluded. Ca<sub>v</sub>2.1 and Ca<sub>v</sub>2.2 are the main channels gating release at PN synapses. The gating kinetics of these channels are very similar and they only differ somewhat in their peak current amplitude (Bornschein et al., Cell Rep. 2019, Figure 4). Since the number of open channels is a fit parameter in our simulations there would only be an effect on the estimate of the number of channels gating release but not for the estimate of the coupling distance. This is all the more true since the EGTA effect depends on the diffusion distance rather than on the gating kinetics. These considerations will also hold for Ca<sub>v</sub>2.3 channels, which have slower closing kinetics, but anyway play only a very minor role for triggering release.

    1. eLife Assessment

      This is a valuable contribution to influenza surveillance. After revisions, addressing methodological and reporting issues, the claims of providing an early warning tool now align more with the reported study results, and ultimately, the evidence presented provide a solid addition to the evidence base.

    2. Reviewer #2 (Public review):

      Summary:

      The study aimed to assess the associations between meteorological drivers and influenza is important although not new. The authors used 6 years of surveillance data and deep learning models, combining distributed lag non-linear models (DLNM) with Bayesian-optimized LSTM neural networks for predictive modeling. The key interest in this area is to explore the subtropical locations, where influenza is less common and circulates year-round. The authors further claimed that such an association could be able to provide an early warning in the community.

      Strengths:

      Study design based on a prospective cohort to analyse the data for retrospective outcomes.

      [Editors' note: The Reviewing Editor has assessed the revised article without further input from the original reviewers. The authors have made considerable efforts to address the methodological and reporting concerns raised by the initial peer review, resulting in a substantially revised manuscript.]

    3. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      This study is a valuable contribution to the evidence base. However, the evidence provided is incomplete as the study results only partially support the study conclusions. Addressing the methodological and reporting issues raised by the peer reviewers and properly aligning the claim made for providing a tool for early warning with the study analysis/results would improve the study quality and usefulness of its findings.

      We are deeply encouraged by the editors’ recognition of this study as a valuable contribution to the evidence base. We fully concur with the eLife assessment that the manuscript, in its original form, required substantive methodological recalibration and rhetorical refinement to ensure the claims were strictly supported by the data. We accept these profound critiques unreservedly and have undertaken a comprehensive, ground-up revision of this study.

      This revision was guided by three overarching principles: (1) systematic decoupling of epidemiological confounders, (2) rigorous elimination of COVID-19-era surveillance bias, and (3) precise alignment of our claims with the actual scope of our predictive framework. Most importantly, we have fundamentally recalibrated our core claim: we have entirely removed all assertions of providing a “ready-to-use early warning tool.” Instead, we accurately reposition our contribution as developing a “robust, climate-informed predictive surveillance framework” that establishes an evidence-based foundation for understanding non-stationary climate-influenza dynamics in a subtropical urban setting.

      To ensure our analytical results definitively support this recalibrated conclusion, we executed a comprehensive remodeling of our entire dataset, rebuilding both the DLNM and LSTM networks. The major structural upgrades include:

      (1) Shift to Positivity Rate to Eliminate Testing Bias: To directly address the reviewers’ incisive concern regarding testing volume bias (i.e., raw case counts were artificially inflated or suppressed by massive fluctuations in PCR testing during the COVID-19 pandemic), we have fundamentally replaced “influenza positive case counts” with the “influenza test positivity rates” as our primary outcome metric. This relative metric mathematically standardizes the denominator, elegantly isolating intrinsic viral transmissibility from the severe artificial fluctuations of healthcare-seeking behaviors and diagnostic intensity.

      (2) Integration of New Epidemiological Covariates: We significantly enhanced our control for critical confounders (as suggested by the reviewers) by incorporating new, highly relevant covariates into our models. These include:

      - Weekly detection volumes: To capture residual testing capacity fluctuations.

      - Mask-wearing stringency indices (OWID): To computationally decouple the artificial suppression of cases caused by strict non-pharmaceutical interventions (NPIs).

      - Proportion of the non-local/transient population: To explicitly account for external viral importation risks.

      - Day of the Week (DOW, weekday vs. weekend): To reflect societal mobility and administrative patterns.

      Simultaneously, we removed redundant variables (e.g., raw COVID-19 case numbers) to minimize noise.

      (3) Comprehensive Re-optimization, Retraining, and Sensitivity Analyses: Following this massive feature engineering, we re-executed the Bayesian hyperparameter tuning process and completely retrained the LSTM networks. We also conducted rigorous sensitivity analyses (detailed in the new Table 2 and Supplementary files) by systematically ablating variables like testing volume and mask-wearing indices, definitively proving the necessity of these covariates in our framework. All corresponding codes, main figures (Figures 1– 6), supplementary figures, and statistical tables have been entirely updated.

      (4) Streamlining the Manuscript: We acknowledge the reviewers’ observation that the original methodology was overly meandering. In this revision, we have ruthlessly streamlined the text. We condensed routine laboratory protocols, reorganized the Methods to logically present the “Study area” prior to the “Study design,” and thoroughly rewrote the Discussion to weave our limitations directly into the interpretation of the results. This ensures conciseness, logical flow, and a sharp focus on the central narrative, tailored for the broad and rigorous readership of eLife.

      We believe these deep methodological revisions have profoundly elevated the scientific rigor, interpretability, and integrity of our study. Below, we provide detailed, point-by-point responses outlining how each specific comment was addressed.

      Reviewer #1 (Public review):

      A major concern is that the model is trained in the midst of the COVID-19 pandemic and its associated restrictions and validated on 2023 data. The situation before, during, and after COVID is fluid, and one may not be representative of the other. The situation in 2023 may also not have been normal and reflective of 2024 onward, both in terms of the amount of testing (and positives) and measures taken to prevent the spread of these types of infections. A further worry is that the retrospective prospective split occurred in October 2020, right in the first year of COVID, so it will be impossible to compare both cohorts to assess whether grouping them is sensible.

      We deeply appreciate this astute epidemiological critique. You have precisely identified the most formidable methodological challenge in pandemic-era time series modeling: the profound non-stationarity of influenza dynamics spanning the 2018–2023 timeline, driven by intensive Non-Pharmaceutical Interventions (NPIs), pandemic-related behavioral shifts, volatile testing volumes, and subsequent immunity debt. We fully agree that 2023 represented an atypical post-restriction “rebound” year and is unlikely to be straightforwardly representative of a stabilized post-2024 epidemiological steady state.

      First, we wish to clarify a minor but crucial methodological detail regarding the October 2020 split. You understandably raised concerns about comparing “both cohorts.” We must emphasize that there is no change in the cohort or the data collection methodology. The surveillance system, sentinel hospitals, and diagnostic protocols (managed by Putian CDC) remained identical and uninterrupted from 2018 to 2023. Although the original manuscript labelled the January 2018 – October 2020 and October 2020 – December 2023 phases as “retrospective” and “prospective” cohorts respectively, this distinction reflected only the administrative timing of ethical approval. It does not indicate any changes in sentinel hospital locations, ILI case definitions, specimen collection procedures, RT-PCR diagnostic protocols, or laboratory quality control. The patient population and clinical criteria are completely homogeneous. We have revised the “Ethics statement” subsection to eliminate this semantic confusion.

      Upon receiving this insightful feedback, our team conducted extensive mathematical evaluation on how best to address this timeline heterogeneity. We initially considered formal stratified temporal analyses (e.g., splitting models into strict Pre-COVID 2018-2019, COVID-disruption 2020-2022, and Post-COVID 2023 periods) and implementing a rolling-window validation scheme. However, after careful evaluation, we concluded that strict time-slicing of this particular dataset would introduce its own substantial methodological problems more severe than the heterogeneity it sought to resolve:

      (1) Statistical power constraints on for stratified Non-linear Lag analysis: The DLNM architecture requires continuous, robust longitudinal data to stably estimate two-dimensional exposure–lag–response surfaces with natural cubic splines, which typically demand approximately 16–25 effective degrees of freedom across the joint exposure × lag space. Stratifying our six-year dataset into the three intervals list above would yield:

      - Pre-COVID stratum (2018–2019): ~730 daily observations and only ~200 influenza B positive events, insufficient to stably identify cross-basis surfaces, with confidence intervals expected to widen to non-informative ranges.

      - COVID-disruption stratum (2020–2022): influenza circulation was substantially suppressed (though, importantly, not eliminated, a point we return to below), reducing signal density and risking that estimated surfaces reflect suppression dynamics rather than climate–transmission relationships.

      - Post-restriction stratum (2023): a single year cannot independently support a DLNM with meaningful lag structure given the typical 0–14-day lag window we examine.

      (2) The “training-domain contamination” problem in rolling-window validation: We carefully examined whether expanding-window rolling validation (e.g., train 2018–2020, validate 2021; train 2018–2021, validate 2022; …) would resolve the non-stationarity concern. We concluded it would not, for a subtle but consequential reason: each such window straddles the abrupt NPI transitions, meaning that within any single training window, the model is exposed to a nonstationary mixture of regimes without any explicit signal indicating which regime each observation belongs to. Implementing a rolling window in this specific context forces the algorithm to repeatedly train on fragmented, incomplete phases of this epidemiological cycle. The model is therefore likely to learn a confounded representation in which the climate signal is partially absorbed into the implicit regime-shift signal, a pathology that is in some respects more difficult to diagnose than that of unified modelling. By employing a single, continuous 88% training block (2018–2022), we structurally guarantee that the LSTM’s memory cell is exposed to the complete sequence of regime shifts, from pre-pandemic natural baseline, through extreme suppression, and ending right at the brink of the NPI relaxation. Reserving the entirely unseen 2023 “rebound year” as a chronological hold-out thereby serves as the ultimate extreme stress-test of the network’s capacity to dynamically synthesize NPI relaxation signals and climate variables to forecast a historically unprecedented surge.

      (3) Architectural mismatch between time-slicing and the LSTM’s design principle: A core motivation for adopting an LSTM rather than period-stratified statistical models was precisely to leverage its long-term dependency memory mechanism, which is designed to allow a unified architecture to learn how predictive relationships are modulated by time-varying contextual conditions. Presegmenting the data into NPI-defined strata forecloses this principal architectural advantage and would, in effect, reduce the analysis to a series of disconnected period-specific models, a design for which the LSTM’s complexity provides no benefit over simpler approaches.

      (4) Tension with the surveillance-bias correction: As we discuss in our response to Public review para 2 below, a separate and equally important critique from you concerns surveillance bias from temporally varying testing intensity. Stratified analysis would compound this problem, because each stratum carries a distinct testing-intensity profile (notably the surveillance surge during 2020–2022), and stratum-specific models cannot leverage cross-period testing-volume normalization. A unified model with explicit testing-volume covariates is protective against bias than period-stratified alternatives.

      Our Methodological Solution: Covariate-Driven Adaptive Learning

      Instead of artificially fracturing the timeline, we substantially restructured the analytical framework to teach the model how to contextualize the pandemic-era disruption while preserving longitudinal continuity:

      - First, we shifted the predictive target to Influenza Positivity Rates (laboratory confirmed cases ÷ ILI specimens tested): The positivity rate is the WHO recommended sentinel surveillance metric and intrinsically corrects for the substantial fluctuations in testing volumes, healthcare-seeking behaviour, and surveillance intensity that characterized the pre-pandemic, pandemic, and post-restriction periods, rendering the outcome metric far more comparable across the timeline.

      - Second, we integrated explicit Epidemiological Context Covariates: We structurally upgraded the LSTM network by feeding it vital time-varying covariates alongside meteorological data. Specifically, we incorporated the Our World in Data (OWID) mask-wearing stringency indices, weekly detection volumes, day of the week (DOW, distinguish between weekdays and weekends to capture administrative reporting patterns) and the proportion of non-local residents among tested patients (to account for population mobility and viral importation risks).

      - Third, we performed a covariate-ablation sensitivity analysis: To prove the model actively utilizes these contextual signals, we performed targeted ablation studies (new Table 2). Removing mask-wearing stringency indices increased the 2023 influenza A forecast MAE to 0.012 and the influenza B forecast MAE to 0.003 (compared with MAEs of 0.009 and 0.002, respectively, obtained when the covariate was retained), indicating a decline in the network’s predictive accuracy. Removing the weekly testing volume had an even more catastrophic impact on Influenza B, surging the MAE by 100% and SMAPE by 199.6%. This provides empirical evidence that the unified network does not passively average across regimes, but dynamically leverages NPI and surveillance signals to adapt to non-stationarity. We have added a paragraph to the revised Discussion interpreting these findings as suggestive (though not definitive) evidence that the unified-with covariates architecture is contextualizing rather than averaging across regimes.

      By explicitly providing the LSTM with these explicit contextual parameters, the network autonomously learned the “regime shifts.” It learned that when the mask-wearing index is high, transmission is dampened despite favorable meteorological conditions.

      Re-interpreting the 2023 Validation

      Guided by your critique, we fully agree that 2023 was an anomalous “rebound” year and not a steady-state reflection of a post-2024 “normal.” We have completely abandoned the framing that our model represents a steady-state tool for the future. Instead, we now explicitly frame the 2023 validation as an “extreme epidemiological stress test.” The revised framing rests on three explicit acknowledgements:

      - The 2023 validation tests forecasting performance during a transitional, post-restriction rebound year, and is best understood as an extreme stress test of the framework’s adaptive capacity, not as evidence of long-term predictive validity under stabilized future conditions.

      - Performance metrics observed in 2023 should not be naively extrapolated to 2024 and beyond.

      - Continuous prospective recalibration as 2024–2025 data accumulate, ideally combined with adaptive learning approaches capable of detecting regime shifts in real time, will be essential before any operational deployment.

      The fact that our updated LSTM network accurately forecasted the explosive, atypical 2023 viral rebound, despite being trained heavily on the suppressed 2020-2022 data, demonstrates its robust capacity to synthesize climate variables and NPI relaxation signals.

      We have substantially rewritten the Discussion section to honestly acknowledge this limitation and contextualize the 2023 results, stating that continuous recalibration will be essential for future forecasting.

      We invite you to review the recently included sentences in the relevant subsections as specified above.

      “A major methodological strength of this study lies in its robust, uninterrupted longitudinal data collection framework spanning January 1, 2018, to December 31, 2023. While the analytical timeline encompasses a “retrospective” phase (January 1, 2018 – October 13, 2020) prior to formal ethical approval, and a “prospective” phase thereafter, we emphasize that this distinction represents a purely administrative demarcation regarding the timing of ethical approval. It does not reflect any shift in demographic cohorts, sentinel hospital locations, or data collection methodologies. Importantly, the historical data (2018– 2020) were not subjected to the recall biases or misclassification risks typical of traditional retrospective chart reviews. Rather, they were systematically extracted from a continuously operating, highly standardized public health sentinel surveillance network. From the inception of data collection through the end of 2023, the local CDC maintained absolute uniformity in clinical influenza-like illness (ILI) definitions, nasopharyngeal swabbing procedures, and real-time reverse transcription polymerase chain reaction (RT-PCR) diagnostic assays. Consequently, the pre-2020 data possess the high-fidelity characteristics of a strict prospective cohort, ensuring unparalleled longitudinal consistency and mitigating temporal measurement bias across the entire pre-pandemic, pandemic, and postrestriction timeline.” (Methods, page 8-9)

      Secondly, the interpretation of our framework’s predictive performance during the 2023 validation period requires careful epidemiological and methodological contextualization. The year of 2023 represented an anomalous, post-restriction “rebound” period characterized by rapid NPI relaxation and the release of accumulated population-level immunity debt, resulting in an atypical influenza surge that exceeded pre-pandemic peaks. The framework’s high accuracy across this period should therefore be interpreted as evidence of algorithmic agility and adaptive capacity during a highly volatile transitional phase, rather than as definitive proof of long-term predictive validity under a stabilized post-2024 epidemiological regime. Methodologically, while our strict chronological OOT data partitioning prevented temporal information leakage, a critical requirement for LSTM integrity, the reliance on a single, fixed chronological split point (December 31, 2022) intrinsically limits our evaluation to one specific structural break. This fixed-split approach may not exhaustively probe the DLNM-LSTM framework’s resilience against all forms of future epidemiological non-stationarity. Consequently, naive extrapolation of the reported 2023 performance metrics to future surveillance years should be avoided absent prospective recalibration. Future studies should consider employing expanding-window or rolling-origin cross-validation frameworks to provide a more continuous characterization of algorithmic robustness. Continuous integration of accumulating 2024 and 2025 data, combined with adaptive learning architectures capable of detecting regime shifts in real time, will be essential before any operational deployment of this, or similar forecasting frameworks, for routine public health surveillance.” (Discussion, page 44-45)

      We believe this integrated, covariate-based approach maintains the mathematical integrity of the time-series analysis while fully addressing your valid concerns regarding epidemiological non-stationarity.

      The outcome of interest is the number of confirmed influenza cases. This is not only a function of weather, but also of the amount of testing. The amount of testing is also a function of historical patterns. This poses the real risk that the model confirms historical opinions through increased testing in those higher-risk periods. Of course, the models could also be run to see how meteorological factors affect testing and the percentage of positive tests. The results only deal with the number of positive (only the overall number of tests is noted briefly), which means there is no way to assess how reasonable and/or variable these other measures are. This is especially concerning as there was massive testing for respiratory viruses during COVID in many places, possibly including China.

      We are exceptionally grateful for this incisive methodological observation. You have accurately identified a fundamental validity threat, “surveillance intensity bias”, that inherently constrains much of the existing climate–infectious disease literature. We fully concur with your assessment that raw case counts are jointly determined by underlying viral transmission and dynamic testing intensity. Furthermore, we recognize your highly valid concern regarding the risk of “circular reasoning,” whereby meteorological factors might simply trigger higher clinical suspicion and testing rates rather than genuine transmission events—a bias severely exacerbated by the massive respiratory testing surges during the COVID-19 pandemic. To systematically dismantle this threat and address your specific recommendations, we executed a ground-up restructuring of our analytical framework, implementing four complementary strategies:

      (1) Primary outcome redefined as Positivity Rates. Throughout the entire revised manuscript, both the DLNM and LSTM pipelines have been completely re-analyzed using influenza positivity rates (laboratory-confirmed cases ÷ total ILI specimens tested) as the primary outcome, rather than raw case counts. The positivity rate is the WHOrecommended metric for sentinel surveillance precisely because it mathematically standardizes the denominator, normalizing the raw testing volume variability. Its adoption effectively neutralizes the circular-reasoning concern raised by you. All predictive models, Figures 4, 5, and 6, and all primary metrics in Table 2 now reflect positivity-based estimates.

      (2) Weekly detection volume integrated as a dynamic LSTM covariate. Even after positivity-rate normalization, residual testing-intensity effects can persist (e.g., if testing patterns shift among demographic subgroups with systematically different positivity profiles). To capture this, our revised LSTM network incorporated weekly detection volumes as an explicit dynamic input feature. As detailed in our covariate-ablation sensitivity analysis (Table 2), removing the testing-volume covariate drastically degraded the forecasting accuracy for 2023, increasing the Mean Absolute Error (MAE) by 22.2% for influenza A and an astounding 100% for influenza B. This indicates that the model’s predictions are not driven solely by meteorological inputs but are appropriately and dynamically conditioned on the surveillance context.

      (3) SHAP quantification of surveillance bias. By incorporating testing volume into the LSTM, we made the surveillance-bias concern concretely visible and quantifiable. As shown in our new SHAP analysis (Figure 6E-H), weekly_detection emerged as the second most impactful predictor of the positivity rate for influenza B, and, when examining the influenza A results, weekly_detection likewise ranked fifth among the most impactful predictors. This proves that the deep learning algorithm autonomously recognized the profound impact of testing intensity and actively utilized it to dynamically adjust and calibrate its epidemiological forecasts.

      (4) Decoupling “Circular Reasoning” via DLNM Sensitivity Analysis. To definitively prove that our model does not merely “confirm historical opinions through increased testing,” we conducted a targeted DLNM sensitivity analysis (detailed in the new Additional file 4 and Figures S3–S4). We compared DLNM exposure-response curves predicting positivity rates with and without adjusting for weekly testing volumes. The results were striking: the non-linear exposure-response curves and extreme-weather lag patterns remained highly consistent across both models. This empirical stability proves that the identified meteorological drivers represent intrinsic biological/environmental triggers of viral transmission, independent of fluctuating surveillance intensity.

      We are deeply grateful that this critique prompted such substantial methodological refinement. The revised analyses are vastly more robust and epidemiologically interpretable than the original case-count-based version. We invite you to review the recently included sentences in the relevant subsections as specified above.

      “To rigorously address the inherent confounding effects of “surveillance intensity bias”, where fluctuations in raw case counts may merely reflect transient surges in clinical testing capacity rather than true community transmission, the primary outcome metric for all DLNM modeling was mathematically defined as the influenza positive rate, with meteorological factors and weekly detection volume serving as independent variables. The formula for calculating the daily influenza positivity rate is as follows:

      where R represents the daily influenza positivity rate, I represents the number of daily influenza positive cases, and N denotes the total number of daily influenza tests performed. By adopting this WHO-recommended surveillance metric, our our analytical framework explicitly standardizes the epidemiological denominator. This mathematical normalization effectively neutralizes the severe surveillance intensity bias caused by dramatic testing volume surges during the COVID-19 pandemic, ensuring that our models capture intrinsic viral transmissibility rather than artificial fluctuations in healthcare-seeking behavior or diagnostic capacity.” (Methods, page 16)

      “Furthermore, to evaluate the robustness of our findings against testing intensity, we conducted a targeted sensitivity analysis by reconstructing the DLNM models without the “weekly detection volumes” covariate. We then compared the non-linear cumulative risks and extreme weather lag effects between these ablated models and the original fully adjusted models. This allowed us to determine whether the identified climate-transmission associations were stable and biologically intrinsic, or merely artifacts of weather-correlated testing behaviors (Supplementary Information Additional file 4).” (Methods, page 17-18)

      “To assess the contribution of pandemic-related confounding variables, we performed targeted sensitivity analyses on the LSTM networks. Specifically, we sequentially removed covariates of mask-wearing stringency indices and weekly detection volumes from the input features while maintaining identical Bayesian-optimized hyperparameters. The performance of these ablated models was evaluated using MAE, RMSE, MAPE, and SMAPE, allowing us to quantify the exact necessity of incorporating testing and behavioral covariates in forecasting models during periods of epidemiological non-stationarity.” (Methods, page 22)

      “Crucially, this LSTM stage explicitly incorporates weekly detection volumes, maskwearing stringency indices, non-local population proportion, and DOW effects alongside meteorological inputs, and uses influenza positivity rates rather than absolute case counts as the modeling endpoint. Together, these design choices are intended to mitigate, rather than fully eliminate, the surveillance-related biases that can distort count-based forecasting during periods of fluctuating testing intensity.” (Discussion, page 40)

      “This study has several limitations that should be considered when interpreting the findings. Firstly, our primary analysis relies on influenza surveillance data collected from seven sentinel hospitals in Putian, which inherently captures only a fraction of all influenza cases occurring in the broader community. Although employing positivity rates as the primary outcome substantially mitigates the surveillance bias inherent in count-based analyses, residual selection effects may persist if testing patterns shift differentially across demographic subgroups with systematically divergent positivity profiles. Our inclusion of weekly detection volumes as an explicit LSTM covariate, complemented by parallel DLNM sensitivity analyses validating the independence of meteorological effects from testing volumes, were designed to characterize and partially account for this residual bias, but cannot fully eliminate it. Consequently, positivity rates likely underestimate the true burden of community influenza infection. Although the surveillance infrastructure in Putian remained uniform and uninterrupted throughout the 2018–2023 timeline, mitigating measurement bias and temporal confounding risks, the transferability of our findings to other global subtropical regions with different socioeconomic structures, healthcare systems, or population behaviors requires cautious, region-specific calibration. Accordingly, our framework should be strictly interpreted as a high-fidelity tool designed to forecast the observable public health surveillance signal, which is the most operationally relevant target for public health agencies, rather than for estimating unobserved, absolute community disease burden.” (Discussion, page 43-44)

      “This limitation was particularly exacerbated by the profound epidemiological disruptions during the COVID-19 pandemic, where raw numbers of confirmed cases became heavily confounded by surveillance intensity (i.e., fluctuating testing volumes) rather than solely reflecting underlying viral transmission. Our adoption of influenza positivity rates as the primary modeling endpoint and incorporation of weekly detection volumes, face-covering stringency, and non-local population proportion as dynamic covariates, rigorously mitigated these aggregate-level biases and linked our methodological design directly to the forecasting outcomes. As unequivocally demonstrated by our covariate-ablation sensitivity analyses (Table 2), failing to account for mask mandates and testing volumes leads to severe, mathematically predictable deviations in absolute forecasting accuracy. Furthermore, while daily case counts in a single city can occasionally be small, sporadic, and driven by external importations, our incorporation of non-local population proportion effectively adjusted for these localized importation risks. Consequently, although our findings characterize population-level associations between meteorological factors and influenza activity, they should not be interpreted as evidence of micro-level causal mechanisms at the individual patient level. Ultimately, the DLNM-LSTM framework’s robust performance across the non-stationary transition out of NPI policies highlights the absolute necessity of integrating behavioral and virological baseline metrics into future climate-driven predictive surveillance systems.” (Discussion, page 45-46)

      We are deeply grateful that this critique prompted such substantial methodological refinement. The revised analyses are vastly more robust and epidemiologically interpretable than the original case-count-based version.

      (1) Although the authors note a correlation between influenza and the weather factors. The authors do not discuss some of the high correlations between weather factors (e.g., solar radiation and UV index). Because of the many weather factors, those plots are hard to parse.

      We sincerely appreciate your constructive feedback regarding the visual clarity of the correlation plots and the methodological implications of highly correlated meteorological variables. We agree that the original 11 × 11 scatterplot matrix was visually overwhelming and could obscure the statistical implications of highly correlated features (such as solar radiation and the UV index, Pearson’s r > 0.9).

      To address this, we have taken two specific revisions:

      (1) Improved Visualization (Revised Supplementary Figure S2): To make complex relationships easier to parse, we have completely redesigned Supplementary Figure S2. It is now logically partitioned into two distinct visual components:

      - Panel A features a high-contrast Pearson correlation heatmap, where color intensity and statistical significance asterisks allow for rapid, intuitive identification of highly correlated pairs (such as solar radiation and UV index).

      - Panel B retains the pairwise scatterplots (lower-left) and density distribution curves (diagonal) to facilitate the detailed visual inspection of non-linear trends, data skewness, and potential anomalies.

      (2) Methodological Defense on Potential Collinearity: We did not arbitrarily eliminate these highly correlated variables because our two-stage modeling approach inherently mitigates multicollinearity risks associated with potential collinearity:

      - For the DLNM analysis: To prevent coefficient instability caused by multicollinearity, all DLNM lagged analyses strictly employed univariate exposure-response models for meteorological factors (i.e., evaluating the relationship between a single meteorological factor and influenza positivity rates at a time, while controlling for long-term temporal trends and including other covariates as fixed terms). Consequently, the independent effect sizes and lag structures derived from the DLNMs are completely unaffected by inter-variable correlations.

      - For the LSTM network: Unlike traditional multiple linear regression (or ARIMA) where potential collinearity inflates standard errors and destabilizes coefficients, Deep Learning architectures (LSTM) are natively robust to redundant features. The network’s non-linear activation functions and gating mechanisms naturally weight overlapping signals during the optimization process, effectively using redundant variables as a form of algorithmic regularization without compromising predictive stability.

      We have expanded the Statistical Analysis and Discussion sections of the manuscript to explicitly articulate this rationale, ensuring maximum methodological transparency.

      We invite you to review the recently included sentences in the relevant subsections as specified above.

      “Prior to analytical modeling, we assessed the pairwise associations and potential collinearity among all meteorological variables using a comprehensive correlation heatmap and scatterplot matrix (Supplementary Figure S2). Notably, certain variables, such as solar radiation and the UV index, exhibited high positive correlations. To avoid coefficient instability typically caused by multicollinearity in regression models, all DLNM analyses adopted a univariate approach for meteorological factors, sequentially evaluating the nonlinear and lagged effects of individual meteorological predictors while adjusting for time trends and including other fixed covariates. Conversely, all variables were retained during the LSTM forecasting phase, as the non-linear gating architecture of recurrent neural networks inherently exhibits robust regularization against potential collinearity among input features.” (Results, page 26)

      “Furthermore, extensive environmental inputs inevitably introduce severe collinearity, such as the strongly correlated solar radiation and UV index. While traditional multivariate models are highly vulnerable to such overlapping variances, the recurrent, weighted representation learned by the LSTM is comparatively tolerant of such redundancy, allowing broader covariate integration than in previous efforts.” (Discussion, page 40)

      We hope that our responses and revisions will meet your expectations and demonstrate our dedication to improving the scientific quality of this study.

      (2) The authors do not actually compare the results of both methods and what the LSTM adds.

      We sincerely appreciate your perceptive critique. We agree that our initial manuscript lacked a sufficiently rigorous, side-by-side comparison, and more importantly, it failed to adequately articulate why the LSTM architecture succeeds where classical statistical baselines fail.

      To rectify this, we have comprehensively overhauled the comparative analysis. We formulated our baseline as a multivariate ARIMA model, supplying it with the exact same multidimensional covariate matrix as the LSTM (including meteorological variables, mask-wearing indices, and weekly testing volumes), focusing on three dimensions:

      (1) Direct Quantitative Comparison: Following the restructuring of our outcome variable to the Influenza Positivity Rate, we re-evaluated both models using identical training (2018–2022) and validation (2023) datasets. The LSTM consistently and substantially outperformed the ARIMA model across all metrics. For Influenza A, the LSTM achieved an MAE of 0.009 and RMSE of 0.035 (vs. ARIMA: MAE 0.136, RMSE 0.238). For Influenza B, the LSTM yielded an MAE of 0.002 and RMSE of 0.011 (vs. ARIMA: MAE 0.049, RMSE 0.057). We have updated Table 2 and the corresponding Results section to explicitly present these side-by-side comparisons.

      (2) Visualizing “Where” the LSTM Outperforms: We have updated Supplementary Figure S3 to map the ARIMA predictions for the 2023 positivity rate, allowing for a direct visual comparison with the LSTM predictions in Figure 6 (A, B). The visualizations explicitly reveal the ARIMA model’s fundamental limitation: it tends to predict relatively flat or conservatively smoothed values, failing entirely to capture the extreme, explosive non-linear peaks of the 2023 viral rebound. Conversely, the LSTM network accurately tracks these sudden epidemic phase transitions.

      (3) Articulating “What the LSTM Adds” (Discussion Expansion): We have significantly expanded the Discussion section to intellectually articulate why the LSTM succeeds where ARIMA fails. To ensure a strictly fair methodological comparison, we formulated our baseline as a multivariate ARIMA model (ARIMAX), supplying it with the exact same multidimensional covariate matrix as the LSTM (including meteorological variables, mask-wearing indices, and weekly testing volumes). Therefore, the LSTM’s superior performance is not due to information asymmetry (i.e., it did not “see” more variables), but stems directly from its algorithmic architecture. The consistent superiority of the LSTM over both the linear sequential baseline (ARIMA) and the non-linear non-sequential baseline (XGBoost) isolates the recurrent gated architecture itself as the source of the predictive gain. This advantage arises from four architectural properties intrinsic to recurrent gated networks yet absent in both tree ensembles and linear autoregressive models:

      a) Sensitivity to temporal ordering, which decision-tree splits and linear regressors cannot natively encode;

      b) Gated propagation of long-range dependencies through forget–input–output mechanisms;

      c) Explicit accommodation of serial autocorrelation, violated by the i.i.d. assumptions underlying gradient-boosted trees;

      d) Paradigmatic comparability with sequential statistical models: the joint failure of ARIMA (linear, sequential) and XGBoost (non-linear, non-sequential) isolates recurrent sequential memory, rather than non-linearity per se, as the critical feature for forecasting under pandemic-era non-stationarity.

      We emphasize that this interpretation applies specifically to the present non-stationary epidemiological forecasting task and does not constitute a general dismissal of gradient-boosted ensembles, which retain competitive performance across many structured prediction domains.

      We invite you to review the recently revised table, figures and sentences in the relevant subsections as specified above.

      “The LSTM networks accurately captured both the timing and magnitude of these nonlinear epidemic surges, including the two outbreak peaks of influenza A during February-March and November-December of 2023, as well as the peak of influenza B in November-December, demonstrating good predictive performance. The predictive performance of the LSTM networks was quantitatively assessed using metrics such as MAE, RMSE, MAPE, and SMAPE. For influenza A, the MAE was 0.009, RMSE was 0.035, MAPE was 0.158, and SMAPE was 0.521; for influenza B, the MAE was 0.002, RMSE was 0.011, MAPE was 0.17, and SMAPE was 0.484 (Table 2).” (Results, page 31)

      “To benchmark the predictive value added by the LSTM architecture, we constructed two baseline models using identical training (2018–2022) and validation (2023) positivity-rate datasets, inclusive of all contextual covariates: the multivariate ARIMA model and the XGBoost gradient-boosting model. The LSTM model demonstrated decisive superiority over both baselines across all evaluated metrics (Supplementary Table S3). For influenza A, the LSTM achieved an MAE of 0.009 and SMAPE of 0.521, compared with the ARIMA’s MAE of 0.136 (SMAPE 1.212) and XGBoost’s MAE of 0.138 (SMAPE 1.081), representing approximately 15-fold error reductions relative to both benchmarks. For influenza B, the LSTM yielded an MAE of 0.002 (SMAPE 0.484), compared with 0.049 (SMAPE 0.810) for ARIMA and 0.070 (SMAPE 0.737) for XGBoost, approximately 25- to 35-fold reductions. Notably, ARIMA and XGBoost produced errors of comparable magnitude despite their disparate assumptions regarding linearity, suggesting that the LSTM’s advantage derives not from non-linear modeling capacity per se, but from architectural properties specific to recurrent sequential processing.

      Beyond global error metrics, visual comparison of forecasting trajectories (Figure 6; Supplementary Figures S5 and S6) reveals critical behavioral disparities. For influenza A, both the linear ARIMA model and the non-linear XGBoost model generated conservatively smoothed forecasts that entirely failed to capture the sudden, explosive peaks of the 2023 post-restriction viral rebound. For influenza B, ARIMA produced continuous spurious fluctuations during non-epidemic periods (when true positivity was near zero), likely overreacting to covariate variations, while XGBoost generated largely flat trajectories that missed the mid-year outbreak peak. In stark contrast, the LSTM’s recurrent gating mechanisms successfully filtered out covariate noise during low-transmission periods while accurately tracking extreme epidemiological phase transitions, demonstrating a qualitative advantage in handling non-stationary regime shifts.” (Results, page 33-34)

      “To rigorously isolate the predictive contribution attributable to the LSTM’s recurrent architecture, we benchmarked it against two covariate-matched baselines representing distinct methodological paradigms: the multivariate ARIMA model and the XGBoost gradient-boosting ensemble. This design controls simultaneously for linearity (ARIMA→LSTM contrast) and for non-linearity without recurrent memory (XGBoost→LSTM contrast), allowing us to attribute observed performance gains to specific architectural inductive biases rather than to informational asymmetry or model non-linearity in general. The LSTM substantially outperformed both baselines on the 2023 validation window. For influenza A, it achieved an MAE of 0.009, compared with 0.136 for ARIMA and 0.138 for XGBoost, approximately 15-fold reductions. For influenza B, the LSTM yielded an MAE of 0.002, versus 0.049 for ARIMA and 0.070 for XGBoost, 25- to 35-fold reductions. Critically, ARIMA and XGBoost produced errors of comparable magnitude despite their disparate assumptions regarding linearity, and both systematically under-predicted the explosive 2023 post-NPI rebound for influenza A while generating flat or spurious trajectories for influenza B (Supplementary Figures S5–S6). That parallel failure indicates the LSTM’s advantage under pandemic-era non-stationarity derives not from non-linearity per se, but from four architectural properties intrinsic to recurrent gated networks yet absent in tree ensembles: (i) threshold-based, sequence-insensitive splits fail to encode present–past dynamics; (ii) XGBoost lacks forget–input–output gates to propagate and re-weight historical states across time lags; (iii) gradient-boosted trees assume near-independence, contradicting the pronounced temporal autocorrelation in epidemiological time series; (iv) ARIMA and LSTM are sequential models with different functional forms, whereas XGBoost is non-sequential. The joint failure of ARIMA and XGBoost, despite differing non-linear treatment, isolates recurrent sequential memory as the critical architectural feature for forecasting under non-stationarity. We emphasize that this interpretation applies specifically to the present non-stationary influenza forecasting task and does not constitute a general dismissal of gradient-boosted ensembles, which retain state-of-the-art performance across many structured prediction domains. Rather, it highlights that for surveillance time series exhibiting pronounced temporal dependencies and abrupt regime shifts, such as the 2023 post-NPI rebound, explicit sequential memory becomes functionally essential. This mechanistic reading, together with the LSTM’s comparative edge over previously reported ARIMA-based (Li et al. 2024) and LSTM-based influenza prediction models (Zhu et al. 2022), positions our framework as a substantive methodological advance in predictive modeling for climate-sensitive diseases.” (Discussion, page 41-42)

      Results

      “The MAE for the covariate-adjusted ARIMA model of influenza A was 0.136, RMSE was 0.238, MAPE was 1.115, and SMAPE was 1.212. For influenza B, the MAE was 0.049, RMSE was 0.057, MAPE was 1.426, and SMAPE was 0.810. Overall, the predictive performance of the ARIMA model was substantially inferior to that of the Bayesian optimized LSTM. As illustrated in Figure S5, the linear model struggled significantly with the non-stationary dynamics of the 2023 viral rebound. It either failed entirely to capture the extreme, explosive non-linear peaks (as seen in Influenza A) or generated continuous spurious predictions during zero-case periods due to mechanical linear reactions to covariate inputs (as seen in Influenza B).” (Supplementary information, page 16)”

      Results

      “The XGBoost model produced substantially higher forecasting errors than the LSTM across both influenza subtypes (Table S3; Figure S6). For Influenza A, XGBoost yielded MAE = 0.1379, RMSE = 0.2161, MAPE = 1.8195, and SMAPE = 1.0805, errors of a magnitude broadly comparable to the multivariate ARIMA baseline (MAE = 0.136) and approximately 15-fold higher than the LSTM (MAE = 0.009). For Influenza B, XGBoost achieved MAE = 0.0704, RMSE = 0.0739, MAPE = 1.9172, and SMAPE = 0.7371, exceeding both the ARIMA benchmark (MAE = 0.049) and the LSTM (MAE = 0.002) by approximately 35-fold relative to the latter.” (Supplementary information, page 19)

      We believe these additions provide a much more rigorous and academically satisfying comparative analysis.

      (3) The methods are long and meandering. They could be cleaned up and shortened. E.g., there is no need for 30 lines on PCR testing; the study area should come before the study design. The authors discuss similar elements in multiple places; this whole section can be shortened considerably without affecting the content.

      We sincerely appreciate the reviewer’s editorial guidance. We agree that the initial Methods section was somewhat disjointed and unnecessarily verbose, reflecting iterations from previous drafts. We have completely restructured and aggressively streamlined this section to ensure a logical and concise flow:

      - Structural Reorganization: As suggested, we have moved the “Study area” section to the very beginning of the Methods, providing the geographical and climatic context before detailing the study design and cohort.

      - Consolidation of Redundancies: We have merged the fragmented descriptions regarding ethical approvals, data anonymization, and sentinel hospital protocols into a single, cohesive “Study design and cohort” subsection.

      - Condensing Laboratory Protocols: We completely agree that 30 lines on standard PCR testing in the main text are unnecessary. We have condensed the “Sample collection and pathogen typing” section into a brief, 4-line summary explicitly stating the use of commercial assays (Da’an Gene Co., Ltd) and strict adherence to China CDC guidelines. The highly technical nuances (e.g., RNA extraction integrity, spectrophotometer ratios, and precise PCR amplification thresholds) have been relocated to the Supplementary Information (Additional file 1) for interested readers.

      We invite you to review the recent revisions in the Methods subsection as specified above.

      “Sample collection and pathogen typing

      Respiratory specimens (nasopharyngeal swabs) were collected from ILI patients during their initial visit, prior to treatment, and stored at 4°C in viral transport medium. All samples were delivered to the Putian CDC laboratory within 24 hours of collection. Nucleic acid extraction and one-step real-time fluorescent RT-PCR for influenza A/B subtyping (H1N1, H3N2, Victoria, and Yamagata lineages) were executed within 24 hours upon sample arrival. All assays utilized commercial diagnostic kits (supplied by Da’an Gene Co., Ltd., Guangzhou, China) and were processed in a Biosafety Level 2 (BSL-2) laboratory, strictly adhering to the manufacturer’s instructions and China CDC’s standardized protocols. Detailed laboratory procedures, including RNA integrity parameters and PCR amplification thresholds, are comprehensively documented in the Supplementary Information (Additional file 1).” (Methods, page 13)

      These revisions have significantly improved the readability of the manuscript without sacrificing methodological transparency.

      (4) How reliable is the "Our Word in Data" website for subnational coverage of restrictions? Some of the authors are from Putian and should be able to confirm the accuracy for both studied areas.

      We are very grateful for this insightful question. You astutely identify a common challenge in geospatial epidemiology in China: the systemic lack of publicly accessible, standardized, daily non-pharmaceutical intervention (NPI) datasets at the municipal (subnational) level. Local CDC policy records are typically maintained as internal administrative documents without standardized time-series data interfaces.

      Given this limitation, we utilized the national Our World in Data (OWID) stringency index as a proxy. To directly address the reviewer's excellent point regarding local verification, the co-authors of this study—who are frontline epidemiologists stationed at the Putian CDC and who directed the local pandemic response, including conducted a rigorous, retrospective cross-validation of the OWID index against local realities.

      Our local experts confirmed a high degree of fidelity between the OWID 0–4 scale and the actual policies enforced in Putian and Sanming:

      - 2020 (Initial Outbreak): Both cities enforced “mandatory face coverings outside the home at all times” (OWID Level 4), aligning perfectly with the dataset.

      - 2021–2022 (Normalized Control): Policies shifted to “required in all shared/public spaces” (OWID Level 3), which accurately reflects the local mandates required for public transit, schools, and commercial venues.

      - 2023 (Post-Pandemic Shift): Following the national policy pivot, mandates were downgraded to “recommended” (OWID Level 1), perfectly mirroring local ground truths.

      Therefore, while OWID provides a national-level index, our local CDC authors have empirically verified that its temporal variations accurately capture the intensity of behavioral restrictions experienced by the populations in our specific study areas. We have incorporated a concise statement regarding this expert validation into the revised Methods section.

      “While OWID stringency indices represent national-level policy, publicly accessible and standardized daily NPI datasets at the municipal level are currently unavailable in China. To ensure the spatial validity of these indices, co-authors from Putian CDC, who actively managed the local epidemic response, conducted a rigorous cross-validation. Our local public health experts confirmed that temporal fluctuations of OWID indices (ranging from Level 4 strict mandates in 2020 to Level 1 recommendations in 2023) exhibited high fidelity with the actual, on-the-ground enforcement of NPIs in both Putian and Sanming. Thus, OWID stringency indices serve as a highly reliable contextual proxy for local social contact restrictions.” (Methods, page 15)

      We hope that our responses and the additional statement will meet your expectations.

      (5) Figure 2A is hard to parse; it would make more sense to plot these as line plots (y=count, x=month).

      We completely agree with you. The original heatmap visualization obscured the temporal dynamics of the distinct influenza subtypes. We have entirely redrawn Figure 2A as multi-line plots mapping the monthly incidence trajectories of Influenza A (H1N1, H3N2) and Influenza B (Victoria, Yamagata) across the six-year study period. This new visualization (included in the revised manuscript) vastly improves readability and explicitly highlights the distinct phase shifts and interruptions caused by the pandemic.

      Reviewer #1 (Recommendations for the authors):

      (1) Figure 3 is hard to parse. The flu counts are repeated in every figure, which makes them look very similar. I would recommend that the authors pick 1 or 2 as subpanels for the main paper and put the rest in the supplement.

      We sincerely appreciate your feedback. We completely agree that the original 3x3 square layout severely compressed the x-axis, making the daily temporal fluctuations difficult to parse and causing the flu curves to look visually redundant.

      To resolve this visual clutter without losing valuable environmental context in the main text, we adopted a highly effective structural solution simultaneously suggested by Reviewer #2. We have completely redrawn Figure 3 into an 8x1 vertically stacked format with a single, shared continuous x-axis across the entire study period. This extended aspect ratio vastly expands the timeline, dramatically revealing the highly distinct, daily microfluctuations of each meteorological factor alongside the epidemiological curves. We believe this new layout completely resolves the parsing difficulty you rightly pointed out. Given that a substantial portion of our readership relies heavily on the main text figures for immediate epidemiological context, retaining these cleanly formatted panels in the main manuscript maximizes the paper’s scientific impact. We hope you find this redesigned visualization satisfactory.

      (2) Using a thousand-separator throughout will make the manuscript more readable.

      We completely agree. We have meticulously applied thousand-separators to all relevant numerical values (e.g., 20,488; 17,333) throughout the revised manuscript to enhance readability.

      (3) Line 556 "quantified calculated", pick 1 word.

      We sincerely apologize for this typographical oversight resulting from the drafting process. However, the original sentence that led to the duplicated phrasing you highlighted has been removed, as we had already undertaken a comprehensive revision of the relevant material in that subsection in response to your earlier remarks. We invite you to review the newly substituted paragraph below.

      “Prior to analytical modeling, we assessed the pairwise associations and potential collinearity among all meteorological variables using a comprehensive correlation heatmap and scatterplot matrix (Supplementary Figure S2). Notably, certain variables, such as solar radiation and the UV index, exhibited high positive correlations. To avoid coefficient instability typically caused by multicollinearity in regression models, all DLNM analyses adopted a univariate approach for meteorological factors, sequentially evaluating the nonlinear and lagged effects of individual meteorological predictors while adjusting for time trends and including other fixed covariates. Conversely, all variables were retained during the LSTM forecasting phase, as the non-linear gating architecture of recurrent neural networks inherently exhibits robust regularization against potential collinearity among input features.” (Results, page 26)

      Reviewer #2 (Public review):

      Summary:

      The study aimed to assess the associations between meteorological drivers and influenza is important although not new. The authors used only 6 years of surveillance data and deep learning models, combining distributed lag non-linear models (DLNM) with Bayesian optimized LSTM neural networks for predictive modeling. The key interest in this area is to explore the subtropical locations, where influenza is less common and circulates year round. The authors further claimed that such an association could be able to provide an early warning in the community. In this direction, the current manuscript has several scopes of improvements and clarification of the claims, as I list here.

      Strengths:

      Study design based on a prospective cohort to analyse the data for retrospective outcomes.

      We sincerely thank you for the careful and constructive evaluation of our manuscript, and in particular for recognising the value of our prospective surveillance design and the importance of investigating influenza–meteorological associations in subtropical settings where year-round circulation patterns differ substantively from those in temperate regions. We are grateful that you have identified four specific dimensions in which the manuscript can be strengthened, rationale clarity, methodological/data-integration transparency, validation reporting, and the calibration of the “early warning” claim. We address each of these four points in detail below, and we have undertaken substantive revisions to the manuscript in response.

      Weaknesses:

      (1) The rationale of the study is not clearly stated.

      We sincerely thank you for this incisive observation. We agree that the original Introduction did not adequately articulate the study’s rationale, specifically, the causal chain linking public-health need, existing methodological limitations, and the incremental contribution of our integrated DLNM-plus-LSTM framework. The Introduction has been substantively rewritten to make this rationale explicit, structured around four logical pillars:

      - Disease burden grounding. We have added quantitative evidence on the global burden of seasonal influenza, such as annual mortality estimates, drawing on solid epidemiological sources, to establish the public-health magnitude that motivates the study.

      - Subtropical-specific knowledge gap. We articulated the distinctive epidemiological challenges of subtropical influenza transmission, including year-round circulation patterns, complex non-linear meteorological associations, and lag-structured exposure-response relationships, that fundamentally differentiate subtropical contexts from temperate epidemiological settings where most existing research has been conducted. This articulation directly motivates our adoption of distributed lag non-linear models (DLNM) as the appropriate analytical framework for capturing these complex non-linear and lag-structured associations.

      - Methodological gap and incremental contribution. We now position our integrated framework against three specific gaps in the existing literature: (a) studies using DLNM alone characterize lag-distributed exposure–response relationships but lack forecasting capability; (b) studies using LSTM alone provide forecasts but typically do not incorporate Bayesian hyperparameter optimization, do not stratify by influenza subtype, and do not account for COVID-19-era non-pharmaceutical interventions; (c) no existing study, to our knowledge, integrates DLNM-based mechanistic interpretation with Bayesian-optimized, subtype-specific LSTM forecasting in a subtropical Chinese setting under pandemic-perturbed surveillance conditions. Our study is positioned to fill this specific gap.

      - Adequacy of the six-year data window. We additionally address your implicit concern regarding study duration. While six years (2018–2023) is shorter than some long-horizon influenza time-series studies, this window was deliberately selected because it brackets a uniquely informative epidemiological transition: two prepandemic baseline years (2018–2019), three Non-Pharmaceutical Interventions (NPI)-suppressed years (2020–2022), and one post-suppression rebound year (2023). This structure allows the model to learn from a structural break that a longer but earlier-only series could not provide. We have made this argument explicit in the revised Introduction.

      - The dual-model rationale. We clarify that our core rationale is to bridge this gap. By utilizing DLNM to uncover the underlying environmental biological triggers and subsequently employing the Bayesian-optimized LSTM network, uniquely upgraded to ingest epidemiological context (mask-wearing stringency index, testing volumes, day and week, viral importation), we provide a comprehensive framework that achieves both mechanistic insight and operational forecasting agility. We have substantially rewritten the Introduction to reflect this explicit storyline.

      We close the revised Introduction by explicitly framing the study’s translational endpoint: providing a methodologically integrated framework (DLNM for mechanistic interpretation; Bayesian-optimised LSTM for forecasting) to support climate-informed influenza preparedness in subtropical settings. We have, however, calibrated the language used to describe this endpoint (see our response to Weakness #4) to avoid overstating the study’s operational readiness as an early-warning tool.

      We have substantially rewritten the Introduction to reflect this sharpened rationale. We invite you to read the whole section of the Introduction in the revised manuscript.

      (2) Several issues with methodological and data integration should be clarified.

      We sincerely appreciate your identification of methodological and data integration ambiguities in the original manuscript. We have substantively addressed this concern through three categories of clarifying revisions: (i) explicit articulation of the covariate framework, (ii) clarification of the analytical relationship between DLNM and LSTM components, and (iii) detailed specification of data sources and quality control procedures:

      - Clarification 1: Comprehensive covariate framework specification. We have explicitly articulated the complete covariate framework integrated into both DLNM and LSTM analyses. Beyond the meteorological factors (mean temperature, maximum temperature, minimum temperature, diurnal temperature range, relative humidity, atmospheric pressure, precipitation, sunshine duration), our analytical framework systematically incorporates: (a) influenza positivity rates as the primary outcome variable (replacing raw case counts to mitigate surveillance intensity bias, as detailed in our response to Reviewer 1’s Public Review Comment 2); (b) weekly testing volumes as an explicit covariate to control for residual surveillance-intensity variations; (c) mask-wearing stringency indices to capture pandemic-era public health intervention effects; (d) day-of-week indicators distinguishing weekdays from weekends to control for healthcare-seeking behavioral cycles; and (e) nonlocal population proportion to account for population mobility-related transmission dynamics.

      - Clarification 2: Articulation of the DLNM-LSTM analytical relationship. We have explicitly clarified the complementary analytical roles of DLNM and LSTM within our integrated framework, addressing potential confusion regarding whether these methods serve redundant or complementary functions.

      - Clarification 3: Data source specification and quality control documentation. We have substantively expanded the data source specification and quality control documentation to ensure full methodological transparency.

      We have revised the “Study design” subsection in the Methods to transparently outline how these multifaceted data streams were temporally aligned and fed into the dual-model architecture. We invite you to review these rewritten paragraphs.

      “The study spanned January 1, 2018, to December 31, 2023, integrating four categories of data sources: (i) ILI and laboratory-confirmed cases from seven influenza sentinel hospitals across Putian’s urban and rural areas, ensuring representative coverage of diverse healthcare-seeking populations; (ii) daily meteorological data; (iii) COVID-19 associated public health intervention indicators (mask-wearing stringency indices) recorded from January 2020 onwards; and (iv) demographic mobility indicators (non-local population proportion) obtained from ILI consultation records. To ensure consistency between meteorological measurements and influenza incidence records across all data sources, we applied rigorous quality control and pre-processing procedures, including: temporal alignment of all data streams to a unified daily resolution; missing value imputation using temporally adjacent observations for sporadic gaps (<5% of records); cross-validation of laboratory-confirmed cases against ILI consultation records to identify and resolve coding inconsistencies; and standardization of meteorological measurements against the regional monitoring network’s established calibration protocols. Final datasets underwent independent verification by two co-investigators to ensure analytical reliability.” (Methods, page 11)

      “DLNM was first constructed to screen meteorological factors and other covariates with substantial influence on influenza seasonality. Subsequently, an LSTM neural network was developed within the same covariate system. The integrated DLNM–LSTM framework was employed as methodologically complementary rather than redundant components, leveraging the distinctive strengths of each approach to address different analytical objectives within a unified investigation. Specifically, DLNM models characterize the non-linear exposure–lag–response relationships between meteorological factors and influenza risk, providing biologically interpretable insights into the temporal structure of weather-influenza associations and identifying meteorological factors with statistically and clinically significant effects on influenza dynamics. Building upon this DLNM-derived foundation, the LSTM network constructs a time-series forecasting tool within an identical covariate framework, evaluating predictive capability for influenza transmission trends. Beyond meteorological factors, our DLNM-LSTM framework systematically incorporated influenza positivity rates as the primary outcome variable, weekly detection volumes, mask-wearing stringency index (indicator of NPIs during the COVID-19 pandemic), day of the week (DOW, distinguishing weekdays from weekends), and non-local population proportion (defined as the ratio of the number of individuals whose reported residential district at the time of testing lies outside Putian city to the total number of tests) as covariates within both DLNM and LSTM modeling pipelines. This comprehensive covariate framework ensures that observed meteorological associations are estimated after controlling for surveillance intensity, public health intervention status, behavioral healthcare-seeking cycles, and population mobility patterns.” (Methods, page 11-12)

      (3) Validation of the models is not presented clearly.

      We sincerely appreciate your identification of insufficient clarity in the validation framework presentation. We acknowledge that the original manuscript inadequately articulated the multi-tiered validation architecture underlying our analytical framework. We have substantively expanded the validation framework documentation through three categories of clarifications: (i) explicit articulation of the three-tier data partitioning architecture, (ii) detailed specification of validation procedures across each tier, and (iii) systematic enumeration of validation evidence supporting each analytical conclusion.

      - Clarification 1: Three-tier data partitioning architecture. Our validation framework employs a rigorously designed three-tier data partitioning architecture that addresses different validation objectives at each tier.

      Tier 1: Internal training and validation (Putian, 2018-2022). We allocated the 2018-2022 Putian surveillance data as the primary training set, within which 10% of samples were further randomly partitioned as an internal validation subset for hyperparameter tuning and overfitting monitoring during LSTM training. Early stopping mechanisms were implemented to terminate training when internal validation loss plateaued, preventing overfitting to training-specific patterns.

      Tier 2: Internal testing (Putian, 2023). We allocated the 2023 Putian surveillance data as the internal testing set, providing a temporally independent assessment of model predictive performance on data not utilized during training or hyperparameter optimization. The chronological partitioning preserves time series modeling validity by ensuring that all training data temporally precede testing data, avoiding data leakage that could artificially inflate performance estimates.

      Tier 3: External validation (Sanming, 2023). We obtained surveillance data from Sanming city for the period January 1, 2023, to December 31, 2023, matching the temporal coverage of the Putian internal testing set. This external validation set provides geographically independent assessment of model transferability across subtropical Chinese contexts, evaluating whether the Putian-derived model architecture generalizes to a different subtropical city sharing comparable climatic characteristics, influenza seasonality patterns, and public health intervention frameworks.

      - Clarification 2: Unified model framework across all validation tiers. A critical methodological feature of our validation framework is that the identical Bayesian-optimized LSTM architecture trained on Putian 2018-2022 data was applied without modification across all three validation tiers. This unified framework approach is methodologically essential because: (a) it tests genuine model transferability rather than evaluating differently-tuned models at each tier, which would conflate validation with re-optimization; (b) it enables direct performance comparison across internal testing and external validation, isolating the marginal performance degradation attributable to geographic transfer; and (c) it aligns with operational deployment scenarios where a trained model must be applied to new contexts without re-training.

      - Clarification 3: Validation evidence enumeration. The validation evidence supporting our analytical conclusions encompasses four complementary dimensions:

      Predictive performance metrics: Across both influenza A and B, the Bayesian-optimized LSTM achieved low error metrics on the internal testing set (influenza A: MAE = 0.009, RMSE = 0.035, MAPE = 0.158, SMAPE = 0.521; influenza B: MAE = 0.002, RMSE = 0.011, MAPE = 0.170, SMAPE = 0.484), substantially outperforming ARIMA benchmark models (Supplementary Figure S3).

      External validation: The Putian-derived LSTM successfully generalized to Sanming external validation data, with performance metrics maintaining comparable magnitudes to internal testing performance, substantiating model transferability across subtropical contexts.

      Sensitivity analyses: We conducted systematic sensitivity analyses across both DLNM and LSTM components. DLNM sensitivity analyses (Supplementary Figures S4-S5) demonstrate substantial concordance in cumulative risk patterns and lag-specific extreme condition responses across models with and without weekly testing volume adjustment, substantiating robustness of meteorological associations. LSTM sensitivity analyses (New Table 2) demonstrate that systematic covariate exclusion produces predictable and biologically plausible performance degradation patterns rather than artificially robust performance, confirming the absence of overfitting characteristics.

      Interpretability verification: SHAP interpretability analysis (Figure 6E-H) substantiates that the LSTM autonomously identified epidemiologically plausible feature importance hierarchies, providing independent verification that model predictions reflect genuine biological signal recognition rather than data artifacts.

      We invite you to review the corresponding manuscript clarifications.

      “The LSTM network was trained on data from Putian corresponding to the four categories described above, with the time period 2018-2022, and Putian’s 2023 data serving as the internal validation set. To rigorously evaluate model transferability beyond the training context, data from Sanming city, a mountainous subtropical city exhibiting comparable climatic characteristics, influenza seasonality, and public health intervention frameworks to Putian, were acquired for the period January 1, 2023, to December 31, 2023, temporally aligned with the Putian internal validation set. This Sanming dataset constituted our external validation set, facilitating a geographically independent assessment of model generalization within subtropical Chinese environments. The predictive performance of the LSTM algorithm for influenza A/B prevalence in 2023 Putian data was benchmarked against a parallel multivariate ARIMA model with exogenous variables. To ensure a fair methodological comparison, this baseline model was supplied with the exact same meteorological and epidemiological covariate matrix as the LSTM.” (Methods, page 12-13)

      “To respect the temporal dependence inherent in LSTM architectures and avoid data leakage, we adopted a strict chronological out-of-time (OOT) validation strategy. Time-series data from January 1, 2018, to December 31, 2022 (88.26% of the Putian dataset) were used for model training, with 10% reserved during Bayesian optimization as an internal validation subset for convergence monitoring and hyperparameter tuning only. Data from January 1, 2023, to December 31, 2023 (11.74%) were held out as a chronologically internal validation set for final performance evaluation, covering a complete annual cycle. External validation was conducted using concurrent 2023 data from Sanming city to assess spatial generalizability. To prevent distributional leakage, all normalization parameters were derived exclusively from the training set and consistently applied to the validation sets, with predictions subsequently transformed back to the original scale. First, we performed data normalization, a crucial step to ensure that training and test set data are compared on a unified scale. We normalized the training and test set data separately within the range [0, 1]. For the test set normalization, we used the maximum and minimum values from the training set as boundaries. This approach ensured consistency between the normalized test set data and the training set data. After the algorithm conducted predictions on the test set data, we performed denormalization to convert the predicted results back to the original data scale and rounded them to integers. These steps ensured that the final prediction results accurately and objectively reflected the LSTM’s performance in real-world scenarios and provided reliable data for subsequent calculation of evaluation metrics.” (Methods, page 18-19)

      (4) The claim for providing tools for 'early warning' was not validated by analysis and results.

      We are grateful for this incisive critique, which identifies a critical mismatch between our research achievements and the terminology employed in the original manuscript. Upon careful re-examination, we acknowledge unreservedly that the original manuscript’s use of “early warning” terminology overstated our actual research contribution. Our research has constructed and validated a methodologically rigorous LSTM-based influenza forecasting framework demonstrating strong predictive performance and external transferability; however, this constitutes a forecasting framework foundation rather than a fully-validated operational early warning tool ready for direct public health implementation.

      We recognize that genuine early warning tools require additional validation dimensions that our current research does not yet comprehensively address. In response to this important critique, we have implemented three categories of substantive corrections:

      - Correction 1: Comprehensive terminology revision throughout the manuscript. We have systematically revised “early warning” terminology throughout the manuscript, replacing it with more accurate descriptors that precisely characterize our actual research contribution. Specifically: “early warning system” has been revised to “forecasting framework” or “forecasting model”; “early warning tool” has been revised to “predictive modeling foundation”; and “early warning capability” has been revised to “predictive capability supporting future early warning system development”. These terminological refinements ensure that manuscript claims precisely correspond to demonstrated research achievements.

      - Correction 2: Manuscript title revision. We have correspondingly revised the manuscript title to remove “early warning” terminology and accurately reflect the study’s actual contributions: Revised title: “Meteorological Drivers of Influenza A and B Positivity in a Subtropical Chinese City: A Six-Year Surveillance Study Integrating Distributed Lag Non-Linear Models and Deep Learning”. This revised title precisely articulates the study’s actual scope: characterization of meteorological drivers (DLNM contribution), focus on positivity rates (methodological refinement addressing surveillance bias), specification of subtropical context (geographic scope), six-year temporal coverage (data scope), and integration of DLNM and deep learning (methodological framework).

      - Correction 3: Explicit articulation of forecasting framework versus operational early warning tool distinction. We have explicitly articulated the distinction between our current achievements and operational early warning tool requirements in both the Discussion and Conclusion sections, framing future research directions for operational early warning system development.

      We invite you to review the corresponding manuscript clarifications.

      “Importantly, however, this framework should be regarded as a methodological foundation for future operational developments rather than as a deployable early-warning system: routine use in public health practice would require prospective recalibration, integration with operational surveillance infrastructure, and additional validation beyond the scope of the present study.” (Discussion, page 43)

      “In conclusion, this study elucidates the distinct, non-linear meteorological drivers of influenza A and B transmission in a subtropical Chinese urban setting through an integrated dual-stage DLNM-LSTM framework. By adopting influenza positivity rates as the primary outcome and integrating socio-behavioral covariates, including mask-wearing stringency indices, weekly detection volumes, and non-local population proportion indicators, our approach mitigates surveillance-related biases and accommodates pandemic-era nonstationarity. It shows lower forecast error than a covariate-matched ARIMA baseline and provides preliminary evidence of portability within southeastern subtropical China, combining the interpretability of distributed lag modeling with the flexibility of deep learning. The framework offers an interpretable, climate-informed methodological foundation for future operational surveillance developments in subtropical settings. (Discussion, page 47)

      Reviewer #2 (Recommendations for the authors):

      (1) The title is not data-driven in different contexts, including 'early warning'; I was expecting substantial analyses in this direction to assess the 'early warning' in the manuscript. But I hardly found them in the text, merely utter as the implication of understanding the associations between meteorological drivers and influenza in advance. I suggest either revising the title or clarifying the claim by providing significant evidence and its impact on the epidemic onset and intensity. Further, revise 'Subtropical China' as 'a Subtropical Chinese city', as the former one is not accounted under this study.

      We completely agree with this constructive feedback. As detailed in our response to your Public Review weakness (4), we acknowledge that claiming an operational “early warning” system requires extensive real-world feasibility and threshold validations that exceed the scope of our current time-series analysis. We understand that achieving early warning capabilities necessitates the completion of at least three additional validation dimensions, as listed below.

      Dimension 1 - Threshold determination: Early warning systems require explicit thresholds defined by integrating predictions with established epidemiological thresholds, typically via ROC analysis to balance sensitivity and specificity for trigger activation; our framework provides predictions but does not define thresholds.

      Dimension 2 - Deployment validation: Validation of warning timeliness, false alarm control, and integration with existing CDC surveillance architectures is required; our work does not assess these operational dimensions.

      Dimension 3 - Robustness across heterogeneous scenarios: Early warning tools need systematic robustness testing across diverse social environments, public health policy contexts, extreme meteorological events, and co-circulation of emerging pathogens; our results cover 2018–2023 Putian-Sanming but not broader operational scenarios.

      We have modified the title to “a Subtropical Chinese city” and systematically revised “early warning” terminology to “forecasting” throughout the manuscript. These revisions ensure accurate representation of the study’s scope and contribution.

      (2) There are several studies establishing the potential association between influenza and the climatic drivers in several locations across the globe. I couldn't find sufficient text on establishing the rationale of this study from the perspective of existing literature. This should clearly be uttered in the introduction section itself.

      We are grateful for this critique, which echoes your Public Review weakness (1). We fully agree that the original Introduction lacked a cohesive narrative connecting the existing literature to our specific methodological innovations.

      To address this, we have comprehensively rewritten the Introduction section. The revised text now systematically establishes our rationale through a clear logical progression:

      - Acknowledging existing studies on climatic drivers but highlighting the unique challenge of non-linear, year-round influenza transmission in subtropical regions.

      - Identifying the methodological gap: existing models either use DLNM purely for retrospective explanation (lacking prediction) or employ deep learning (LSTM) purely for prediction (lacking epidemiological interpretability).

      - Highlighting the critical failure of current literature to mathematically adjust for the profound non-stationarity and surveillance intensity biases introduced by the COVID-19 pandemic (fluctuating testing volumes and NPIs).

      - Introducing our dual-stage solution: integrating DLNM and LSTM to forecast the Influenza Positivity Rate (rather than raw cases), structurally augmented with masking and testing volume covariates.

      We believe this robust literature review now unequivocally establishes the necessity and novelty of our study.

      (3) In connection with the above point, why the authors required the prospective cohort to assess a historical outcome should be highlighted clearly, which is one of the selling points of the study.

      You astutely highlight one of the core methodological strengths of our study design, and we appreciate the opportunity to emphasize this “selling point.”

      As noted in our response to Reviewer #1 regarding timeline splits, the phrase “prospective cohort to assess historical outcomes” reflects the administrative timeline of our study, but mathematically, the data stream is a continuous, longitudinal ecological surveillance.

      The critical selling point here, which we have now explicitly highlighted in the revised Methods section, is that our “historical” data (2018–2020) was not collected via traditional, unstructured retrospective chart reviews. Instead, it was derived from an already operational, highly standardized public health sentinel surveillance system. Because this system utilized identical clinical case definitions, swabbing protocols, and RT-PCR diagnostic assays continuously from 2018 through 2023, the historical data inherently possesses the high fidelity, standardized quality, and lack of recall bias typically reserved for strict prospective cohorts. We have modified the Ethics statement subsection to explicitly underscore this epidemiological advantage.

      “A major methodological strength of this study lies in its robust, uninterrupted longitudinal data collection framework spanning January 1, 2018, to December 31, 2023. While the analytical timeline encompasses a “retrospective” phase (January 1, 2018 – October 13, 2020) prior to formal ethical approval, and a “prospective” phase thereafter, we emphasize that this distinction represents a purely administrative demarcation regarding the timing of ethical approval. It does not reflect any shift in demographic cohorts, sentinel hospital locations, or data collection methodologies. Importantly, the historical data (2018–2020) were not subjected to the recall biases or misclassification risks typical of traditional retrospective chart reviews. Rather, they were systematically extracted from a continuously operating, highly standardized public health sentinel surveillance network. From the inception of data collection through the end of 2023, the local CDC maintained absolute uniformity in clinical influenza-like illness (ILI) definitions, nasopharyngeal swabbing procedures, and real-time reverse transcription polymerase chain reaction (RT-PCR) diagnostic assays. Consequently, the pre-2020 data possess the high-fidelity characteristics of a strict prospective cohort, ensuring unparalleled longitudinal consistency and mitigating temporal measurement bias across the entire pre-pandemic, pandemic, and postrestriction timeline.” (Methods, page 8-9)

      (4) The authors retrieved the daily data on the cases, which is usually small in number for most of the time, can often be driven by the importation (by population mobility with risk of infections) for particularly in a small location like a city.

      We deeply appreciate this incisive epidemiological observation. We fully agree that in a municipal-scale study, relying solely on daily absolute case counts presents significant mathematical and epidemiological vulnerabilities: absolute numbers can be small, highly stochastic, and susceptible to sudden spikes driven by imported cases rather than indigenous climate-driven transmission.

      Driven directly by your comment, we have implemented two fundamental, structural upgrades to our study design:

      - Shift to Positivity Rates: As detailed in our previous responses, we have entirely abandoned daily absolute case counts. Our DLNM and LSTM models now strictly utilize daily influenza positivity rates (positive cases ÷ total daily tested samples) for influenza A and B as the primary outcome. Positivity rates inherently smooth out the stochastic noise of small daily counts and provide a robust, normalized metric of true transmission intensity.

      - Explicit Modeling of Importation Risk: To directly address the risk of importation via population mobility, we have integrated a novel covariate into our LSTM network: the daily proportion of non-local population/residents tested (labeled as outsiders_proportion in our SHAP analysis). By explicitly feeding this mobility proxy into the deep learning algorithm, the model is now mathematically equipped to contextualize and partial out the influence of imported infections when forecasting local transmission trends.

      (5) The authors have not considered this extrinsic factor in the account and not even discussed it.

      We apologize for previously neglecting this critical extrinsic factor. As outlined in our response to Recommendation 4, we have now explicitly operationalized this extrinsic factor by incorporating daily non-local population proportion (outsiders_proportion) as a dynamic input feature in our revised LSTM architecture.

      Furthermore, to empirically assess the magnitude of this importation risk, we conducted a retrospective analysis of the demographic data spanning our six-year study period. Our descriptive statistics reveal that days where the non-local population accounted for >50% of the daily tested cohort represented less than 1% of the total study days.

      This empirical finding allows us to draw two important conclusions: First, while importation undoubtedly occurs (and is now accounted for by our LSTM covariate), indigenous transmission remains the overwhelmingly dominant driver of the observed epidemic curves in Putian. Second, massive importation shocks are rare enough that they do not systematically skew the overarching climate-disease associations identified by our models. We have thoroughly integrated both the methodological adjustment and this empirical discussion into the revised Methods and Discussion sections.

      We invite you to review the corresponding manuscript clarifications.

      “The study spanned January 1, 2018, to December 31, 2023, integrating four categories of data sources: (i) ILI and laboratory-confirmed cases from seven influenza sentinel hospitals across Putian’s urban and rural areas, ensuring representative coverage of diverse healthcare-seeking populations; (ii) daily meteorological data; (iii) COVID-19 associated public health intervention indicators (mask-wearing stringency indices) recorded from January 2020 onwards; and (iv) demographic mobility indicators (non-local population proportion) obtained from ILI consultation records.” (Methods, page 11)

      “Beyond meteorological factors, our DLNM-LSTM framework systematically incorporated influenza positivity rates as the primary outcome variable, weekly detection volumes, mask-wearing stringency indices (indicator of NPIs during the COVID-19 pandemic), day of the week (DOW, distinguishing weekdays from weekends), and non-local population proportion (defined as the ratio of the number of individuals whose reported residential district at the time of testing lies outside Putian city to the total number of tests) as covariates within both DLNM and LSTM modeling pipelines. This comprehensive covariate framework ensures that observed meteorological associations are estimated after controlling for surveillance intensity, public health intervention status, behavioral healthcare-seeking cycles, and population mobility patterns.” (Methods, page 12)

      Multivariate DLNMs efficiently expose transparent, lag-resolved main-effect surfaces for each meteorological variable, but cannot accommodate high-dimensional interactions among meteorological, autoregressive, and socio-behavioral factors without parameter inflation and severe multicollinearity. The LSTM stage was therefore not intended to replace DLNM inference, but to complement it by learning joint non-linear structure across concurrent covariates.

      Crucially, this LSTM stage explicitly incorporates weekly detection volumes, maskwearing stringency indices, non-local population proportion, and DOW effects alongside meteorological inputs, and uses influenza positivity rates rather than absolute case counts as the modeling endpoint. Together, these design choices are intended to mitigate, rather than fully eliminate, the surveillance-related biases that can distort count-based forecasting during periods of fluctuating testing intensity.” (Discussion, page 39-40)

      (6) Further, how could such a small number of cases (which can be sporadic) define the epidemic onset and its uncertainty?

      You are absolutely correct: defining an epidemic onset using a small, sporadic number of absolute daily cases introduces severe statistical uncertainty and false-positive onset triggers. This specific methodological vulnerability was a primary catalyst for our decision to fundamentally pivot our analytical framework from absolute cases to Influenza Positivity Rates.

      Unlike absolute counts, where a jump from 1 to 5 sporadic cases might artificially trigger an “onset” definition, positivity rates provide a continuous, normalized epidemiological signal. By assessing the proportion of positive tests against the total testing denominator, positivity rates mathematically stabilize the variance caused by sporadic daily testing. Consequently, an upward trajectory in positivity rates provides a highly reliable, low uncertainty signal of true epidemic onset and acceleration. The exceptional validation metrics of our revised LSTM model (e.g., MAE of 0.009 for Influenza A positivity rate) demonstrate that utilizing this normalized metric virtually eliminates the noise and uncertainty associated with sporadic small-number counts.

      (7) Figure 3 presents the time series of the cases. I wonder whether the data for these factors and outcomes are daily or aggregated by week/month? I suggest representing it in 9x1 format with a single x-axis to compare, instead of 3x3 format. Authors can refer similar plot in https://doi.org/10.1371/journal.pcbi.1012311 in Figure 1.

      We are extremely grateful for this specific and highly constructive visualization suggestion. To answer your query: the data plotted for both the meteorological factors and the influenza outcomes are indeed daily observations.

      We fully agree that the original 3x3 format severely compromised the readability of this daily data. Following your excellent advice and referencing the suggested literature, we have entirely redesigned Figure 3 into an 8x1 vertically stacked format with a single shared continuous x-axis. This structural upgrade has completely transformed the figure, eliminating the horizontal compression and elegantly exposing the fine-grained, daily temporal alignments between climatic extremes and viral surges. The newly rendered Figure 3 is now much more intuitive and analytically valuable. Due to space constraints and given that the revised Figure 3 has been included in our response to a similar comment from Reviewer #1, we will not reproduce the figure in this response. We invite you to review the revised manuscript or refer to our response to Reviewer #1’s recommendation 1 for Figure 3.

      (8) "Additionally, we plotted the loss function curves for the network on the training and validation sets to monitor LSTM convergence and the risk of overfitting." The authors validated the model with predefined training and validation sets. I suggest providing more details on the techniques and the length of the sets. How are these considerations safe for the assumptions and limitations of the models?

      We sincerely thank you for this crucial request for methodological transparency. You are absolutely correct that the techniques used for data partitioning are fundamental to the safety and validity of time-series modeling assumptions. To strictly respect the temporal dependencies of LSTM networks and avoid future-to-past data leakage, we avoided standard random train/test splitting. Instead, we implemented strict Chronological Out-of-Time (OOT) Three-Tier Validation Architecture.

      We have extensively expanded the Methods and Discussion sections to detail this partitioning scheme, our rationale, and its associated limitations:

      - The Three-Tier Validation Architecture and Length of Sets:

      Tier 1 — Training and Internal Cross-Validation (2018–2022, ~88%): Used for initial model fitting. Within this phase, 10% of the samples were held out during Bayesian optimization as an internal validation fold strictly for monitoring convergence, controlling overfitting (via early stopping), and guiding the hyperparameter search. This internal fold never contributed to the final reported performance metrics.

      Tier 2 — Internal Hold-out Test Set (2023, ~12%): A continuous 365-day block reserved as a strictly chronological hold-out, used solely for final performance evaluation. No information from this period influenced training or hyperparameter tuning.

      Tier 3 — External Independent Validation (Sanming 2023): The full surveillance time series from a geographically distinct subtropical city, providing the strongest evidence of cross-location generalizability.

      - Justification of the 88% / 12% Partition Ratio:

      This specific ratio was deliberately chosen to balance two competing epidemiological and computational considerations:

      Sufficient Training Memory: The model required a multi-year training continuum (2018–2022) to autonomously learn the structural breaks and complex non-stationarities introduced by the COVID-19 pandemic and strict NPIs.

      Epidemiological Gold Standard for Testing: Allocating exactly one year (2023) for testing is the epidemiological gold standard for seasonal infectious diseases. A full 365-day cycle ensures that model performance is evaluated across all seasonal phases (spring peaks, summer lulls, winter rebounds) rather than a biased, partial-year fragment.

      - Methodological Safeguards Against Information Leakage (Safe Assumptions):

      To ensure these partitions were safe for the model’s assumptions, we implemented strict safeguards:

      No Temporal Leakage: Tier 2 strictly follows Tier 1 in calendar time, preserving the sequential integrity assumed by LSTM architectures.

      No Distributional Leakage: All feature scaling and normalization parameters (means, standard deviations, min-max ranges) were derived exclusively from Tier 1 (Training) and applied unchanged to Tiers 2 and 3.

      - Explicit Acknowledgement of Limitations:

      We honestly acknowledge that while fixed chronological partitioning is the methodological standard for LSTM forecasting, a single fixed split point (Dec 31, 2022) fundamentally tests only one structural break. It does not exhaustively probe all potential future non-stationarities. Alternative strategies, such as expanding-window or rolling origin cross-validation, could offer additional robustness characterizations. We have integrated this crucial point into the Limitations section of the Discussion.

      We invite you to review the corresponding manuscript clarifications.

      “To respect the temporal dependence inherent in LSTM architectures and avoid data leakage, we adopted a strict chronological out-of-time (OOT) validation strategy. Timeseries data from January 1, 2018, to December 31, 2022 (88.26% of the Putian dataset) were used for model training, with 10% reserved during Bayesian optimization as an internal validation subset for convergence monitoring and hyperparameter tuning only. Data from January 1, 2023, to December 31, 2023 (11.74%) were held out as a chronologically internal validation set for final performance evaluation, covering a complete annual cycle. External validation was conducted using concurrent 2023 data from Sanming city to assess spatial generalizability. To prevent distributional leakage, all normalization parameters were derived exclusively from the training set and consistently applied to the validation sets, with predictions subsequently transformed back to the original scale.” (Methods, page 18)

      “Secondly, the interpretation of our framework’s predictive performance during the 2023 validation period requires careful epidemiological and methodological contextualization. The year of 2023 represented an anomalous, post-restriction “rebound” period characterized by rapid NPI relaxation and the release of accumulated population-level immunity debt, resulting in an atypical influenza surge that exceeded pre-pandemic peaks. The framework’s high accuracy across this period should therefore be interpreted as evidence of algorithmic agility and adaptive capacity during a highly volatile transitional phase, rather than as definitive proof of long-term predictive validity under a stabilized post-2024 epidemiological regime. Methodologically, while our strict chronological OOT data partitioning prevented temporal information leakage, a critical requirement for LSTM integrity, the reiance on a single, fixed chronological split point (December 31, 2022) intrinsically limits our evaluation to one specific structural break. This fixed-split approach may not exhaustively probe the DLNM-LSTM framework’s resilience against all forms of future epidemiological non-stationarity. Consequently, naive extrapolation of the reported 2023 performance metrics to future surveillance years should be avoided absent prospective recalibration. Future studies should consider employing expanding-window or rolling-origin cross-validation frameworks to provide a more continuous characterization of algorithmic robustness. Continuous integration of accumulating 2024 and 2025 data, combined with adaptive learning architectures capable of detecting regime shifts in real time, will be essential before any operational deployment of this, or similar forecasting frameworks, for routine public health surveillance.” (Discussion, page 44-45)

      (9) The authors considered the DLNM analysis without considering the potential interactions among meteorological factors, which can't be avoided in real-world environmental contexts. I would suggest constructing such models by incorporating more reasonable interaction terms to reflect the complex relationships between these variables. Although the impact of COVID-19 was considered on the outcome of influenza directly.

      We sincerely thank you for raising this vital conceptual point. We completely agree that meteorological factors exhibit physically real interactions under real-world environmental conditions (e.g., the synergistic effect of extreme heat and high humidity on viral viability and aerosol dynamics). We welcome the opportunity to clarify how our analytical framework structurally addresses this precise complexity without compromising mathematical stability.

      Rather than forcing interaction terms into a single statistical model, we designed our dual-stage architecture specifically to create a functional division of labor between the DLNM and LSTM frameworks:

      - Methodological Constraints on Explicit DLNM Interaction Terms: While conceptually appealing, explicitly incorporating pairwise interaction terms across eight meteorological variables within the DLNM stage would generate 28 two-way interaction cross-bases, each with its own non-linear and lag-distributed spline structure. This parameter explosion leads to the “curse of dimensionality,” causing: (a) severe multicollinearity given the strong baseline correlations among weather variables; (b) profound instability of the cross-basis estimates and inflated standard errors; (c) a massive risk of overfitting; and (d) the complete loss of visual interpretability, which is the principal value proposition of DLNM. Consequently, as is standard practice in environmental epidemiology (Gasparrini et al., 2010), we restricted our DLNM stage to isolating interpretable, lag-distributed main-effect exposure–response surfaces.

      - The LSTM Stage as the Engine for Complex Interactions: This is precisely where the deep learning architecture provides its unique methodological value. The LSTM network does not require analysts to manually pre-specify rigid interaction terms. Instead, its multi-layer, non-linear gating architecture is intrinsically capable of autonomously extracting and representing arbitrary, high-dimensional interactions among all input variables simultaneously.

      Therefore, complex meteorological interactions are absolutely not ignored in our study; rather, they are absorbed into and resolved by the LSTM stage to maximize predictive accuracy, while the DLNM stage provides the lag-resolved, interpretable backbone for individual main effects. Our multi-stage variable integration ensures that the limitations of traditional statistical models do not artificially bottleneck the deep learning framework’s capacity to synthesize real-world complexities.

      We have now added explicit paragraphs to both the Methods and Discussion sections articulating this functional division of labor, ensuring maximum methodological transparency regarding how meteorological interactions are accommodated within our framework.

      We invite you to review the corresponding manuscript clarifications.

      “In the first stage, DLNMs characterize subtype-specific, non-linear, and lag-distributed associations between meteorological variables and influenza A and B positivity. In the second stage, a Bayesian-optimized LSTM network integrating meteorological, autoregressive, and socio-behavioral covariates is used for short-horizon forecasting, benchmarked against a covariate-matched multivariate ARIMA model and evaluated in an independent subtropical city (Sanming) as a preliminary test of model portability. Influenza positivity rate is used as the primary modeling endpoint to mitigate testing-related surveillance bias. Our aim is to provide an interpretable, climate-informed forecasting approach for subtropical influenza that can serve as a methodological foundation for future operational surveillance developments.” (Introduction, page 7-8)

      “DLNM was first constructed to screen meteorological factors and other covariates with substantial influence on influenza seasonality. Subsequently, an LSTM neural network was developed within the same covariate system. The integrated DLNM–LSTM framework was employed as methodologically complementary rather than redundant components, leveraging the distinctive strengths of each approach to address different analytical objectives within a unified investigation. Specifically, DLNM models characterize the nonlinear exposure–lag–response relationships between meteorological factors and influenza risk, providing biologically interpretable insights into the temporal structure of weather influenza associations and identifying meteorological factors with statistically and clinically significant effects on influenza dynamics.” (Methods, page 11)

      “Building on the significant non-linear and lagged effects through DLNM analysis of real world environmental exposures, this study further constructed multi-factor influenza A and B prediction LSTM networks. Leveraging a recurrent architecture with non-linear gating mechanisms, these networks are able to automatically capture and represent complex, high dimensional interactions among meteorological variables without the need for manual prespecification. This functional division of labor between the DLNM and LSTM models enhances predictive performance while preserving the interpretability and inferential stability established in the DLNM stage.” (Methods, page 18)

      “A central methodological feature of our framework is the deliberate division of labor between the DLNM and LSTM components. Multivariate DLNMs efficiently expose transparent, lag-resolved main-effect surfaces for each meteorological variable, but cannot accommodate high-dimensional interactions among meteorological, autoregressive, and socio-behavioral factors without parameter inflation and severe multicollinearity. The LSTM stage was therefore not intended to replace DLNM inference, but to complement it by learning joint non-linear structure across concurrent covariates. Furthermore, extensive environmental inputs inevitably introduce severe collinearity, such as the strongly correlated solar radiation and UV index. While traditional multivariate models are highly vulnerable to such overlapping variances, the recurrent, weighted representation learned by the LSTM is comparatively tolerant of such redundancy, allowing broader covariate integration than in previous efforts (Zhu et al. 2022). Crucially, this LSTM stage explicitly incorporates weekly detection volumes, mask-wearing stringency indices, non-local population proportion, and DOW effects alongside meteorological inputs, and uses influenza positivity rates rather than absolute case counts as the modeling endpoint. Together, these design choices are intended to mitigate, rather than fully eliminate, the surveillance-related biases that can distort count-based forecasting during periods of fluctuating testing intensity.” (Discussion, page 39-40)

      (10) In context with the above points, although COVID-19-related variables are included, important confounding factors such as population mobility, vaccination coverage, and school calendar (e.g., school openings/closings) are not adequately considered. Additionally, there is a potential risk of overfitting due to an imbalanced data split-too much data is allocated to the training set, while the validation set is relatively small on the other hand.

      We sincerely appreciate your comprehensive evaluation regarding confounding control and the risk of overfitting. These are highly pertinent methodological concerns, and we have implemented multiple refinements and empirical justifications to address each of them systematically.

      - Comprehensive Control of Confounding Factors

      To address the omitted confounders you rightfully identified, we have substantially expanded our covariate framework in the revised models:

      Population Mobility: We introduced the proportion of the migrant/non-local population as a new quantitative covariate (computed as the ratio of tested individuals reporting non-Putian residential addresses). This explicitly captures the extrinsic transmission pressure and viral importation risk exerted by mobile populations.

      Social/School Routines: To capture cyclical social contact patterns and surveillance reporting dynamics, we incorporated the Day-of-the-Week (DOW) indicator (distinguishing weekdays from weekends).

      Vaccination Coverage & School Calendar (Limitations): We honestly acknowledge that highly granular, municipal-level daily vaccination registry data and official macro-school holiday timelines were unavailable for integration into our daily time-series framework. However, for essential epidemiological context, the overall influenza vaccination coverage in mainland China during the 2018–2023 study window is historically estimated at a mere 2% to 3% of the general population, substantially lower than the 40–60% coverage typically observed in high-income temperate countries. Given this exceedingly low baseline, the population-level confounding contribution of vaccination on our predictive accuracy is expected to be minimal in absolute magnitude. We have now explicitly addressed this specific regional epidemiological context in the Discussion section.

      - Justification of the Data Split Ratio (88% vs. 12%)

      Regarding the perceived imbalance in the data partition, the chronological 88% / 12% split was not arbitrary; it was designed to balance two competing epidemiological necessities:

      Preserving Training Memory: The 88% training block (2018–2022) was strictly required for the LSTM to autonomously learn the multi-year seasonal cycles, the structural breaks induced by COVID-19 NPIs, and the suppressed-regime dynamics.

      The Epidemiological Gold Standard: The 12% testing block equates exactly to the 2023 calendar year (365 days). In seasonal infectious disease forecasting, evaluating performance across a complete, unbroken annual cycle is the epidemiological gold standard, ensuring the model is tested across all phases (spring peaks, summer lulls, winter rebounds) rather than a biased, partial-year fragment.

      - Safeguards Against Overfitting & Empirical Proof (Sensitivity Analysis)

      To definitively mitigate and disprove the risk of overfitting, we implemented a Chronological Three-Tier Validation Architecture:

      During the training phase, we randomly partitioned a 10% internal validation subset exclusively for hyperparameter tuning (via Bayesian Hyperopt) and implementing early stopping to terminate training the moment validation loss plateaued.

      We utilized the full 2023 surveillance data from an entirely distinct city (Sanming) as an external independent validation set. The fact that our model generalized excellently to Sanming is the strongest empirical proof against localized overfitting.

      Finally, to further substantiate model robustness, we conducted a controlled covariateablation sensitivity analysis (new Table 2). If a deep learning model is severely overfitted (i.e., memorizing noise), removing covariates often yields chaotic or random performance changes. However, when we explicitly removed the “mask-wearing stringency indices” or the “weekly detection volumes”, our model’s performance degraded in a predictable, biologically plausible manner (e.g., MAE increased by 33.3% to 100% across subtypes).

      This structurally proves that our LSTM architecture is not achieving artificially robust performance through overfitting, but exhibits appropriate sensitivity to the exact epidemiological features driving true viral transmission.

      We have extensively documented these justifications, safeguards, and limitations in the revised Methods, Results, and Discussion sections.

      We invite you to review the newly added limitation statement paragraph.

      “Thirdly, beyond the surveillance coverage limitation noted above, the aggregate-level nature of the available data further constrained our ability to adjust for individual-level confounders, including personal vaccination status, detailed comorbidities, healthcare seeking behavior, and socioeconomic status. This limitation was particularly exacerbated by the profound epidemiological disruptions during the COVID-19 pandemic, where raw numbers of confirmed cases became heavily confounded by surveillance intensity (i.e., fluctuating testing volumes) rather than solely reflecting underlying viral transmission. Our adoption of influenza positivity rates as the primary modeling endpoint and incorporation of weekly detection volumes, face-covering stringency, and non-local population proportion as dynamic covariates, rigorously mitigated these aggregate-level biases and linked our methodological design directly to the forecasting outcomes. As unequivocally demonstrated by our covariate-ablation sensitivity analyses (Table 2), failing to account for mask mandates and testing volumes leads to severe, mathematically predictable deviations in absolute forecasting accuracy. Furthermore, while daily case counts in a single city can occasionally be small, sporadic, and driven by external importations, our incorporation of non-local population proportion effectively adjusted for these localized importation risks. Consequently, although our findings characterize population-level associations between meteorological factors and influenza activity, they should not be interpreted as evidence of micro-level causal mechanisms at the individual patient level. Ultimately, the DLNMLSTM framework’s robust performance across the non-stationary transition out of NPI policies highlights the absolute necessity of integrating behavioral and virological baseline metrics into future climate-driven predictive surveillance systems.” (Discussion, page 45-46)

      (11) The authors should justify why the baseline model selection was made by comparing the LSTM model only with ARIMA? How the outcomes could be sensitive to other commonly used machine learning methods, such as Random Forest or XGBoost, etc, as a benchmark for their performance.

      We are deeply grateful for this methodologically incisive suggestion, which prompted us to substantially strengthen the manuscript’s benchmarking framework. Recognising the methodological importance of this recommendation, our team initiated the construction of an eXtreme Gradient Boosting (XGBoost) benchmark model in parallel with the Round 1 revision submission, and has now completed its full validation and interpretive analysis. XGBoost, one of the leading non-deep-learning machine-learning frameworks for structured predictive tasks, was implemented using the identical covariate matrix employed in the LSTM. The complete methodology, results, and interpretive analysis have been incorporated as Supplementary Additional File 6, with corresponding updates to the Methods (Study Design and Model Construction), Results, and Discussion sections of the main manuscript.

      Methodological design. The XGBoost model was implemented in R (xgboost package, version 3.2.1.1) and received a covariate set strictly matched to the LSTM: eight meteorological variables, weekly testing volumes, the mask-wearing stringency index, a day-of-week indicator, and the proportion of non-local (migrant) population. Critically, no additional lag features, sliding-window statistics, or autoregressive terms were introduced, ensuring that XGBoost’s information set was identical to the contemporaneous covariate stream consumed by the LSTM at each time step. This design deliberately isolates the predictive contribution of architectural inductive bias, the LSTM’s intrinsic sequential memory versus XGBoost’s memory-free tree-partitioning geometry, from any confounding due to hand-engineered temporal features. The training (2018 – 2022) and validation (2023) partitions and the evaluation metrics (MAE, RMSE, MAPE, SMAPE) followed the LSTM and ARIMA protocols precisely.

      Empirical findings.

      A complete performance matrix including MAE, RMSE, MAPE, and SMAPE for all three models is presented in Supplementary Table S3 (Additional File 6), which confirms the same rank-ordering (LSTM ≪ ARIMA ≈ XGBoost) across all four error metrics.

      Despite receiving identical inputs, XGBoost yielded errors approximately 15-fold higher than the LSTM for Influenza A and 35-fold higher for Influenza B, with errors broadly comparable in magnitude to the ARIMA baseline. Visual inspection of the 2023 forecasting trajectories (Supplementary Figure S6) revealed that XGBoost systematically under-predicted the explosive post-NPI Influenza A rebound in late 2023 and generated largely flat forecasts for Influenza B that failed to reproduce the mid-year peak.

      Interpretation. These findings are mechanistically informative and reinforce, rather than merely confirm, the manuscript’s central methodological claim. Four architectural properties of XGBoost account for the observed performance gap:

      Insensitivity to temporal ordering. Decision-tree splits operate on feature-dimensional thresholds and cannot structurally encode “the present depends on the past” in the manner natively accommodated by recurrent architectures.

      Absence of gated recurrence. XGBoost lacks any mechanism analogous to the LSTM’s forget-input-output gating for propagating, filtering, and dynamically re-weighting historical states, and therefore cannot represent long-range non-linear temporal dependencies.

      Violated independence assumption. Tree ensembles implicitly assume independent and identically distributed (i.i.d.) observations, an assumption fundamentally at odds with the strong serial autocorrelation of epidemiological time series.

      Non-comparable modelling paradigms. Both ARIMA and LSTM are explicit sequential models within a common paradigm, whereas XGBoost belongs to a fundamentally non-sequential family. The LSTM–ARIMA contrast therefore isolates the specific contribution of deep sequential learning within a paradigmatically comparable framework, while the LSTM–XGBoost contrast demonstrates that even a state-of-the-art memory-free non-linear learner cannot substitute for a genuinely sequential architecture under epidemiologically non-stationary conditions.

      The consistent superiority of the LSTM over both a linear statistical benchmark (ARIMA) and a non-linear memory-free machine-learning benchmark (XGBoost), all receiving identical inputs and evaluated on the identical validation window, isolates the recurrent gated architecture itself as the source of the predictive gain. We are indebted to the Reviewer for prompting this three-way benchmark analysis, which has materially strengthened the methodological rigour and interpretive depth of the manuscript.

      This additional benchmark analysis, though completed subsequent to the Round 1 submission, has now been fully integrated into the Round 2 revision at the appropriate positions across the Abstract, Importance Statement, Methods, Results, Discussion, and Conclusion, together with the new Additional File 6. We are grateful to you for prompting this three-way benchmark, which we believe has materially strengthened the methodological rigour and interpretive depth of the manuscript.

      We invite you to review the newly added Discussion paragraphs.

      “The optimized LSTM algorithm demonstrated strong predictive performance, with loss curves for both subtype-specific networks exhibiting favourable convergence, consistent with prior work (Du et al. 2023). To rigorously isolate the predictive contribution attributable to the LSTM’s recurrent architecture, we benchmarked it against two covariate-matched baselines representing distinct methodological paradigms: a multivariate ARIMA model and an XGBoost gradient-boosting ensemble. This design controls simultaneously for linearity (ARIMA→LSTM contrast) and for non-linearity without recurrent memory (XGBoost→LSTM contrast), allowing us to attribute observed performance gains to specific architectural inductive biases rather than to informational asymmetry or model non-linearity in general. The LSTM substantially outperformed both baselines on the 2023 validation window. For influenza A, it achieved an MAE of 0.009, compared with 0.136 for ARIMA and 0.138 for XGBoost, approximately 15-fold reductions. For influenza B, the LSTM yielded an MAE of 0.002, versus 0.049 for ARIMA and 0.070 for XGBoost, 25- to 35-fold reductions. Critically, ARIMA and XGBoost produced errors of comparable magnitude despite their disparate assumptions regarding linearity, and both systematically under-predicted the explosive 2023 post-NPI rebound for influenza A while generating flat or spurious trajectories for influenza B (Supplementary Figures S5–S6). That parallel failure indicates the LSTM’s advantage under pandemic-era non-stationarity derives not from non-linearity per se, but from four architectural properties intrinsic to recurrent gated networks yet absent in tree ensembles: (i) threshold-based, sequence-insensitive splits fail to encode present–past dynamics; (ii) XGBoost lacks forget–input–output gates to propagate and re-weight historical states across time lags; (iii) gradient-boosted trees assume near-independence, contradicting the pronounced temporal autocorrelation in epidemiological time series; (iv) ARIMA and LSTM are sequential models with different functional forms, whereas XGBoost is non-sequential. The joint failure of ARIMA and XGBoost, despite differing non-linear treatment, isolates recurrent sequential memory as the critical architectural feature for forecasting under non-stationarity. We emphasize that this interpretation applies specifically to the present non-stationary influenza forecasting task and does not constitute a general dismissal of gradient-boosted ensembles, which retain state-of-the-art performance across many structured prediction domains. Rather, it highlights that for surveillance time series exhibiting pronounced temporal dependencies and abrupt regime shifts, such as the 2023 post-NPI rebound, explicit sequential memory becomes functionally essential. This mechanistic reading, together with the LSTM’s comparative edge over previously reported ARIMA-based (Li et al. 2024) and LSTM-based influenza prediction models (Zhu et al. 2022), positions our framework as a substantive methodological advance in predictive modeling for climate-sensitive diseases.” (Discussion, page 40-42)

      “Fourthly, our benchmarking strategy was deliberately structured as a paradigmatically layered three-way comparison (linear-sequential ARIMA, non-linear non-sequential XGBoost, non-linear sequential LSTM) rather than an exhaustive algorithmic survey. This design prioritized methodological clarity, isolating the contribution of recurrent sequential inductive bias, over horizontal coverage. Nonetheless, our evaluation does not extend to Transformer-based attention architectures, nor to hybrid ensemble strategies (e.g., LSTM–XGBoost stacking or multi-model Bayesian model averaging) that may offer complementary strengths. Systematic benchmarking against these emerging architectures, alongside prospective recalibration on additional subtropical surveillance streams and operational stress-testing under real-time data latency, will be essential next steps as the framework evolves toward operational deployment.”(Discussion, page 46)

      (12) I was expecting the statement of generalizability in terms of locations with the link to the results of the study. I suggest including such statements in the text along with limitations.

      We sincerely thank you for this highly constructive suggestion. We fully agree that an explicit, results-linked generalizability statement is crucial for appropriately interpreting and safely deploying our forecasting framework. To address this, we have substantially expanded the Limitations subsection within the Discussion to articulate a rigorous, three tier generalizability framework directly linked to our study findings:

      - Empirically Validated Generalization: The successful external validation of our framework in Sanming, a geographically distinct, inland prefecture-level city, constitutes direct empirical evidence of cross-location generalizability within the subtropical south eastern Chinese context. As detailed in our results, the framework retained high operational forecasting accuracy (e.g., Influenza A MAE of 0.015, SMAPE of 0.610) without any retraining on Sanming’s local data, proving that the meteorological exposure–response patterns and LSTM architecture are transferable across similar subtropical climates under analogous public health intervention frameworks.

      - Plausible Inferential Generalization: Beyond the directly validated Putian–Sanming pair, the framework’s architecture is plausibly generalizable to other subtropical cities across adjacent regions (e.g., Guangdong, eastern Guangxi) that share comparable monsoon climates, year-round multi-peak influenza circulation patterns, and similar sentinel surveillance infrastructures.

      - Boundaries of Transferability (Non-generalizability): We explicitly state the boundaries where our specific model parameters should not be directly extrapolated without rigorous local recalibration. These include: (a) regions located in subtropical latitudes but characterized by fundamentally distinct climate regimes (e.g., monsoonal systems with characteristics that are markedly incomparable), where the temperature–humidity coupling structures and seasonality with multiple winter peaks differ qualitatively from those in the study setting; and (b) regions with substantially different public-health intervention landscapes, such as settings with high population-level vaccination coverage or school-closure-based mitigation strategies.

      By explicitly defining these boundaries in the revised Discussion, we ensure transparent communication of the model’s appropriate application scope and methodological limitations.

      We invite you to review the newly added limitation statement paragraph.

      “Furthermore, the Bayesian-optimized LSTM architecture was directly applied, without re-training or hyperparameter adjustment, to both the Putian internal test set and the Sanming external validation set. This unified modeling framework ensures that performance differences between the two evaluation contexts primarily reflect geographic transferability rather than model re-optimization.” (Methods, page 21)

      Finally, it is imperative to explicitly delineate the generalizability of our forecasting framework within geographic and epidemiological contexts. The successful external validation in Sanming provides direct empirical evidence that our LSTM architecture and the identified meteorological thresholds are robustly transferable across the subtropical southeastern Chinese context, sharing comparable monsoon climates, year-round influenza circulation, and standardized public health intervention frameworks. However, we explicitly define the boundaries of this transferability. The specific meteorological coefficients, lag structure, and predictive parameters derived in this study should not be indiscriminately extrapolated to other regions, even within subtropical latitudes, where climatic regimes (e.g., monsoon regimes with markedly non-comparable characteristics) or public health contexts (e.g., higher baseline influenza vaccination coverage or distinct non-pharmaceutical intervention strategies) differ substantially. In such disparate settings, directly applying our pretrained model may yield substantial systematic biases. While the underlying modeling framework remains methodologically transferable, its parameterization requires rigorous local recalibration using region-specific surveillance data. Acknowledging these boundaries ensures that the framework can be deployed more safely and appropriately to realize targeted, climate-sensitive infectious disease forecasts.” (Discussion, page 44-47)”

      (13) The flow of the manuscript should be revised for general readers, for example author spent a lot of text on limitations in general without linking them to the outcomes and study design. The English language has a huge scope for improvement. I suggest paying attention to the presentation of the text for the English language use and continuity.

      We sincerely appreciate your candid feedback on the manuscript’s readability. We recognize that in our previous draft, integrating diverse critiques from multiple rounds of peer review inadvertently resulted in disjointed narrative flows, particularly in the Methods and Limitations sections.

      To address this, we have undertaken a massive structural overhaul and linguistic refinement of the entire manuscript:

      - Logical Flow: The Introduction has been sharply refocused; the Methods section has been streamlined (e.g., moving lengthy PCR protocols to the Supplement); and the Results have been reorganized with clear subheadings.

      - Contextualized Limitations: We completely rewrote the Discussion section. Rather than listing generic limitations, we have deeply anchored them to our specific study design and outcomes. For instance, we now explicitly discuss how our shift to predicting the “Positivity Rate” directly mitigates the “Surveillance Bias” limitation, and how the 2023 “rebound” limits steady-state generalizations.

      - Language Polish: The manuscript has undergone comprehensive editing by a native English-speaking academic expert to ensure grammatical precision, sophisticated vocabulary, and seamless continuity.

      We believe these revisions have dramatically elevated the clarity and scholarly tone of the text.

    1. eLife Assessment

      This study explores the phylogeny and macroevolutionary patterns of water striders, linking sexual conflict-driven morphological traits to species diversification. It provides robust and compelling evidence for ILS, introgression, and differential evolutionary rates of male and female antagonistic traits, offering valuable insights into sexual conflict and speciation.

    2. Reviewer #1 (Public review):

      Summary:

      This work explores the relationship between sexual conflict and species diversification in a clade of small water striders. They use multiple phylogenetic methods to establish a potential phylogenetic tree and explore incomplete lineage sorting and introgression. They then compare their calibrated species tree with phenotypic measures of many species to try to infer the relationship between species diversity and sexual conflict. They show evidence for both ILS and introgression within the broader clade. They also find evidence that male leg diversity is associated with speciation, potentially due to sexual conflict, and that male genital grasping structures as well as female anti-grasping traits have less evidence for an association with speciation. This paper seems like a generally rigorous and interesting addition to the literature on sexual conflict and speciation, and I believe the conclusions are well supported with only a few minor concerns with methodology.

      Strengths:

      This paper verifies results using multiple methodologies to ensure robustness to changes in software use, and presents convincing evidence for the hypotheses tested.

      Weaknesses:

      (1) Lines 299-304: For the categorization of sexual conflict traits, were categories chosen by a blinded participant or by a researcher who knew the species they were observing? If any of these traits are subtle, this could add bias to the categorisation of traits.

      (2) Lines 325-327: It is unclear to me if you control for phylogenetic relationships in this model. Diversification rate could be lineage-specific regardless of sexual conflict, so it seems like potentially including phylogenetic relationships in a model would control for that. And similarly, lines 331-333, I may be mistaken, but it sounds like you are using lm() to model a binary outcome (presence/absence), but lm() doesn't do logistic regression as far as I know, so it may be better to model this as a logistic regression (controlled for phylogeny) using glm().

      (3) There are a few areas where the reporting of inference or statistics could be improved, and throughout the manuscript there are often mentions of 'significant' without any measure of uncertainty such as confidence intervals (example on lines 449-451). In lines 385-390, because the intervals are so wide, I would suggest saying these clades diverged between X-mya and Y-mya, rather than giving an actual estimate. It seems to me like giving a specific date is a bit overconfident when the authors have such wide intervals. On line 436, I think the authors should report confidence intervals for their 5mya estimate. In lines 455-456, what is this correlation and what are the authors' uncertainties around the estimate?

      (4) Lines 549-545: When the two possible explanations for this result are reported, it seems like the authors are saying that because they can't think of how to test this possibility, the other possibility is more likely. But I don't think that is an argument against the first possibility.

      (5) Lines 591-602: This paragraph confuses me. I thought that much of the introduction and Figure 1 seemed to be putting forth that the terminal segment complexity was a conflict structure that we were interested in. However, here it is stated that it does not strongly influence mating success, so is it necessarily a conflict trait? Is it demonstrated that the leg structures influence mating success?

    3. Reviewer #2 (Public review):

      Summary:

      This study uses comparative phylogenetic methods to examine the evolution of male and female antagonistic traits in a group of small water striders. Water striders have long been a model system for studies into the sexual conflict that arises through anisogamy, the differential investment in gametes by males and females. Here, the authors aimed to reveal the evolutionary rates and trajectories of male grasping and female anti-grasping traits across species of the minute water-strider subgenus Pseudovelia. This was done by combining multiple genomic techniques to generate phylogenies to test trait evolution, quantify rates of evolution, and identify instances of incomplete lineage sorting (a result of rapid diversification) and introgression (the result of interbreeding between genetically different populations/species).

      Strengths:

      The strengths of this study lie in its comparative macroevolutionary framework, in particular the generation of multiple phylogenetic hypotheses using different methods (mitochondrial genes, USCOs, and SNPs), and contrasting these to glean insights into evolutionary patterns across species.

      Weaknesses:

      The main weakness of the study is the lack of underlying experimental evidence to explicitly show the grasping and anti-grasping functions of the various male and female traits, relying instead on studies of similar structures in more distantly related taxa. Without explicitly showing the functional mechanisms and reproductive costs of these traits, the resulting interpretations are wholly speculative. However, I would argue that such macroevolutionary studies are still very useful, and provide the groundwork for future studies untangling the relative roles of sexual conflict, cryptic female choice, sperm competition and reproductive interference in trait evolution and ultimately in speciation.

    4. Author response:

      Response to Reviewer 1:

      We thank the reviewer for their valuable and constructive suggestions.

      (1) The categorization of morphological traits was performed by researchers who are familiar with the taxonomy and morphology of this taxa. We acknowledge that this may introduce some degree of subjectivity, and we will provide photographs of other morphological traits for each species in the Supplementary data.

      (2) In our original analysis, we treated the PCAmix values derived from multiple discrete traits as continuous variables and used the lm () function to test the correlation between these values and net diversification fates (lines 352-357). We agree that a phylogenetic comparative approach would be more appropriate. We plan to re-analyze the data using either glm () or phylogenetic generalized least squares (PGLS) to properly account for phylogenetic relationships. In lines 331-333, we would clarify that this part refers to the HiSSE analysis based on discrete traits, which is independent of the linear regression analysis mentioned above. Nevertheless, we will ensure that both analyses are clearly distinguished and properly described in the revised Methods section. We will update the relevant sections accordingly.

      (3) We will carefully re-examine the entire manuscript and add appropriate measures of uncertainty.

      (4) We agree with the reviewer that two hypotheses are not mutually exclusive. In the revised manuscript, we will rephrase this paragraph to present both possibilities more neutrally.

      (5) We have observed mating behaviors in this group and found that ASE function occurs after the male has successfully grasped the female using its legs. However, we acknowledge that our study did not include direct experiments to quantitatively test the effect of ASE complexity on mating success. Therefore, our discussion in this section is indeed somewhat speculative.

      Response to Reviewer 2:

      We thank the reviewer for their encouraging and constructive comments. We agree that our study lacks direct experimental evidence to explicitly demonstrate the grasping and anti-grasping functions of the various male and female traits in Pseudovelia, and that our interpretations currently rely on comparisons with functional studies in more distantly related taxa. In the revised manuscript, we will explicitly state this limitation and refer to these traits as “putative” grasping or anti-grasping traits throughout the text where appropriate.

      We will also carefully address all minor editorial suggestions, including clarifying terminology, correcting typos, and improving figure legends.

    1. eLife Assessment

      This important study elucidates the roles of two lytic transglycosylase (LTG) enzymes in the progression of spore formation in the research model species Myxococcus xanthus. Consistent with previous work indicating loss of peptidoglycan during the laboratory-induced rapid spore formation process, the new data demonstrate a similar loss of peptidoglycan during the normal, slow sporulation process. Solid evidence is provided for the roles of the two LTGs in spore formation and for interplay between the functions of these enzymes and peptidoglycan synthetic systems. These findings may have broader implications for understanding PG metabolism across a range of bacterial species.

    2. Reviewer #2 (Public review):

      The authors initial goal was to demonstrate loss of PG during the slow sporulation process of Myxococcus xanthus, with examination of the PG degradation products in order to implicate possible enzymes involved. Upon finding a predominance of LTG products, they examined sporulation in strains lacking each of the 14 candidate LTGs encoded in the genome, leading to the identification of two sporulation-linked LTGs. An extensive characterization of the roles played by these LTGs. One LTG is responsible for the slow sporulation PG degradation, while another is required for the rapid sporulation process. Interestingly, the "slow" LTG seems to provide an important regulatory brake on the rapid enzyme. Single molecule fluorescent tracking of these enzymes was used to develop a model for their interaction with PG that mimics their observed activity. The rate of PG synthesis activity was also shown to impact the rate of PG degradation, suggesting potential interplay between the synthetic and degradative enzymes.

      Strengths:

      The genetic analysis to identify sporulation-linked LTGs and their effects on growth sporulation, and spore properties was well done and productive. The fluorescence microscopy to track LTG mobility, presumably tied to activity, produced a convincing argument about the mechanism of regulation of one LTG by another. The authors have responded well to most points of the previous review.

      Weaknesses:

      While the impact of LTGs on sporulation was clearly demonstrated, the PG analysis that resulted in the study of LTGs raised some important unanswered questions. The analyses suggest that the PG is degraded to quite small fragments, which would normally be lost during the purification of PG. The conclusions concerning the PG degradation during sporulation needs to be clarified, as described below. The authors suggest a "new mechanism of sporulation" when they have actually simply identified an important factor (PG degradation by LTGs) within a complex "process of sporulation". This needs to be reflected also in title of the paper.

    3. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      Ramirez Carbo et al. use the powerful M. xanthus spore morphogenesis model to address fundamental mechanisms in coordinated peptidoglycan remodeling and degradation. As peptidoglycan is an essential macromolecule and difficult to study in vivo, the authors use indirect but important methodology. The authors first identify two lytic transglycosylase (Ltg) enzymes necessary for spore morphogenesis using mutant phenotypic studies. They characterize these mutants for their role in coordinating spore morphogenesis induced either in fruiting bodies (starvation-dependent) or in liquid-rich media conditions (chemical-dependent). They conclude from these phenotypic and epistatic analyses that LtgA is necessary for morphogenesis during chemical-induced sporulation, and LtgB appears to be necessary to coordinate LtgA activity by interfering with LtgA function. Under starvation-induced sporulation, the absence of LtgB interferes with the building of fruiting bodies. LtgA does not appear to play a primary role in promoting aggregation into fruiting bodies, nor in degradation of peptidoglycan as assayed by loss of signal in anti-PG immunofluorescence. The authors demonstrate that the purified periplasmic domain of LtgA is highly active in degrading purified PG sacculi in vitro, while that of LtgB is highly reduced (relative to LtgA or lysozyme). The authors use photoactivated mCherry Lyt fusions and PALM to track the fusion protein mobility, which they state correlates with activity as immobilization results from PG binding. They demonstrate that in vegetative cells, a greater proportion of LtgA-PAmCh is more immobile (more active) than LtgB-PAmCh, but that directly after chemical-induction of sporulation, LtgB-PAmCh becomes more immobile (active). These analyses in the partner mutant backgrounds suggest that LtgA-PAmCh is more immobile (less active) in the absence of LtgB, but the reverse is not observed. Finally, the authors demonstrate that overexpression of LtgA in vegetative conditions leads to cell rounding, likely because of uncontrolled PG degradation, while overexpression of LtgB displays no phenotype.

      Strengths:

      This paper capitalizes on a novel spore morphogenesis mechanism to define proteins and mechanisms involved in peptidoglycan reorganization. The authors use the powerful PALM microscopy technique to assess Ltg activity in vivo by assaying for immobility as a proxy for PG binding. The authors elucidate a novel mechanism by which two Ltg's function together- with one (LtgB) seeming to regulate the activity of the other (the primary Ltg).

      Despite some weaknesses, there is no question that this study provides important insight into mechanisms of peptidoglycan remodeling- a difficult but highly impactful area of study with implications for the development of novel therapeutics and the discovery of mechanisms of fundamental bacterial physiology.

      Weaknesses:

      In many places, the authors do not adequately justify interpretations of their assays, leading to some apparently unjustified conclusions. Many of these are minor and may just require citations to demonstrate that the interpretations are justified by previous studies (detailed in recommendations below), but two bigger concerns are as follows:

      (1) It is not clear how the muropeptides listed in Figure 1 were assigned, and it is missing in the methods. In the sporulating conditions, the spectra look like combinations of multiple peaks, and the data, as stated, is not convincing to the non-specialist eye.

      We thank the reviewer for raising this point. We've expanded the Methods section to give a fuller account of how muropeptides were identified. In particular, we now describe the chromatographic separation and the assignment process, which relies on comparison with published data, fragmentation patterns, retention times, and accurate mass values. We acknowledge that the chromatograms show several peaks; nevertheless, only those muropeptides that we could clearly identify by MS/MS analysis were annotated in Figure 1. This clarification is now made explicit in the figure legend in the revised version of the manuscript.

      (2) The observation that the lytB mutant prevents appropriate aggregation into fruiting bodies does not allow the interpretation that the absence of LtgB prevents PG morphogenesis in the starvation-induced sporulation pathway, per se. It is more likely that in the LtgB mutant, the morphogenesis program is not even triggered. This is because signaling proteins and regulators (specifically, C-signal accumulation/activated FruA), which are dependent on increased cell-cell signaling in the fruiting body, do not accumulate appropriately in shallow aggregates. C-signal/FruA are necessary to trigger the sporulation program in FBs. BTW: A hypothesis to explain the indirect effect of ltgB absence on aggregation could be that UDP-precursors are not regulated appropriately (unregulated LtyA (LtgA [sic])??), so polysaccharides necessary for motility are not properly produced.

      Along these lines, fruiting body formation does not equal sporulation, and even "darkened" fruiting bodies can be misleading, as some mutants form polysacchariderich fruiting bodies (that appear dark under certain light conditions in the stereomicroscope) but do not sporulate efficiently. The wording in the text suggests that the authors assume that sporulation levels are normal because fruiting bodies are produced (see specific comments for details).

      We deeply appreciate this question. Seeking the answer, we repeated the fruiting body assay and found that both the ΔltgA and ΔltgB mutants formed dark aggregates that were comparable to wild-type fruiting bodies. However, these “fruiting body-like” aggregates did not contain sonication-resistant spores. Thus, regardless of the signals, either glycerol or starvation, sporulation requires both LtgA and LtgB. We have corrected the mistakes in the first submission.

      (3) The authors repeatedly state that production of spore coat polysaccharides likely affects the PG IP staining (see below), but this is not well justified. A citation is needed if this has already been directly shown, or the language needs to be softened.

      We agree with the reviewer. We have softened our language as “However, we cannot exclude the possibility that the polysaccharide spore coats (Voelz & Dworkin, 1962) hinder antibody access to PG.”

      (4) Better justification for the immobility of Ltg proteins in vivo as an assay for activity may be required. If this is well known in the field, it should be explicitly stated. The authors address this better in the discussion - but still state it is a correlation.

      We elaborated the justification, “Thus, when diffusive enzymes bind to PG, their mobility decreases (Lee et al., 2016; Zhang et al., 2023). For instance, DacB, another PG hydrolase, reduces its single-particle mobility in the conditions where its activity is activated (Zhang et al., 2023). By tracking single fluorescently-labeled enzyme particles, we can approximate their PG-binding in different physiological conditions and genetic backgrounds (Ramirez Carbo et al., 2024, Zhang et al., 2023, Ramírez Carbó & Nan, 2026).”

      We further discussed the correlation in discussion, “The simultaneous occurrence of reduced LtgA mobility and PG degradation during glycerol-induced sporulation indicates that the molecular dynamics of LtgA accurately mirrors its enzymatic activity. Such correlation between decreased particle mobility and increased enzymatic activity applies to many other PG-related enzymes, including multiple PG polymerases in E. coli and the endopeptidase DacB in M. xanthus (Lee et al., 2016, Zhang et al., 2023, Yang et al., 2021).”

      Reviewer #2 (Public review):

      Summary:

      The authors' initial goal was to demonstrate loss of PG during the slow sporulation process of Myxococcus xanthus, with examination of the PG degradation products in order to implicate possible enzymes involved. Upon finding a predominance of LGT products, they examined sporulation in strains lacking each of the 14 candidate LTGs encoded in the genome, leading to the identification of two sporulation-linked LTGs. An extensive characterization of the roles played by these LTGs. One LTG is responsible for the slow sporulation PG degradation, while another is required for the rapid sporulation process. Interestingly, the "slow" LTG seems to provide an important regulatory brake on the rapid enzyme. Single-molecule fluorescent tracking of these enzymes was used to develop a model for their interaction with PG that mimics their observed activity. The rate of PG synthesis activity was also shown to impact the rate of PG degradation, suggesting potential interplay between the synthetic and degradative enzymes.

      Strengths:

      The genetic analysis to identify sporulation-linked LTGs and their effects on growth, sporulation, and spore properties was well done and productive. The fluorescence microscopy to track LTG mobility, presumably tied to activity, produced a convincing argument about the mechanism of regulation of one LTG by another.

      Weaknesses:

      While the impact of LTGs on sporulation was clearly demonstrated, the PG analysis that resulted from the study of LTGs raised some important unanswered questions. The analyses suggest that the PG is degraded to quite small fragments, which would normally be lost during the purification of PG. How these small fragments were thus detected is unclear, and this suggests a more complex story concerning PG metabolism during sporulation. An anti-PG antibody is used to quantify PG in the spores, but it is not made clear what the specificity of this antibody is, and thus whether it would recognize the LTG -altered PG of the spore. The authors suggest a "new mechanism of sporulation" when they have actually simply identified an important factor (PG degradation by LTGs) within a complex "process of sporulation".

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Details on places in the text that could be improved:

      (1) Line 77: more appropriate "homologs of the sporulation genes"

      Corrected.

      (2) Line 137-139. Where are the 14 KO mutant data? Only showing the ones with phenotypes? Needs a citation if published elsewhere.

      The phenotypes of other mutants are shown in the new Fig. S1.

      (3) Line 144 and later: why "ORF"? Are they not homologous to Ltg's??. Also, the orf designation is not in the figure, so it is hard to follow the data the authors are presenting.

      Following the reviewer’s recommendation, we deleted “ORF” and presented the ORF designation in the figure legend.

      (4) Figure 2B - add the y-axis legend to the figure (also S1B).

      Added

      (5) Line 156: What does "same below" mean?

      Deleted “same below”.

      (6) Line 161: emtA is not indicated on Figure 2D - can you reword or add to the legend??

      Changed to “MltE, an LTG encoded by Escherichia coli emtA”

      (7) Line 163: "opposite roles" could be a bit better defined... do you mean ltgA induces rounding and ltgB delays rounding???

      We agree with the reviewer. “opposite” was replaced by “different”.

      (8) Line 172: But did the ltgA mutant make the same number of viable spores (not just FBs)? Can't conclude "the slow sporulation pathway only requires ltgB" if ltgA mutant was not tested.

      We agree with the reviewer. In the revised manuscript, we quantified the starvation-induced spores and found that both LtgA and LtgB are required for forming mature spores. The data are shown in Figure 3 and supplement figures.

      (9) Line 202: might be appropriate to indicate the degree of homology (% identity over protein length).

      Added.

      (10) Line 178, 181, and thereafter: DIC microscopy.

      Corrected.

      (11) Line 182/183 and 190: What is the basis for the conclusion that "polysaccharides sustained unflattened structures"??

      We changed the description to “likely due to the deposition of spore coat polysaccharides that sustained unflattened cell structures (Wartel et al., 2013; Holkenbrink et al., 2014).”

      (12) Line 185/186: What is the basis for the conclusion that ltgB only contained a small amount of PG? If it is based on reduced intensity, it is not obvious from the images presented. Quantitative analysis would better support this statement.

      We changed the description to “Sacculi from the ∆ltgB pseudospores still contained PG but showed lower fluorescence intensity”. We also updated the figure to show the typical fluorescence intensities.

      (13) Figure legend 3 line 245/6: It is not totally clear what the difference is, in that the white arrows are pointing to in ltgA vs ltgB mutant. Is it the proportion of large spherical objects in ltgB? What is the significance of the puncta in A vs. B?

      We updated the description in the text as “While the remaining PG sacculi of these pseudospores were largely spherical, they lost integrity during purification, with many sacculi displaying irregular shapes in the fluorescence channel (Figure 4A).” In the figure legend, we changed the description to White arrows point to the sacculi in irregular shapes.”

      (14) Line 190-192. Again, not really seeing what the authors are identifying to conclude "irregular shapes and ruptures" could authors pin-point more specifically and quantify such structures?

      We updated the description in the text as “While the remaining PG sacculi of these pseudospores were largely spherical, they lost integrity during purification, with many sacculi displaying irregular shapes in the fluorescence channel (Figure 4A).”

      (15) Line 194: if the authors want to make such a strong conclusion, they really should demonstrate this. Anti-spore coat antibodies are available in the field. Or take out these statements.

      We took this statement out.

      (16) Line 210: either they lack PG, OR the antibody can't gain access? How to conclude both? Might be more accurate to say "although we can't rule out that the polysaccharide spore coat prevents access to PG"

      Following the reviewer’s recommendation, we changed the description to “These fruiting body spores lacked PG-specific fluorescence (Fig. 3A), consistent with their markedly reduced PG content (Fig. 1). However, we cannot exclude the possibility that the polysaccharide spore coats (Voelz & Dworkin, 1962) hinder antibody access to PG.

      (17) Line 217-19: This is often stated in the literature, but full glycerol-induced spore maturation requires much more than 2 hours (note the authors are using O/N glycerol induction in Figure 1). And the long starvation-induced sporulation process is likely due to the differential start of sporulation, because the cells don't all enter the aggregate at the same time.... so they are not triggered to induce sporulation at the same time.

      We agree with the reviewer on the first statement and changed “two hours” to “four hours”. For the second statement, we do not completely agree. Much research showed that spore development in fruiting bodies is well synchronized, for example, in (Dworkin & Voelz, 1962), starvation-induced spores did not at 48 h.

      (18) Line 222: What is the evidence that it is really variation in production from the van promoter? Could it also just be due to differences in the cell cycle in the population? The authors may be correct, but should be less definitive about those conclusions since they haven't measured LtgA levels directly.

      We agree with the reviewer. We changed the description to, “Upon induction with 200 μM vanillate, cells exhibited heterogeneous morphology, likely resulting from variations in LtgA expression or differences in cell cycle stages within the population.” Following this suggestion, we also changed the description on murA overexpression, “Similar to the cells that overexpressed LtgA (Fig. 3B), this heterogeneity likely reflects variable murA induction or the unsynchronized growth stages within the population.”

      (19) Line 224: Where is the over-expressed LtgA data in Figure 2A? Is it that the authors are referring to over-expressed murA data, and the point is that overexpression of this kind of gene can lead to veg cell rounding? If the latter, this should be specifically stated in the text.

      We apologize for this mistake. The data were shown in Fig. 3B, rather than Fig. 2A, 3B.

      (20) Line 224: Did the authors test that LtgB is stably overproduced in veg cells? It could be that LtgB is turned over while LtgA is not. From the purified protein blot in Figure 3C, it does look like LtgB contains a degradation product (which is perhaps inhibiting the LtgB activity).

      (21) Line 229: AgmT? Do the authors mean Ltg?

      Corrected.

      (22) Line 238: Is "rate" the right word? Technically, kinetics haven't been measured. Could state "LtgB less active" or "less efficient"?

      Changed.

      (23) Line 255: "ltgB shows a slight increase in expression during slow sporulation". Do the authors mean over the entire dev time course or specifically during the sporulation phase (which will be different timing for DZ2 vs DK1622)?

      We clarified the description as “Consistent with its role in PG degradation, in a microarray-based transcriptome analysis, ltgA transcription was found to increase about twofold during rapid sporulation (4 h) but remain unchanged during slow sporulation (96 h). in the closely related DK1622 strain (Muller et al., 2010). Conversely, ltgB expression gradually rises during slow sporulation, reaching 1.8 times the vegetative level at 96 h, while remaining stable during rapid sporulation (Muller et al., 2010, Munoz-Dorado et al., 2019).”

      (24) Line 278: This needs to be corrected. The authors have shown that production of fruiting bodies is not affected by the fusions. (also in Fig. S1 legend). This is really not the same as sporulation efficiency.

      We changed the description to “the PAmCherry tags did not affect the formation of either glycerol-induced spores or starvation-induced fruiting bodies”.

      (25) Figure S1A: panel A: Is there a deg product partially cut off at the bottom of the gel? (or is this a non-specific cross-reactive band). What is the predicted molecular mass for both proteins with the PAmCherry fusion? What conditions were these lysates generated from: veg prior to glycerol induction? I understand there are probably no antibodies available to LtgA or B, but it is important to note that it is not possible to know if there is simultaneously wt LtgA or B produced (by cleavage and degradation of the mCh fusion). Panel B: Were the differences in l/w between wt and the fusion strains tested to see if there really were no significant differences? Please state in the text (it looks like the data variance is higher in the fusion strains relative to the wt at 1 hr).

      The bands at the bottom of the gel are the running front that appear in both lanes. They are not mCherry because there estimated molecular weight is much lower than that of mCherry (26.3 kDa). To avoid confusion, we cut these bands from the figure. The expression of both fusion proteins was detected from vegetative cells, which was clarified in the legend. The predicted molecular weights of them were provided in the legend too.

      (26) Figure S1C: The ability to make fb is not the same as the production of spores. The authors should test the number of viable spores produced under starvation conditions, if they want to state starvation-induced sporulation is not affected by the fusions.

      We agree with the reviewer. We quantified starvation-induced sporulation in the revised manuscript.

      (27) Line 399: Figure S2 looks at fb formation in the absence of vanillate, not sporulation.

      We changed the description to “cells grown without vanillate progressed normally through glycerol-induced sporulation and starvation-induced fruiting body formation”.

      We also quantified starvation-induced sporulation in the revised manuscript.

      (28) Line 281: How many fold is the ltgA transcript reduced compared to the ltgB? (i.e., If it is 1.2 fold reduced that may not be as worth mentioning as if it was 5-10 fold reduced).

      ltgA transcription in vegetative cells was detected in a microarray (Muller et al., 2010) but not reported in RNAseq (Munoz-Dorado et al., 2019), significantly different from that of ltgB. We pointed this out in the revised manuscript.

      (29) Figure 4B legend line 358. Define D (should it be italicized?).

      Corrected.

      (30) Line 363: Please define how significance was calculated. Is this the p-value?

      We deleted the word “significant”.

      (31) Line 298: How do the authors know immobility is from binding to PG? Provide a reference if this is well-known.

      The rationale and references have been mentioned at the beginning of this section.

      (32) Line 302 and thereafter: suggest "2.62 × 10-2 (plus minus) 2.0 × 10-3 μm2/s" is presented as "2.62 (plus minus) 0.20 × 10-2 μm2/s" for easier reading; switch to past tense (Were not are).

      Changed following the reviewer’s recommendation.

      (33) Line 310: "suggest" not "indicate", because binding of PG was directly tested.

      Corrected.

      (34) Line 341: "confirming it restricts access" seems very strong wording. Suggest: may compete with.

      Changed.

      (35) Line 392: SOME fb are larger- many are significantly smaller.

      Because fruiting body sizes do not reflect sporulation efficiency, we removed this description.

      (36) Figure S2 legend. Leaky expression; or on fruiting body formation.

      Corrected.

      (37) Line 436: stationary phase cells decrease PG synthesis- does this increase PG degradation? Perhaps it does lead to the death phase, which is striking in M. xanthus....

      How do cells die, either through death phase or under antibiotic stresses, is not well understood (Baquero & Levin, 2021) Very likely, cell death is due to the accumulation of oxidative damages (Kohanski et al., 2007) and cell lysis could be a byproduct of cell death, when cells lose control of the enzymes that break PG. While cell death is a great topic to investigate, it is beyond the scope of this study.

      (38) Line 507: washed.

      Corrected.

      (39) Line 524 (514 [sic]) and thereafter: sacculi.

      Corrected.

      (40) Line 587: reference for cell lysis procedure?

      The procedures of cell lysis, column loading and elusion were described in details “…cells were harvested by centrifugation at 6,000 × g for 20 min and lysed by sonication in buffer A (20 mM Tris-HCl pH 8.0, 200 mM NaCl), (Nan et al., 2010, Nan et al., 2006). Proteins were loaded to an NGC™ Chromatography System (BIO-RAD) and 5-ml HisTrap™ columns (Cytiva) and eluted by buffer B (20 mM Tris-HCl pH 8.0, 200 mM NaCl, 500 mM immidazole) (Pogue et al., 2018, Nan et al., 2010).”

      (41) Line 269: by microscopy (or "under THE microscope").

      Corrected.

      (42) Line 271: THE cell/PG.

      “The” added.

      (43) Line 603: OR not and.

      Corrected.

      (44) Line 497 (and elsewhere): Is it really CFU? If determined by OD, then not technically CFU because some cells will not grow into colonies.

      We agree with the reviewer. Cell concentrations were determined by OD. “CFU” was deleted.

      (45) Line 506: washed.

      Same as recommendation (38). Corrected.

      (46) Line 521: min.

      Corrected.

      (47) Line 532: How were peaks assigned?

      We described peak assignment in details in the revised manuscript, “Muropeptides were assigned based on: (i) accurate mass matching to theoretical monoisotopic masses of expected M. xanthus PG building blocks (Bui et al., 2009, White et al., 1968) and (ii) comparison of retention times with those reported in previous analysis with similar PG compositions.”

      Reviewer #2 (Recommendations for the authors):

      (1) If almost all the muropeptides detected in spores are anhydro products of LTGs, then it might be expected that these are all very small peptidoglycan fragments in the spores. If the anhydro units were at the ends of short PG chains, then muramidase digestion would release similar amounts of non-anhydro products, but none are detected. So, is muramidase digestion doing anything to the PG derived from the spores? Is muramidase digestion required to observe the spore muropeptide pattern? A control sample in which muramidase digestion is omitted would answer these questions.

      Muramidase digestion is essential for solubilizing PG into different muropeptides for UPLC analysis. In each sample, both anhydro and non-anhydro products were detected. In our original submission, we pointed out that “The two spore types showed similar profiles of a discernible presence of muropeptides that resembled those found in vegetative cells, albeit in significantly reduced quantities (Fig. 1).” We clarified the PG analysis. “The purification procedure yields only sedimentable PG, as all soluble fragments are removed during the washing steps. The resulting sacculi were then digested with muramidase, and the solubilized muropeptides were analyzed by UPLC (see Materials and Methods). The chromatograms in Fig. 1 reflect the muropeptides released specifically from the sedimented sacculus fraction.” Because muramidase release polysaccharides or disaccharides, so it’s digestion does not release nonanhydro GlcNAc species, which is why we did not see equal amounts of anhydro and non-anhydro products. In our case, over 90% of the polysaccharides and disaccharides contain Anhydro-MurNAc. The dominance of anhydro products indicates that LTGs cut very frequently on glycan chains.

      (2) This also raises the question of how these very small peptidoglycan fragments are even retained in the spores. They would be expected to be lost during spore purification or during PG purification prior to muramidase digestion. How do you even purify sacculi when the spores have no PG chains? One could theorize that the polysaccharide coats hold everything in, but then the PG would be protected from muramidase digestion.

      We thank the reviewer for this question. Our analysis suggests that M. xanthus spores do not lack PG entirely but retain a residual PG mesh that is still crosslinked. Such crosslinked material sediments during PG purification and remains accessible to muramidase, which hydrolyses internal glycosidic bonds and releases anhydro muropeptides. Non-crosslinked fragments would indeed be washed away during purification steps, so the detected muropeptides reflect the structure of this residual sacculus (Fig. 1).

      (3) What is the anti-PG antibody recognizing, the glycan backbone, the peptide side chain, or both? Can the antibody recognize the very short anhydro-containing disaccharides proposed to be predominant in spores? If not, then the PG quantification using the antibody is not accurate.

      The structures recognized by the anti-PG antibodies are unknown. Our samples do not contain small degradation products, which was clarified in the revised manuscript, “To answer this question, we purified cell sacculi and used immunofluorescence and an anti-PG serum (de Pedro et al., 1997) to visualize the remaining PG.” We used immunofluorescence to display the PG scaffolds remained in each sample and we did not perform any quantitative analysis based on the images.

      (4) Lines 325-327: This is a speculative conclusion and should be stated as such, i.e., "may control the pace.." This conclusion could be somewhat more strongly stated at the end of the next section, around lines 345-350.

      Moved following the reviewer’s recommendation.

      (5) Lines 403-404: This first sentence of the discussion seems completely dissociated from the topic of the paper; it should be deleted.

      Deleted.

      (6) Lines 404-405. I am not convinced that these findings "elucidate a new mechanism of sporulation." The study was undertaken because PG degradation was already tied to sporulation in a previous study. Furthermore, I am not sure that PG degradation is a "mechanism of sporulation." It is clearly an important step in sporulation of this species, but is it the driving "mechanism"?

      We tuned down our statement as “Our findings demonstrate that M. xanthus, a nonfirmicute bacterium, relies on PG degradation to change cell shape during sporulation”.

      (7) Lines 510-516 describe PG purification from vegetative cells. How was this process modified for spores?

      We apologize for the confusion. The process was clarified as “For PG analysis, samples were processed as previously described for Gram-negative bacteria (Alvarez et al., 2016; Desmarais et al., 2013). Vegetative cells were harvested at mid-stationary phase by centrifugation (30 min, 8,000 g). Vegetative cells and purified spores (as described in the previous section) were resuspended…”.

      Alvarez, L., Hernandez, S.B., de Pedro, M.A., and Cava, F. (2016) Ultra-Sensitive, High-Resolution Liquid Chromatography Methods for the High-Throughput Quantitative Analysis of Bacterial Cell Wall Chemistry and Structure. Methods Mol Biol 1440: 1127.

      Baquero, F., and Levin, B.R. (2021) Proximate and ultimate causes of the bactericidal action of antibiotics. Nat Rev Microbiol 19: 123-132.

      Bui, N.K., Gray, J., Schwarz, H., Schumann, P., Blanot, D., and Vollmer, W. (2009) The peptidoglycan sacculus of Myxococcus xanthus has unusual structural features and is degraded during glycerol-induced myxospore development. J Bacteriol 191: 494505.

      de Pedro, M.A., Quintela, J.C., Holtje, J.V., and Schwarz, H. (1997) Murein segregation in Escherichia coli. J Bacteriol 179: 2823-2834.

      Desmarais, S.M., De Pedro, M.A., Cava, F., and Huang, K.C. (2013) Peptidoglycan at its peaks: how chromatographic analyses can reveal bacterial cell wall structure and assembly. Mol Microbiol 89: 1-13.

      Dworkin, M., and Voelz, H. (1962) The formation and germination of microcysts in Myxococcus xanthus. J Gen Microbiol 28: 81-85.

      Holkenbrink, C., Hoiczyk, E., Kahnt, J., and Higgs, P.I. (2014) Synthesis and assembly of a novel glycan layer in Myxococcus xanthus spores. J Biol Chem 289: 32364-32378.

      Kohanski, M.A., Dwyer, D.J., Hayete, B., Lawrence, C.A., and Collins, J.J. (2007) A common mechanism of cellular death induced by bactericidal antibiotics. Cell 130: 797-810.

      Lee, T.K., Meng, K., Shi, H., and Huang, K.C. (2016) Single-molecule imaging reveals modulation of cell wall synthesis dynamics in live bacterial cells. Nature communications 7: 13170.

      Muller, F.D., Treuner-Lange, A., Heider, J., Huntley, S.M., and Higgs, P.I. (2010) Global transcriptome analysis of spore formation in Myxococcus xanthus reveals a locus necessary for cell diberentiation. BMC Genomics 11: 264.

      Munoz-Dorado, J., Moraleda-Munoz, A., Marcos-Torres, F.J., Contreras-Moreno, F.J., MartinCuadrado, A.B., Schrader, J.M., Higgs, P.I., and Perez, J. (2019) Transcriptome dynamics of the Myxococcus xanthus multicellular developmental program. Elife 8.

      Nan, B., Liu, X., Zhou, Y., Liu, J., Zhang, L., Wen, J., Zhang, X., Su, X.D., and Wang, Y.P. (2010) From signal perception to signal transduction: ligand-induced dimeric switch of DctB sensory domain in solution. Mol Microbiol 75: 1484-1494.

      Nan, B., Zhou, Y., Liang, Y.H., Wen, J., Ma, Q., Zhang, S., Wang, Y., and Su, X.D. (2006) Purification and preliminary X-ray crystallographic analysis of the ligand-binding domain of Sinorhizobium meliloti DctB. Biochim Biophys Acta 1764: 839-841.

      Pogue, C.B., Zhou, T., and Nan, B. (2018) PlpA, a PilZ-like protein, regulates directed motility of the bacterium Myxococcus xanthus. Mol Microbiol 107: 214-228.

      Ramirez Carbo, C.A., Faromiki, O.G., and Nan, B. (2024) A lytic transglycosylase connects bacterial focal adhesion complexes to the peptidoglycan cell wall. Elife 13.

      Ramírez Carbó, C.A., and Nan, B. (2026) Using Single-Particle Fluorescence Microscopy to Quantify Substrate Binding of Peptidoglycan-Modification Enzymes. Bio-protocol 16: e5696.

      Voelz, H., and Dworkin, M. (1962) Fine structure of Myxococcus xanthus during morphogenesis. J Bacteriol 84: 943-952.

      Wartel, M., Ducret, A., Thutupalli, S., Czerwinski, F., Le Gall, A.V., Mauriello, E.M., Bergam, P., Brun, Y.V., Shaevitz, J., and Mignot, T. (2013) A versatile class of cell surface directional motors gives rise to gliding motility and sporulation in Myxococcus xanthus. PLoS Biol 11: e1001728.

      White, D., Dworkin, M., and Tipper, D.J. (1968) Peptidoglycan of Myxococcus xanthus: structure and relation to morphogenesis. J Bacteriol 95: 2186-2197.

      Yang, X., McQuillen, R., Lyu, Z., Phillips-Mason, P., De La Cruz, A., McCausland, J.W., Liang, H., DeMeester, K.E., Santiago, C.C., Grimes, C.L., de Boer, P., and Xiao, J. (2021) A two-track model for the spatiotemporal coordination of bacterial septal cell wall synthesis revealed by single-molecule imaging of FtsW. Nat Microbiol 6: 584-593.

      Zhang, H., Venkatesan, S., Ng, E., and Nan, B. (2023) Coordinated peptidoglycan synthases and hydrolases stabilize the bacterial cell wall. Nature communications 14: 5357.

    1. eLife Assessment

      This valuable paper describes the regulation of the association of meiotic chromosome axis proteins on chromosome ends with sub-telomeric elements in budding yeast. The genome-wide analyses of binding of chromosome components as well as chromatin regulators, complemented with the mapping of meiotic DNA double-strand breaks on chromosome ends, provided solid evidence to support the authors' conclusion. The results in the paper are of interest to researchers studying meiotic recombination and the structure of genomes and chromosomes.

    2. Reviewer #1 (Public review):

      [Editors' note: this version has been assessed by the Reviewing Editor without further input from the original reviewers. The authors have addressed the comments raised in the previous round of review.]

      Meiotic recombination at chromosome ends can be deleterious, and its initiation-the programmed formation of DSBs-has long been known to be suppressed. However, the underlying mechanisms of this suppression remained unclear. A bottleneck has been the repetitive sequences embedded within chromosome ends, which make them challenging to analyze using genomic approaches. The authors addressed this issue by developing a new computational pipeline that reliably maps ChIP-seq reads and other genomic data, enabling exploration of previously inaccessible yet biologically important regions of the genome.

      In budding yeast, chromosome ends (~20 kb) show depletion of axis proteins (Red1 and Hop1) important for recruiting DSB-forming proteins. Using their newly developed pipeline, the authors reanalyzed previously published datasets and data generated in this study, revealing here-to-fore-unseen details at chromosome ends. While axis proteins are depleted at chromosome ends, the meiotic cohesin component Rec8 is not. Y' elements play a crucial role in this suppression. The suppression does not depend on the physical chromosome ends but on cis-acting elements. Dot1 suppresses Red1 recruitment at chromosome ends but promotes it in interior regions. Sir complex renders subtelomeric chromatin inaccessible to the DSB-forming machinery.

      The high-quality data and extensive analyses provide important insights into the mechanisms that suppress meiotic DSB formation at chromosome ends.

      Comments on latest version:

      I have checked the authors' responses and the revised analyses. I think they have adequately addressed my main concerns, particularly regarding the quantitative analyses of the chromosome fusion and SK1/S288c comparisons. I have no further comments and am content for you to proceed.

    3. Reviewer #2 (Public review):

      Summary:

      In this manuscript, Raghavan and his colleagues sought to identify cis-acting elements and/or protein factors that limit meiotic crossover at chromosome ends. This limitation is important for avoiding chromosome rearrangements and preventing chromosome mis-segregation.

      By comparing protein axis recruitment in SK1 and S288C background, which differ in their number and distribution of Y' elements, the authors show that Y' element have a limited impact on axis protein enrichment. Genetic analyses coupled with ChIP experiments revealed that the differential binding of the Red1 protein in subtelomeric regions requires the methyltransferase Dot1. Interestingly, the lack of Red1 depletion in subtelomeric regions in this mutant does not impact DSB formation. Another surprising finding is that deleting DOT1 has no effect on Red1 loading in the absence of the silencing factor Sir3. Unlike Dot1, Sir3 directly impacts DSB formation, probably by limiting promoter access to Spo11. As now clearly stated in the abstract and the discussion, this explains only a small part of the low levels of DSBs forming in subtelomeric regions and the main mechanisms suppressing crossover close to the ends of chromosomes remain to be deciphered.

      Strengths:

      This work provides intriguing observations, such as the impact of Dot1 and Sir3 on Red1 loading and the uncoupling of Red1 loading and DSB induction in subtelomeric regions.

      The separation of axis protein deposition and DSB induction observed in the absence of Dot1 is interesting because it rules out the possibility that the binding pattern of these proteins is sufficient to explain the low level of DSB in subtelomeric regions.

      The demonstration that Sir3 suppresses the induction of DSBs by limiting the openness of promoters in subtelomeric regions is convincing.

      Weaknesses:

      Sir3's impact on DSB induction is compelling, yet it only accounts for a small proportion of DSB depletion in subtelomeric regions. Thus, the main mechanisms suppressing crossover close to the ends of chromosomes remain to be deciphered. [Update: these limitations have been added to the text.]

    4. Author response:

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #1 (Public review):

      The revised manuscript includes several useful additions, and I appreciate the efforts to clarify parts of the analysis. The dataset remains valuable. However, several key issues raised previously are not yet fully resolved and continue to limit the clarity of the main conclusions.

      (1) I appreciate that the authors guide the reader to the relevant regions in the analysis of chromosome fusions (Fig. 2b). However, these subtelomeric regions are not clearly visualized, making it difficult to compare fused and unfused profiles, even though the conclusions rely largely on visual inspection of them. A more direct comparison between fused and unfused ends, together with quantitative summaries (e.g., binned Red1 enrichment and comparisons with internal regions), would make this experiment more convincing.

      Thank you for this suggestion. Figure 2 – figure supplement 1 now shows Red1 enrichment in 20-kb bins tiling in from the fusion points to more clearly show that there is no significant difference in Red1 enrichment between fused and unfused chromosomes. These data are consistent with the model that Red1 enrichment is not affected by the presence of telomeres and imply that Red1 under-enrichment near telomeres is primarily encoded in cis.

      (2) The SK1/S288c comparison (Fig. 2c) is an excellent approach, but is currently presented just as profiles, which again requires substantial effort from the reader to extract the relevant information. A systematic analysis across all informative chromosome ends-for example, comparing Red1 levels in syntenic regions using binned log2 fold-change-would more directly test the proposed in cis effect (L168) and clarify the contribution and range of Y'-associated effects. Other factors (e.g. distance from chromosome ends) could also be assessed within this framework.

      Thank you. Figure 2 – figure supplement 2 now shows the profiles placed in register using peak distribution. This analysis demonstrates that registered S288c and SK1 profiles have the same enrichment of Red1, indicating that there are no detectable long-range effects of Y’ elements or other telomere-associated sequences on the neighboring axis binding sites. Figure 2 – figure supplement 3a further quantifies the effect of registering profiles and separates the data based whether Y’ elements are present. These data are consistent with the interpretation that the presence of Y’ elements primarily affects the average axis protein enrichment profiles by displacing strong axis protein binding sites towards the chromosome interior. Our analyses also indicate that this effect is not limited to the Y’ elements as other telomere-associated sequences have a very similar effect on axis protein distribution near chromosome ends.

      Related to this, it is unclear if Y' elements themselves exhibit lower Red1 binding than the genome average. Providing the mean Red1 signal per Y' element would clarify this point and may also aid interpretation of the relationship between coding density and Red1 enrichment.

      Figure 1 – figure supplement 4a and Figure 2 – figure supplement 3b now show that the mean Red1 enrichment on Y’ elements is on average lower than in the rest of the genome. However, as shown in Figure 2 – figure supplement 3b, this effect is not unique to Y’ elements as other telomere-associated sequences show a very similar level of depletion.

      (3) The Dot1-Sir3 section is now simpler. However, I still find it difficult to follow the underlying rationale. In particular, it is unclear why a Dot1 function dependent on H3K79 methylation is introduced, given that the data in the previous section suggest H3K79 methylation is dispensable for subtelomeric Red1 depletion. A clearer statement of the authors' working model would be helpful.

      We apologize for this confusion. We restructured this section in an attempt to clarify the link between Dot1 activity and Sir3.

      Reviewer #2 (Public review):

      Summary:

      In this manuscript, Raghavan and his colleagues sought to identify cis-acting elements and/or protein factors that limit meiotic crossover at chromosome ends. This limitation is important for avoiding chromosome rearrangements and preventing chromosome mis-segregation.

      By comparing protein axis recruitment in SK1 and S288C background, which differ in their number and distribution of Y' elements, the authors show that Y' element have a limited impact on axis protein enrichment. Genetic analyses coupled with ChIP experiments revealed that the differential binding of the Red1 protein in subtelomeric regions requires the methyltransferase Dot1. Interestingly, the lack of Red1 depletion in subtelomeric regions in this mutant does not impact DSB formation. Another surprising finding is that deleting DOT1 has no effect on Red1 loading in the absence of the silencing factor Sir3. Unlike Dot1, Sir3 directly impacts DSB formation, probably by limiting promoter access to Spo11. As now clearly stated in the abstract and the discussion, this explains only a small part of the low levels of DSBs forming in subtelomeric regions and the main mechanisms suppressing crossover close to the ends of chromosomes remain to be deciphered.

      Strengths:

      This work provides intriguing observations, such as the impact of Dot1 and Sir3 on Red1 loading and the uncoupling of Red1 loading and DSB induction in subtelomeric regions.

      The separation of axis protein deposition and DSB induction observed in the absence of Dot1 is interesting because it rules out the possibility that the binding pattern of these proteins is sufficient to explain the low level of DSB in subtelomeric regions.

      The demonstration that Sir3 suppresses the induction of DSBs by limiting the openness of promoters in subtelomeric regions is convincing.

      Weaknesses:

      The section examining the impact of Dot1 and Sir3 remains complex, which is partly inherent to the intricate relationship between Dot1 and Sir3. However, the authors conclude that Dot1 acts independently of its catalytic activity based on the phenotype of the H3K79R mutant phenotype. Although this is possible it is not fully demonstrated as the H3K79R mutant may exhibit its own phenotype independently of Dot1. Unless the authors test the impact of the catalytic dead mutant Dot1-G401R on axis protein enrichment at subtelomeres they cannot claim that Dot1 act independently of its catalytic activity.

      Thank you. We softened the relevant statements and do not invoke Dot1 catalytic activity.

      Sir3's impact on DSB induction is compelling, yet it only accounts for a small proportion of DSB depletion in subtelomeric regions. Thus, the main mechanisms suppressing crossover close to the ends of chromosomes remain to be deciphered.

      We explicitly state the fact that further regulation remains to be discovered in the abstract, results, and discussion.

    1. eLife Assessment

      This work provides a reassessment of VBIT-4, a compound previously proposed to inhibit oligomerization of the crucial protein known as the mitochondrial voltage-dependent anion channel. Combining complementary experimental approaches with molecular dynamics simulations, the authors provide compelling evidence that VBIT-4 primarily disrupts lipid membranes and induces channel-independent cytotoxicity. The study has fundamental implications for interpreting previous work using VBIT-4 as a probe of channel function and highlights the need to consider membrane-disruptive effects when evaluating drug mechanisms.

    2. Reviewer #1 (Public review):

      Summary:

      The Voltage-Dependent Anion Channel 1 (VDAC1) is the most abundant β-barrel protein in the outer mitochondrial membrane and the main conduit for metabolite and ion exchange between the cytosol and mitochondria. Its oligomerization has been proposed to control mitochondrion-mediated apoptosis, making it a prime target for therapeutic intervention in diseases associated with excessive cell death, such as neurodegenerative disorders and autoimmunity. VBIT-4 is a small molecule developed to inhibit VDAC oligomerization and has shown therapeutic potential in various preclinical models. Despite its widespread use, the mechanism of action of VBIT-4 has not yet been fully elucidated. In this paper, Ravishankar et al. combine a suite of biophysical approaches with computer simulations to demonstrate that VBIT-4 forms water-permeable defects in membrane bilayers without any detectable effects on VDAC1 channel properties or oligomerization. Furthermore, cytotoxicity assays revealed identical VBIT-4 IC50 values in wild-type and VDAC1-KO cells, indicating that its activity does not depend on VDAC1. Collectively, these findings cast significant doubt on the widely held assumption that VBIT-4 is a specific inhibitor of VDAC1 oligomerization. Instead, it appears that VBIT-4 functions as a membrane-active compound.

      Strengths:

      This is a carefully conducted and well-written study that highlights potential side effects of VBIT-4, a compound that has been used to study the role of VDAC1 in a range of physiological and pathological conditions. The work is of interest to a broad readership by showcasing the importance of a systematic assessment of drug-membrane interactions to identify potential off-target membrane-driven effects of small molecules that may be mistakenly attributed to the inhibition of specific proteins. Its strength lies in the variety of complementary approaches the authors used to rigorously challenge the effect of VBIT-4 on VDAC1 organization and function. Overall, the experimental data are compelling and of high quality.

      Weaknesses:

      The authors used high-speed atomic force microscopy (HS-AFM) to study the impact of VBIT-4 on VDAC1 oligomerization in real time at nanoscale resolution. Toward this end, they adsorbed POPC:POPE:cholesterol membranes reconstituted with or without VDAC1 on mica. This revealed that addition of VBIT-4 produced small perforations in the bilayer that were independent of VDAC1. In the absence of VBIT-4, VDAC1 showed the characteristic honeycomb topography that the authors described in a previous study (Ref. 17). To quantitatively assess whether VBIT-4 affects VDAC1 organization, they analyzed protein compaction within clusters using inter-protein distance measurements. This analysis revealed no significant difference in VDAC1 organization between control conditions, 1 uM and 10 uM VBIT-4, supporting a model in which VBIT-4 primarily perturbs the lipid matrix rather than VDAC1 assemblies. This conclusion is based on the assumption that VDAC channels retain some lateral mobility in bilayers adsorbed onto mica, for which the manuscript does not provide direct evidence. In their rebuttal, the authors cite previous studies indicating that VDAC channels and beta-barrel proteins with larger extracellular domains exhibit measurable lateral diffusion in supported lipid bilayers formed on mica. Based on this, they conclude that their methodology does not constitute a limiting factor for lateral diffusion. It would be appropriate to cover this point and cite the corresponding references in the manuscript.

    3. Reviewer #2 (Public review):

      Summary

      This manuscript re-evaluates the mechanism of action of VBIT-4, a compound widely used as a putative inhibitor of VDAC1 oligomerization. The authors test whether VBIT-4 acts directly on VDAC1 assemblies or instead perturbs lipid membranes more generally. Using high-speed atomic force microscopy, electrophysiology, liposome leakage assays, Laurdan fluorescence, microscale thermophoresis, coarse-grained molecular dynamics simulations, and cell-based assays in wild-type and VDAC1-knockout HeLa cells, they show that VBIT-4 partitions into lipid bilayers, induces membrane defects and leakage, and causes VDAC1-independent cytotoxicity at concentrations commonly used in the literature to infer VDAC1-specific effects.

      Strengths

      The main strength of the study is the convergence of multiple independent approaches on the same central conclusion. Atomic force microscopy directly visualizes VBIT-4-induced defects in lipid regions while VDAC1 assemblies remain apparently intact. Electrophysiology separates VDAC1 channel behavior from background membrane conductance and shows that VBIT-4 does not measurably alter VDAC1 conductance or voltage gating, while increasing nonspecific membrane permeability. Lipid-only membranes, lipid nanodiscs lacking VDAC1, and VDAC1-knockout cells provide important controls supporting a VDAC1-independent mechanism.

      The wild-type versus VDAC1-knockout cytotoxicity comparison is a particularly strong test of VDAC1 independence at concentrations above 10 µM. The manuscript also usefully emphasizes that VBIT-4 is poorly soluble, aggregation-prone, pH-dependent, membrane-partitioning, and storage-sensitive. These properties are important for interpreting variability across previous studies using this compound.

      The manuscript is careful in defining the scope of its conclusions. It distinguishes AFM- and simulation-based measurements of VDAC1 cluster organization from cross-linking-defined proximity, which is important because these are related but non-equivalent readouts of VDAC1 organization. It also explicitly discusses how VBIT-4 solubility, aggregation, protonation, membrane partitioning, and storage sensitivity complicate comparisons based on nominal compound concentration. These points help readers interpret both the current data and the broader literature using VBIT-4.

      Limitations

      The cellular data strongly support VDAC1-independent cytotoxicity above 10 µM, but the lower-dose mitochondrial functional phenotypes, including effects on respiration, mitochondrial calcium, and mitochondrial membrane potential, were not directly compared between wild-type and VDAC1-knockout backgrounds. The manuscript appropriately avoids overinterpreting these mitochondrial effects as directly VDAC1-independent, but readers should note that VDAC1 independence is more firmly established for cytotoxicity than for the lower-dose mitochondrial phenotypes.

      The coarse-grained simulations provide useful mechanistic support for membrane partitioning, aggregation, and defect formation. However, the partitioning validation relies on the neutral VBIT-4 species and comparison with empirical partition-coefficient predictors rather than a matched all-atom octanol-water transfer calculation using the same atomistic model. This is a reasonable modeling choice, but it does not eliminate the likely importance of atomistic-level details for accurately describing pore formation. This is especially relevant for a compound with pH-dependent protonation, aggregation, and interfacial membrane localization. The simulation-derived partitioning and pore-formation results should therefore be interpreted as strong qualitative and mechanistic support rather than as a definitive quantitative description of VBIT-4 behavior across all protonation states, concentrations, and membrane environments.

      Overall assessment

      Overall, this is an important and timely study that provides a strong reassessment of VBIT-4 as a tool compound. The evidence that VBIT-4 perturbs lipid membranes independently of VDAC1 is compelling and should be useful for researchers interpreting past and future studies that use VBIT-4 as a probe of VDAC1 function.

    4. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      The authors used high-speed atomic force microscopy (HS-AFM) to study the impact of VBIT-4 on VDAC1 oligomerization in real time at nanoscale resolution. Toward this end, they adsorbed POPC:POPE:cholesterol membranes reconstituted with or without VDAC1 on mica. This revealed that the addition of VBIT-4 produced small perforations in the bilayer that were independent of VDAC1. In the absence of VBIT-4, VDAC1 showed the characteristic honeycomb topography that the authors described in a previous study (Reference 17). To quantitatively assess whether VBIT-4 affects VDAC1 organization, they analyzed protein compaction within clusters using inter-protein distance measurements. This analysis revealed no significant difference in VDAC1 organization between control conditions, 1 uM and 10 uM VBIT-4, supporting a model in which VBIT-4 primarily perturbs the lipid matrix rather than VDAC1 assemblies. This conclusion is based on the assumption that VDAC channels retain some lateral mobility in bilayers adsorbed onto mica. Do the authors have evidence that this is indeed the case? Did they also perform HS-AFM on VDAC1-containing membranes treated with VBIT-4 prior to adsorption onto mica?

      We thank the reviewer for this important question regarding the lateral mobility of VDAC1 in supported lipid bilayers (SLBs).

      It is well established that membrane proteins can retain lateral mobility in SLBs formed on mica. This is enabled by the presence of a thin interstitial water layer (typically on the order of 10–20 Å) between the substrate and the bilayer, which reduces frictional coupling and preserves membrane fluidity. This property is a key advantage of SLB systems and has been extensively described [1].

      Furthermore, HS-AFM studies provide experimental evidence supporting such mobility. For example, Casuso et al. [2] showed that OmpF trimers—whose extracellular domain is larger than that of VDAC1— exhibit measurable lateral diffusion in supported membranes. This indicates that even relatively bulky membrane proteins are not immobilized by the mica support.

      In the specific case of VDAC1, our data provide direct evidence of mobility. As shown in Figure 4J of Ref. 17, individual VDAC1 pores display lateral displacements within the membrane, despite the presence of strong protein–protein interactions. This observation indicates that VDAC1 is not rigidly immobilized upon adsorption.

      Regarding the reviewer’s question about VBIT-4 treatment prior to membrane adsorption onto mica, we also performed experiments in which VDAC1-containing proteoliposomes were pre-incubated with VBIT4 before deposition onto mica and formation of supported lipid bilayers. These samples were not imaged at sufficiently high resolution to allow the same quantitative analysis of VDAC1 cluster organization as performed for the experiments shown in the main text. However, at the resolution obtained, we did not observe obvious large-scale changes in membrane organization compared with samples in which VBIT-4 was added after supported bilayer formation.

      Finally, in light of VDAC1 mobility observed under our experimental conditions and the well-established properties of SLBs, the mica support does not constitute a major limiting factor for lateral diffusion.

      Reviewer #2 (Public review):

      (1) The main limitation is that the conclusion that VBIT-4 does not affect VDAC1 oligomerization is strongest for the specific readouts used here: atomic force microscopy measurements of cluster compaction, VDAC1 channel properties, and simulated assembly behavior. These are direct and informative measurements, but they are not identical to the chemical cross-linking readouts used in much of the prior VBIT-4 literature. Readers should therefore distinguish between VDAC1 cluster organization in membranes, as measured here, and cross-linking-defined VDAC1 proximity.

      We agree with the reviewer that AFM-defined VDAC1 cluster organization and cross-linking-defined VDAC1 proximity are related but non-equivalent readouts, and we have revised the manuscript to make this distinction explicit. However, this distinction also highlights an important limitation in interpreting changes in cross-linking efficiency as direct evidence of altered VDAC1 oligomerization. Chemical cross-linking primarily reports the proximity and accessibility of reactive residues and does not directly provide information on the number, size, stability, or supramolecular organization of VDAC1 assemblies. Moreover, VDAC1 organization is strongly influenced by the lipid environment, as shown in giant proteoliposomes with reconstituted VDAC1 by fluorescence correlation spectroscopy [3] and AFM [4], and changes in protein spacing or orientation within dynamic lipid–protein clusters could alter cross-linking efficiency without disrupting the assemblies themselves.

      Cross-linking can provide valuable information when interpreted in the context of independently defined oligomeric structures or interfaces, as recently illustrated by Takeda et al. for yeast Por1 [5]. However, in most studies reporting inhibition of VDAC1 oligomerization by VBIT-4, the evidence relies primarily on SDSPAGE analysis of chemically cross-linked species, frequently quantified as changes in VDAC1 dimers. In contrast, our HS-AFM measurements directly assess the spatial organization and compaction of VDAC1 assemblies in lipid membranes, while our simulations independently assess assembly behavior. Although these approaches do not measure cross-linking efficiency, neither reveals measurable disruption of VDAC1 assemblies by VBIT-4 under the conditions tested. Our observations are therefore difficult to reconcile with the interpretation that VBIT-4 inhibits VDAC1 assembly formation. Instead, we propose that previously reported changes in cross-linking efficiency could reflect changes in protein proximity, orientation, or residue accessibility within dynamic lipid–protein clusters, particularly given the membrane-perturbing properties of VBIT-4 demonstrated here, rather than disruption of VDAC1 assemblies themselves.

      We have therefore revised the Discussion to acknowledge that these approaches probe distinct aspects of VDAC1 organization, while also clarifying that changes in cross-linking efficiency alone cannot be interpreted as direct evidence that VBIT-4 inhibits VDAC1 assembly formation.

      “Cross-linking—the assay most commonly used to monitor VDAC1 “oligomerization”—does not report on oligomer number or stability but rather on the proximity of proteins within these adaptable clusters in MOM. In contrast, HS-AFM directly measures the spatial organization of VDAC1 assemblies in lipid membranes, while molecular dynamics simulations provide an independent description of assembly behavior. These approaches therefore probe related, but non-equivalent, aspects of VDAC1 organization. The organization of VDAC1 in the MOM is extremely sensitive to lipid composition [17]; any hydrophobic compound that perturbs membrane properties may therefore influence cross-linking efficiency through changes in protein spacing, orientation, or residue accessibility, without necessarily altering the overall organization of VDAC1 assemblies. Accordingly, although our results do not directly address cross-linking efficiency, AFM quantification (Supplementary Figure 2) and molecular dynamics simulations (Supplementary Figure 8) consistently show that VBIT-4 does not measurably alter VDAC1 cluster compaction or prevent assembly formation under the conditions examined.”

      (2) A second limitation is the uncertainty around effective VBIT-4 concentration. Because VBIT-4 is poorly soluble, aggregation-prone, pH-dependent, membrane-partitioning, and storage-sensitive, nominal added concentration may differ substantially from the concentration of active compound available in each assay. This complicates comparisons across the different in vitro, simulation, cellular, and previously published assays.

      We thank the reviewer for this important comment. We agree that the nominal concentration of VBIT-4 does not necessarily reflect the effective concentration of active compound available in solution or within lipid membranes. We have therefore expanded the Discussion to explicitly distinguish nominal from effective concentration and to emphasize that, because of VBIT-4's poor solubility, aggregation, membrane partitioning, pH-dependent protonation, and limited stability during storage, the effective concentration cannot be readily determined or compared across experimental systems. “As a consequence, the nominal concentration of VBIT-4 added to an experiment is unlikely to correspond to the effective concentration of active compound available in solution or within lipid membranes. The effective concentration is expected to vary substantially with pH, storage conditions, formulation, and membrane composition, complicating direct comparisons between different assays and across studies. “

      Importantly, to facilitate comparison with the existing literature, we deliberately used the same nominal concentration range as previous studies investigating VBIT-4. Thus, although the effective membrane concentration is inherently uncertain, this limitation applies equally to previous studies using VBIT-4 and represents an intrinsic limitation of the compound rather than of our experimental approach. Measuring the effective membrane concentration is currently not feasible and is beyond the scope of the present study. We further note that this intrinsic uncertainty likely contributes to the variability observed between published studies.

      (3) The coarse-grained simulations provide a coherent mechanistic framework for membrane partitioning, aggregation, and defect formation. However, the VBIT-4 coarse-grained model is newly parameterized and is used to support a quantitative partitioning argument. The manuscript would be easier to interpret if the coarse-grained-derived partition coefficient were reported with uncertainty, convergence information, and protonation state, and compared with a matched all-atom octanol-water partition estimate from the same atomistic model used to build the coarse-grained mapping. This matters because the partitioning argument is used quantitatively to relate micromolar aqueous VBIT-4 to millimolar concentrations in the bilayer.

      Following the Reviewer’s comments, we have expanded the Methods section to add further detail and references on the transfer free-energy calculations and include the requested details. The protonation state used throughout is neutral VBIT, as alchemical free energy calculations of charged solutes require additional corrections, and reliable reference logP values for charged species are scarce given that most empirical predictors are parameterised for neutral molecules; this is the standard Martini pathway for nonbonded term validation.

      The calculated octanol/water logP for neutral VBIT, averaged over three independent replicates, is 3.52 ± 0.01 (replicate values: 3.54, 3.53, 3.51). Convergence and overlap diagnostics confirmed well-sampled simulations across all lambda windows; the corresponding forward/backward convergence plots and MBAR overlap matrices are provided in the Supporting Information.

      Regarding the suggestion to compare against a matched atomistic free-energy calculation: we chose not to pursue this route, as atomistic MD-based logP estimates are not a more reliable reference than empirical predictors for this purpose. Benchmark studies have shown that empirical consensus methods generally outperform atomistic free-energy calculations when compared against experiment, while being substantially less computationally demanding (See [6,7]). We therefore benchmark against a consensus of five established empirical logP predictors (iLOGP, XLOGP3, WLOGP, MLOGP, SILICOS-IT via SwissADME), yielding a consensus logP of 3.42 ± 0.63 for neutral VBIT, in good agreement with our CG estimate of 3.52 ± 0.01.

      We replaced the following manuscript text in the Methods section:

      “These choices were validated by estimating CG octanol/water partitioning free energies, which were compared to predictors obtained via SwissADME[80] (iLOGP[81], XLOGP3[82], WLOGP[83], MLOGP[84], SILICOS-IT). The calculated partitioning free energies were obtained by thermodynamic integration as described elsewhere[77].”

      by the following:

      “These choices were validated by calculating CG octanol/water partitioning free energies for neutral VBIT and comparing them against a consensus of reference values obtained by theoretical predictors. Rather than comparing our Martini logP measurements against atomistic molecular dynamics calculations, we benchmark against established logP prediction methods. For equilibrium octanol/water partitioning, empirical predictors have generally demonstrated accuracy comparable to, or better than, atomistic free energy calculations when evaluated against experiment [6]. We further use a consensus of five empirical models, as consensus predictions have been shown to outperform individual predictors [7]. Reference logP values for neutral VBIT were obtained using the prediction methods available through SwissADME [8] (iLOGP [9], XLOGP3 [10], WLOGP [11], MLOGP [12], SILICOS-IT), yielding predicted logP values of 3.66, 3.85, 4.06, 2.25, and 3.28, respectively. The average of these predictions was used as a consensus estimate, giving a logP value of 3.42 ± 0.63.”

      “The calculated CG partitioning free energies were obtained for neutral VBIT by thermodynamic integration as described elsewhere [13]. In short, the solute is alchemically decoupled from each solvent environment independently across 12 lambda windows, gradually turning off all non-bonded interactions between the solute and its surroundings. The free energy change along this path corresponds to the solvation free energy in that solvent, and taking the difference between octanol and water (ΔG_octanol − ΔG_water) yields the transfer free energy, which is converted to a partition coefficient. Three independent replicates were run, yielding octanol-water logP values of 3.54, 3.53, and 3.51, with an average of 3.52 ± 0.01. Convergence and overlap analysis confirmed that all lambda windows were well-sampled and that the free energy estimates were statistically reliable as required per the guidelines for the analysis of free energy calculations [14]; the corresponding forward/backward convergence plots and MBAR overlap matrices are provided in the Supplementary Figures 11 and 12.”

      (4) Finally, the cellular data strongly support VDAC1-independent cytotoxicity, but the lower-dose mitochondrial functional phenotypes were not directly compared between wild-type and VDAC1-knockout backgrounds. VDAC1 independence is therefore more directly established for cytotoxicity than for the lower-dose mitochondrial phenotypes.

      We agree that our cellular data cannot confirm that the mitochondrial effect of VBIT-4 is independent of VDAC1. However, a previous study by Belosludtsev et al. shows a similar decrease in membrane potential upon VBIT-4 treatment due to inhibition of electron transport chain complexes [15]. Following the Reviewer’s advice, we narrowed the wording accordingly in the Results and Discussion sections.

      Overall, this work provides a valuable and timely reassessment of VBIT-4, and its central conclusion will be useful for researchers interpreting studies that use this compound as a probe of VDAC1 function.

      Suggestions for authors:

      (1) Soften categorical statements such as "VBIT-4 does not alter VDAC1 oligomerization" by specifying the tested readouts: VDAC1 cluster compaction, channel properties, and simulated assembly behavior under the conditions used here. A matched cross-linking experiment under the authors' own VBIT-4 handling and concentration conditions could be useful, but is not essential; the essential point is to make clear that cross-linking-defined VDAC1 proximity and AFM/simulation-defined membrane cluster organization are related but non-equivalent readouts.

      We modified the manuscript to soften the tone. In the discussion, we now emphasize that cross-linking and HS-AFM/simulation probe related but non-equivalent aspects of VDAC1 organization, and that differences in cross-linking efficiency previously reported could be due to alteration of protein spacing or orientation.

      These findings demonstrate that VBIT-4 acts by perturbing lipid bilayers rather than through detectable direct modulation of VDAC1,

      Change title: VBIT-4 Does Not Alter VDAC1 Oligomerization

      To “VBIT-4 Does Not Measurably Alter VDAC1 Cluster Organization or Assembly Behaviour”

      This indicates that VBIT-4 neither prevents nor disrupts VDAC oligomerization.

      To “This indicates that VBIT-4 does not measurably prevent or disrupt VDAC1 assembly under the simulated conditions.”

      “Cross-linking—the assay most commonly used to monitor VDAC1 “oligomerization”—does not report on oligomer number or stability but rather on the proximity of proteins within these adaptable clusters in MOM. In contrast, HS-AFM directly measures the spatial organization of VDAC1 assemblies in lipid membranes, while molecular dynamics simulations provide an independent description of assembly behavior. These approaches therefore probe related, but non-equivalent, aspects of VDAC1 organization. The organization of VDAC1 in the MOM is extremely sensitive to lipid composition [17]; any hydrophobic compound that perturbs membrane properties may therefore influence cross-linking efficiency through changes in protein spacing, orientation, or residue accessibility, without necessarily altering the overall organization of VDAC1 assemblies. Accordingly, although our results do not directly address cross-linking efficiency, AFM quantification (Supplementary Figure 2) and molecular dynamics simulations (Supplementary Figure 8) consistently show that VBIT-4 does not measurably alter VDAC1 cluster compaction or prevent assembly formation under the conditions examined.”

      (2) More explicitly distinguish nominal added VBIT-4 concentration from effective available concentration, given the solubility, aggregation, pH-dependence, membrane partitioning, and storagesensitivity observations.

      We modified the Discussion (see public review)

      (3) Report the coarse-grained-derived octanol/water partition coefficient or transfer free energy numerically, with uncertainty, convergence information, and protonation state. Consider providing the corresponding all-atom octanol/water transfer free energy or partition coefficient for the same protonation state(s).

      We answered this comment and added two Supplemental figures 11 and 12. (see public review)

      (4) Either repeat the oxygen consumption rate, TMRM, and Rhod-2 assays in VDAC1-knockout cells, or narrow the wording so that only cytotoxicity is described as directly shown to be VDAC1-independent.

      Thank you for pointing this out. Following the Reviewer’s advice, we narrowed the wording accordingly in the Results (suppression of “This demonstrates that the cytotoxicity is due to a loss of membrane integrity.”) and Discussion sections.

      We removed the direct link to VDAC1: At concentrations below 10 μM, VBIT-4 decreased mitochondrial calcium, respiration, and membrane potential in HeLa cells without affecting mitochondrial mass. They align with reports that VBIT-4 also accumulates in the mitochondrial inner membrane, where it inhibits respiratory complexes I, III, and IV and decreases mitochondrial membrane potential [15,16].

      (5) Consider moving the storage-stability observation into the main text, given its likely importance for interpreting variability in the broader VBIT-4 literature. It would also be useful to include clearer information on stock age, storage temperature, freeze-thaw history, solvent conditions, and whether precipitation or turbidity was observed.

      The Supplemental Figure 9C was moved the main text as new Figure 6, and additional information about storage conditions is added to the Methods section.

      (6) Minor correction: the parenthetical "10^3.5 = 3.2" should be corrected. Since 10^3.5 is approximately 3,162, the intended statement appears to be that a 1 µM aqueous concentration corresponds to approximately 3.2 mM in the bilayer.

      Thank you for pointing it out, it is corrected.

      References:

      (1) Castellana, E. T. & Cremer, P. S. Solid supported lipid bilayers: From biophysical studies to sensor design. Surf. Sci. Rep. 61, 429–444 (2006).

      (2) Casuso, I. et al. Characterization of the motion of membrane proteins using high-speed atomic force microscopy. Nat. Nanotechnol. 7, 525–529 (2012).

      (3) Betaneli, V., Petrov, E. P. & Schwille, P. The role of lipids in VDAC oligomerization. Biophys. J. 102, 523–531 (2012).

      (4) Lafargue, E. et al. Lipid composition of the membrane governs the oligomeric organization of VDAC1. 2024.06.26.597124 Preprint at https://doi.org/10.1101/2024.06.26.597124 (2024).

      (5) Takeda, H. et al. Oligomer-based functions of mitochondrial porin. Nat. Commun. 16, (2025).

      (6) Işık, M. et al. Assessing the accuracy of octanol–water partition coefficient predictions in the SAMPL6 Part II log P Challenge. J. Comput. Aided Mol. Des. 34, 335–370 (2020).

      (7) Calculation of molecular lipophilicity: State‐of‐the‐art and comparison of log P methods on more than 96,000 compounds - Mannhold - 2009 - Journal of Pharmaceutical Sciences - Wiley Online Library. https://onlinelibrary.wiley.com/doi/10.1002/jps.21494.

      (8) Daina, A., Michielin, O. & Zoete, V. SwissADME: a free web tool to evaluate pharmacokinetics, druglikeness and medicinal chemistry friendliness of small molecules. Sci. Rep. 7, 42717 (2017).

      (9) Daina, A., Michielin, O. & Zoete, V. iLOGP: a simple, robust, and efficient description of noctanol/water partition coefficient for drug design using the GB/SA approach. J. Chem. Inf. Model. 54, 3284–3301 (2014).

      (10) Cheng, T. et al. Computation of octanol-water partition coefficients by guiding an additive model with knowledge. J. Chem. Inf. Model. 47, 2140–2148 (2007).

      (11) Wildman, S. A. & Crippen, G. M. Prediction of Physicochemical Parameters by Atomic Contributions. J. Chem. Inf. Comput. Sci. 39, 868–873 (1999).

      (12) Lipinski, C. A., Lombardo, F., Dominy, B. W. & Feeney, P. J. Experimental and computational approaches to estimate solubility and permeability in drug discovery and development settings. Adv. Drug Deliv. Rev. 46, 3–26 (2001).

      (13) Souza, P. C. T. et al. Protein-ligand binding with the coarse-grained Martini model. Nat. Commun. 11, 3714 (2020).

      (14) Klimovich, P. V., Shirts, M. R. & Mobley, D. L. Guidelines for the analysis of free energy calculations. J. Comput. Aided Mol. Des. 29, 397–411 (2015).

      (15) Belosludtsev, K. N. et al. Effect of VBIT-4 on the functional activity of isolated mitochondria and cell viability. Biochim. Biophys. Acta Biomembr. 1866, 184329 (2024).

      (16) Belosludtsev, K. N. et al. Pharmacological and Genetic Suppression of VDAC1 Alleviates the Development of Mitochondrial Dysfunction in Endothelial and Fibroblast Cell Cultures upon Hyperglycemic Conditions. Antioxid. Basel Switz. 12, 1459 (2023).

    1. eLife Assessment

      The study provides important insight into the development of CD8 T cell virtual memory cells by identifying a dose-sensitive role for Bcl11b. The authors define the developmental stages affected by reduced Bcl11b expression and establish its role in promoting the virtual memory fate. The evidence is compelling, supported by multiple genetic models and complementary approaches.

    2. Reviewer #1 (Public review):

      Summary:

      In this study, Sidwell and Rothenberg report a genetically rigorous study whose central finding - reduction of Bcl11b at the positive selection stage reroutes CD8 T cells to the T cell virtual memory (TVM) fate - is well supported by the convergence of elegant mouse models (WT, Bcl11bΔEnh, Bcl11b+/-, Bcl11bR3S). In wildtype mice, Bcl11b acts as a transcriptional repressor that limits the reprogramming of CD8 T cells to the TVM cell type. Reducing Bcl11b dosage partially relieves this repression, resulting in increased development toward the TVM fate. Their analysis reveals several interesting aspects: 1)<br /> Strengths:

      Factors that drive CD8 TVM fate are important and relatively poorly understood. This study makes the unexpected and interesting observation that Bcl11b gene dosage impacts TVM during T cell selection in the thymus. The conclusions are strongly supported by multiple independent lines of evidence.

      Weaknesses:

      The direct genomic targets impacted by Bcl11b heterozygosity were not identified.

    3. Reviewer #2 (Public review):

      This manuscript by Sidwell and Rothenberg demonstrates that commitment of CD8 T cells to the virtual memory TVM cell lineage is fine-tuned in a dose-dependent manner by the transcription factor Bcl11b during intrathymic positive selection. Using multiple mouse models, the authors show that a subtle, less than two-fold reduction in Bcl11b expression or disruption of its corepressor-recruitment domain biases developing CD8 single-positive thymocytes toward a TVM cell fate without requiring peripheral activation, lymphopenia, or external cytokine signaling. Mechanistically, this modest decrease in Bcl11b does not alter global chromatin accessibility but instead enhances downstream T-cell receptor (TCR) signal responsiveness, effectively mimicking a high-affinity selection response to divert late-cycling CD8SP thymocytes into the TVM pathway. These data suggest that Bcl11b essentially serves to attenuate the interpretation of TCR (and cytokine) mediated signals to prevent the excessive differentiation characterised by virtual memory T cells and the CD44int naïve T cells. This is distinct from alternative pathways of Tvm development that are driven predominantly by exposure to cytokines, namely IL-4, in the thymus, and serves to reinforce our understanding that Tvm cells are an alternate lineage of T cells that arise during development, in part as a consequence of strong TCR signalling. There are some issues arising, not least of which is why the attenuated Bcl11b expression is insufficient to drive negative selection rather than Tvm formation.

      This paper was an absolute pleasure to read given its engaging narrative style. However, in some parts it was a bit long-winded and took a while to get to the destination. Some effort should go into making the narrative more concise, while retaining the thoroughly clear explanation and interpretation of the data.

    4. Reviewer #3 (Public review):

      Summary:

      The authors explore the impact of a modest (<2-fold) reduction in the expression of Bcl11b on the differentiation of CD8+ T cells, with a special focus on the generation of memory-like cells (sometimes called "virtual" memory cells - or TVM). The manuscript covers a lot of ground, but highlights are using diverse models to show that reduced Bcl11b expression during thymic development (but not in naïve CD8+ T cells that have accessed the periphery) leads to enhanced generation of cells with phenotypic, transcriptional, epigenetic and functional characteristics of TVM; that this is a cell-intrinsic effect, but not apparently driven by enhanced responsiveness to cytokines (which promote TVM in some models); decreased Bcl11b improves T cell sensitivity at the mature (and likely in immature thymocytes). Myriad approaches and controls are used, providing a very thoroughly explored model.

      This is a tour-de-force in applying the geneticists tool-box for investigating how a tantalizingly modest decrease in Bcl11b expression impacts the generation of TVM-like cells during thymic development. It is unreasonable to request additional data, but clarification of some key conclusions is needed.

    5. Author response:

      We sincerely thank the editors and the three reviewers for their thorough, highly constructive, and positive evaluation of our manuscript. We are gratified by the reviewers’ recognition of the genetic rigor of our study and the compelling nature of our findings regarding the dose-dependent role of Bcl11b in virtual memory CD8 T cell differentiation.

      We agree that the reviewers have raised fair and addressable points that will undoubtedly strengthen the final manuscript. Below, we outline our planned revisions to address the primary themes raised in the public reviews:

      (1) Genomic Targets and Bcl11b Occupancy (Reviewers #1 & #3)

      To address the request for direct genomic targets of Bcl11b, we will incorporate our existing Bcl11b ChIP-seq data from Bcl11b haploinsufficient and control peripheral naïve CD8+ T cells. We will provide comparative analyses demonstrating that Bcl11b ChIP-seq read density per region is highly concordant with population-matched ATAC-seq accessibility profiles, highlighting that direct Bcl11b occupancy closely mirrors the accessible chromatin landscape, which itself we have assayed thoroughly in thymic precursors. Furthermore, we will integrate this ChIP-seq analysis to cross-reference our bulk and pseudobulk differentially expressed gene lists to better define target overlaps.

      (2) Phenotypic Definitions, Cytokine Independence, and Quantification (Reviewers #2 & #3)

      We appreciate the reviewers’ suggestions to further solidify the phenotypic definitions of our populations. In our revision, we will:

      Perform targeted flow cytometry utilizing our Bcl11b haploinsufficient models to provide explicit quantification of CD5 expression across mature thymic (DP to CD8SP) and peripheral CD8 T cell populations.

      Include supplementary flow cytometry panels for CD122 alongside our standard CD44, CD62L, and CD49d gating strategies used throughout the manuscript in order to comprehensively lock down the T<sub>VM</sub> vs. T<sub>CM</sub> phenotypic definitions.

      Provide absolute cell counts (rather than just relative frequencies) for neonatal CD8 populations to explicitly confirm absolute expansion.

      Expand our evaluation of cytokine-independence by analyzing ImmGen-derived cytokine-response signatures against our scRNA-seq datasets.

      (3) Single-Cell Transcriptomic Alignments (Reviewer #3)

      To contextualize our findings within the broader literature, we will score our scRNA-seq datasets against the derived T<sub>VM</sub> transcriptional signatures recently published by Zhang et al. (2024). While exact cluster-to-cluster matching may be limited by differences in experimental models (e.g., steady-state ontogeny versus influenza infection), this alignment will allow us to demonstrate where our newly minted thymic T<sub>VM</sub>(/precursor) cells map along the established peripheral T<sub>VM</sub>-state continuum.

      Additionally, we will generate targeted split-violin visualizations of T<sub>VM</sub> module scores specifically within scRNA-seq Cluster 8. This will visually clarify the transcriptomic shifts driven by Bcl11b haploinsufficiency within this specific cluster, supplementing the DEG/GSEA tables currently provided.

      (4) Conceptual Clarifications and Discussion Expansions (Reviewers #2 & #3)

      We will expand our discussion to address several excellent conceptual points raised by the reviewers:

      Negative Selection vs. Fate Diversion: We will clarify why attenuated Bcl11b alters TCR-induced gene programs without triggering negative selection, emphasizing that Bcl11b dose reduction drives a portional failure of specific transcriptional repression rather than a general deregulation of global TCR-dependent signaling.

      Haploinsufficiency vs. Knockout: We will explicitly contrast our haploinsufficient (<2 fold reduction) virtual memory phenotype with the innate-like T (and ex-T) cell phenotypes previously reported in complete Bcl11b loss-of-function models.

      Mitochondrial Dynamics: We will refine our text regarding mitochondrial biology to more clearly distinguish between compensatory nuclear transcription (the mitochondrial gene module) and physical organelle performance (mitochondrial membrane potential).

      We look forward to submitting the fully revised manuscript and a detailed point-by-point response in the near future.

    1. eLife Assessment

      This fundamental study advances our understanding of the neural substrate of planning trajectories towards a goal by using recurrent neural networks. The manuscript provides compelling evidence for the proposed mechanism by linking a computational model to previous findings in neural data, designing a hand-crafted implementation of the mechanism, and showing that trained recurrent networks converge to computational principles consistent with the proposed mechanism. The work will be of broad interest to theoretical and systems neuroscientists and to cognitive scientists.

    2. Reviewer #1 (Public review):

      Summary:

      This work builds a theory to implement planning trajectories towards a goal in a known environment, inspired by analyses of prefrontal neural recordings. Unlike standard neural architectures for this task, such as value-based learning and successor representations, their proposed theory is able to adapt to novel goal locations within-trial. The key to the theory is that future times and locations are represented by disjoint groups of neurons. The recurrent connectivity between groups of neurons selective to specific future time and locations reflects the learned knowledge of the task. Finally, the authors show that standard networks trained on the task approximate their proposed theory.

      Strengths:

      The structure of the work is clear, and the article is very well written, which is particularly noticeable given the consequential amount of results presented. The authors are able to link their theory to experimental findings in neural recordings. The reverse-engineering of trained recurrent neural networks is very thorough, by analyzing both dynamics and connectivity. The assumptions and predictions of their model are clearly stated.

      Weaknesses:

      I believe the article shows no major weaknesses. There are important points that are beyond the scope of this study, that may limit its impact. For instance, while very little previous literature has linked planning to RNNs, their proposed theory of "space-time attractors" in RNNs is about input-driven stable attractors. The authors clarify this in the text, highlighting differences between classic attractor networks and the mechanism studied here. Related to this, an assumption of this mechanistic theory is that rewards are flexibly rerouted to the corresponding neural populations after every action, which seems difficult to implement biologically. Both aspects are discussed in the current manuscript.

    3. Reviewer #2 (Public review):

      This well-written manuscript proposes to use attractors in space and time (STA) as a mechanistic explanation for planning in the prefrontal cortex. The main conceptual hypothesis is that planning is implemented as attractor dynamics in a representation that encodes states at each time step jointly. Depending on inputs the network relaxes to a trajectory that already contains future states that will be visited at each time step, rather than computing a scalar value at each point in time and space like other classical approaches from RL. The authors compare this approach to implementations such as TD learning and successor representation, and further show that trained recurrent neural networks on specific tasks involving planning develop structured subspaces resembling the ones postulated in STA.

      The idea of treating attracting trajectories unfolding in time as the computational substrate for planning is very interesting and potentially important. The explicit construction of a state x time representational space and its implementation via recurrent dynamics are appealing and convincing in the idealized tasks considered. I found the ms to be refreshingly explicit regarding several of the assumptions and limitations of the models, for example the fact that certain advantages can be viewed as properties of the state space itself and not necessarily of a fundamentally new planning mechanism.

      I thank the authors for their reply and their thorough rebuttal. It answered most of my previous questions and greatly enhanced the understanding of the paper.

      I have just two remaining concerns:

      (1) The ms shows attractor dynamics in the trained RNN during planning, but it is less clear how these relate to the execution phase. It would be helpful to clarify whether the network state during execution is expected to effectively be close to a FP or at a FP for each input, or whether the RNN implements transient dynamics shaped by the underlying attractor landscape.

      (2) Regarding the previously raised point of calling their result a "Mechanistic theory of planning", I did not mean to suggest that a theory cannot be mechanistic, or that "mechanistic theory" is not a valid term, especially in the context of this paper (although I believe this topic would deserve an entire separate discussion in the neuroscience field).

      My point was about whether STA should primarily be interpreted as a mechanistic theory of planning, or as a candidate neural mechanism for implementing the planning as inference theory. I am aware that mechanistic theory and mechanistic models are often used interchangeably in neuroscience, and I certainly do not claim that my interpretation is the only valid one. My opinion is that the manuscript presents a convincing and interesting candidate neural mechanism for planning, which can be strongly related to planning as inference. The reason why I am not fully convinced about the framing as a mechanistic theory of planning is mainly that the adjacency-based connectivity isn't emerging or derived, but is instead introduced based on practical and empirical considerations. It's not a major issue, but I would personally frame it as a mechanistic account or model of planning (and/or planning-as-inference), rather than a theory, mechanistic or not.

    4. Author response:

      The following is the authors’ response to the current reviews.

      We appreciate the additional clarifications suggested by the reviewers, and we will include these in the final Version of Record.


      The following is the authors’ response to the original reviews.

      Reviewing Editor Comments:

      The reviewers are very enthusiastic about this study, but have pointed out a central issue: are "space-time attractors" really attractors?

      The reviewers would be willing to increase the assessment of significance if the comments are properly addressed, and in particular, the issue about space-time attractors.

      We thank the editors and reviewers for the feedback on our manuscript and have revised the paper to address their questions and concerns. This document includes (i) an overview of the major changes to the paper, and (ii) point-by-point responses to the reviewers. We have also attached a version of the revised paper that highlights substantial changes to the text.

      Briefly, the reviewers asked for improved intuition about the STA representation and dynamics, and its relationship to attractor networks. To address their questions, we have restructured the paper. It now starts by introducing the STA, which has a representation and connectivity that are both handcrafted. The revised manuscript characterises the resulting dynamics and fixed points in more detail, both empirically and analytically. We then introduce a new model that directly optimises the fixed points of a neural network to represent an explicit plan of the future. The optimal weights for inferring such representations resemble the STA connectivity empirically. Finally, we analyse our unconstrained recurrent neural network, which learns both optimal representations and connectivity. As also shown in the original paper, this network learns to implement an algorithm that closely resembles an STA. Together, these results show that attractor networks can infer PFC-like representations of the future, and this is an efficient solution to dynamic planning problems known to depend on PFC.

      RE1: Expanded theory of STA dynamics

      We have now formalised how the STA relates to a formulation of planning as an inference process over future trajectories, which has been previously proposed in cognitive science and reinforcement learning. We show in the revised paper that the STA dynamics resemble an algorithm for approximate inference in the corresponding probabilistic graphical model. This allows us to characterise the fixed points of the algorithm analytically and relate them directly to a well-established cognitive theory of planning. These analyses help bridge the gap between neural implementation and cognitive computation. They shine new light on previous results in the paper while also providing more intuition for the STA dynamics.

      We have also included a new model that directly optimises the fixed points of an attractor network to resemble a posterior distribution over future locations from planning-as-inference. This analysis complements the handcrafted STA, where we impose both the representation and connectivity, and the RNN, where both the representation and connectivity are learned. The new model imposes (i) an explicit spacetime representation, and (ii) the multiplicative structure of a message passing algorithm. We then train the weights associated with the forward and backward messages by gradient descent on the KL divergence between (i) the true posterior marginals and (ii) the approximate distribution over future locations implied by the network representation at the fixed point. Supplementary Figure S2 of the revised manuscript shows that the optimal weights reflect the transition structure of the environment, similar to the handcrafted STA model and the task-optimised RNN. This makes the connection between attractor dynamics and planning-as-inference more explicit by showing that the fixed points of an attractor network can be optimised directly for planning.

      RE2: Improved characterisation of fixed points

      We have clarified how and why the STA is an attractor network. Attractor networks are defined by the existence of stable fixed points. In ring and grid attractors, there is a continuum of such fixed points in the absence of structured inputs (but often with tonic excitation). In contrast, the STA has a discrete set of input-dependent fixed points. We show explicitly in the revised manuscript how these fixed points depend on the reward inputs to the network, and also how they relate to planning-as-inference.

      We are not claiming that the STA is exactly equivalent to continuous ring and grid attractors. Instead, we want to convey the intuition that the connectivity of the STA constrains the possible fixed points to be plausible trajectories through space and time. The reward inputs determine which of these possibilities is an actual fixed point in a given planning problem. This is not unlike ring attractors in the presence of strong visual inputs. The connectivity enforces a single bump of activity, and the visual input ‘yokes’ the bump to an appropriate orientation. These similarities and differences are highlighted in the revised paper.

      Finally, we have added a new Figure 3 to the main text that characterises the STA fixed points empirically. This figure:

      (a) Shows the evolution of the STA dynamics and convergence to different fixed points in different environments (panels A-B).

      (b) Shows that the network can converge to different fixed points on different trials in the same environment. This happens when there are multiple equally good paths to a goal (panels B-D).

      (c) Shows that other fixed points also exist that correspond to longer trajectories, but the dynamics of the network bias it towards representations of shorter paths. The STA reliably converges to fixed points representing longer trajectories if it is initialised within their basin of attraction (panel F).

      Updated main text:

      “Unlike ring and grid attractors, the fixed points of the spacetime attractor depend on tonic inputs. However, the connectivity constrains the fixed points to represent continuous trajectories for any combination of inputs. In this section, we show this empirically. Later, we will see that such connectivity is optimal for planning-as-inference.

      To compute a plan, it is necessary to know which states will be rewarding in the future. This reward information is provided as an input to the STA and enables fast adaptation without rewiring the synaptic connections. It alters the fixed points of the recurrent dynamics to only include trajectories that are also associated with high cumulative reward (Figure 3; Methods).”

      RE3: Ground truth rewards as an input to the network

      Both reviewers asked about the external input to the STA that specifies the reward available at different states in the future. In reinforcement learning and cognitive science, ‘planning’ is usually defined as the problem of computing a trajectory that maximises cumulative future reward, given an initial state and a reward function (e.g. Mattar & Lengyel, 2022). This is similar to many real-life situations, where we have a known but distant goal (win a game of chess, finish our paper before a deadline, …). When such a reward function is known, it remains challenging to determine the sequence of actions to get there. This has been the topic of much previous work in neuroscience, including (i) the successor representation, which combines a trial-specific reward function with stable transition statistics; and (ii) different types of sequential search, which use a known reward function to evaluate different possible future trajectories.

      To highlight the importance of planning, even when the reward function is known, Figure 4 of the revised manuscript shows that the STA performs better than a greedy baseline that acts according to the immediate reward input instead of planning to maximise cumulative reward. Planning is therefore distinct from learning or inferring a reward function, which is itself a major open question in cognitive science. While undoubtedly interesting, a solution to this problem is beyond the scope of our paper. That is why we decided to simply provide ground truth rewards as an input to the STA. We have clarified the distinction between planning and ‘reward learning’ in the revised paper, and the supplementary material now includes a discussion of where the reward input to the STA could come from.

      Author response image 1.

      All performance quantifications in Figure 4 now include an additional ‘greedy’ baseline (grey bars). This is an agent that acts according to the immediate future reward. The performance improvement of the STA over this baseline highlights the importance of planning to maximise cumulative reward.

      Updated main text:

      “There are several possible sources of reward input to a spacetime attractor (Supplementary Note). We focus on planning under a known reward function and therefore assume access to ground-truth rewards.”

      Reviewer #1 (Public review):

      Summary:

      This work builds a theory to implement planning trajectories towards a goal in a known environment, inspired by analyses of prefrontal neural recordings. Unlike standard neural architectures for this task, such as value-based learning and successor representations, their proposed theory is able to adapt to novel goal locations within a trial. The key to the theory is that future times are represented by orthogonal groups of neurons. The recurrent connectivity between groups of neurons selective to specific future times and locations reflects the learned knowledge of the task. Finally, the authors show that standard networks trained on the task approximate their proposed theory.

      Strengths

      The structure of the work is clear, and the presentation of the results is very well written, which is particularly noticeable given the consequential amount of results presented. The authors are able to link their theory with experimental findings in neural recordings. The reverse-engineering of trained recurrent neural networks is very thorough, by analyzing both dynamics and connectivity. The assumptions and predictions of their model are clearly stated.

      We appreciate the encouraging comments and hope our revised manuscript addresses the reviewer’s questions.

      Weaknesses

      (1.1) It is unclear whether their proposed theory, "space-time attractors", actually is an attractor network. The authors used recurrent neural networks with very few timesteps, and long single neuron time constants with respect to the task time scales. Attractor networks, as the ones the authors cite, refer to networks that generate nontrivial patterns of activity through recurrent interactions, after long periods of time.

      See RE1 & RE2 for a comprehensive response to this question. Briefly, we show in the revised manuscript how the fixed points of the STA dynamics relate to planning-as-inference, and we clarify the similarities and differences between the STA and other attractor networks in the main text. We show in the new Figure 3 that (i) representations of future paths are stable over long periods of time, and (ii) multiple fixed points can exist when there are multiple paths to the goal. It is also worth noting that the RNN representation in Figure 6H remains stable for 75 time constants and recovers from perturbations. This is substantially longer than during training, where ‘planning’ lasted up to 14 network time constants, and it suggests that the network representation is a stable fixed point.

      (1.2) The authors gloss over how the reward inputs are calculated. Computing these reward inputs should be part of the planning process, and the authors are implicitly leaving this problem aside. How does the reward input, which includes future time and location, depend on the actions that have not yet been taken by the agent? It feels like most of the planning computation is already provided by these reward inputs at the beginning of the trial. It could be that the network is only learning to process the planned sequence of actions present in the inputs.

      See RE3 for a comprehensive response to this question. Briefly, ‘planning’ is often defined as the problem of computing a trajectory that maximises future reward, given a reward function, initial state, and transition function. The reward function provided to the agent indicates which future states it would be desirable to reach, but not how to reach them. To make this point clearer, we show in Author response image 1 that the representations computed by the STA generate better behaviour than an agent acting greedily according to the reward function specified by the inputs. This highlights the importance of considering distant goals when choosing immediate actions.

      Reviewer #1 (Recommendations for the authors):

      The text is very nicely written, and I appreciated the way in which methods are presented, with a clear structure and a logical chaining of the different sections. My comments and suggestions refer mostly to the methods and the RNN implementation. Please find below a list of issues.

      Relatively major:

      (1.3) All the equations of the dynamics should be written in discrete time and not in continuous time. There is no notion of "iteration" in continuous time, so it is currently very hard to understand how the RNN works, and what the different epochs are ("the RNN performed 10 network iterations..."?).

      We have rewritten all equations in discrete time and clarified the notion of ‘iterations’.

      (1.4) This is pointed out in the public review, but the authors insist on making an analogy between the spacetime attractor implementation of planning and attractor networks. It seems to me that these two types of models are very different. What defines attractor networks (such as grid- or ring-attractor networks) is that recurrent connections internally generate stable states of activity for long periods of time, in the absence of inputs. Nothing like that is shown here. Robust input-driven trajectories are neither necessary nor sufficient for showing that an RNN is an attractor network. The fact that, given the inputs, networks are run for very short periods of time in this work seems to indicate that this is a very different type of network compared to the attractor networks mentioned previously.

      See RE1 & RE2 for a comprehensive response to this question. Briefly, it is correct that the fixed points of the STA depend on the inputs, which is different from canonical ring and grid attractors.

      We show that the fixed points of the recurrent STA dynamics are reward-maximising paths when conditioned on those inputs. Briefly, we now (i) analytically characterise the input-dependence of the fixed points and show how they relate to planning-as-inference; and (ii) show empirically that STA representations remain stable for long periods of time (Figure 3). The revised paper also clarifies the similarities and differences to previous attractor models.

      Minor

      (1.5) It would be nice to show more clearly what the inputs are in a given trial, and how they change over time in the RNN (specifying the planning and execution phases). A supplementary figure may help.

      How is the information about walls provided to the network exactly? More generally, it would be nice to clearly indicate in the Methods what all the inputs "x" to the RNN are, and how they change over time (during planning, during execution, and how they change as the environment steps are updated).

      We have made a new Supplementary Figure S3 that illustrates the inputs to and outputs from the different models. Briefly, the information about the walls is provided to the RNN as a binary vector x<sub>w</sub> ∈ ℝ <sup>2𝑁</sup>. The elements of this vector indicate for each of the N states whether there is a wall (i) to the right of, and (ii) above it. In each trial, a subset of these is present, and a subset is absent. All ‘present’ walls are assigned a value of +1 in x<sub>w</sub> , and all absent walls are assigned a value of 0 in x<sub>w</sub> . We have also clarified this in the revised Methods.

      (1.6) It would help to clarify, at least in the methods, the shape of all the matrices and vectors that are trained.

      We have clarified the shapes of all matrices and vectors in the Methods.

      (1.7) N in the methods is not defined (I think N = 16, the total number of locations on the grid).

      N is indeed the total number of locations in the state space. This is 16 for almost all analyses in the paper, which involve planning on a 4x4 grid. The updated manuscript includes a few analyses in larger environments, where N is larger. We have clarified this in the Methods.

      Reviewer #2 (Public review):

      This well-written manuscript proposes to use attractors in space and time (STA) as a mechanistic explanation for planning in the prefrontal cortex. The main conceptual hypothesis is that planning is implemented as attractor dynamics in a representation that encodes states at each time step jointly. Depending on inputs, the network relaxes to a trajectory that already contains future states that will be visited at each time step, rather than computing a scalar value at each point in time and space like other classical approaches from RL. The authors compare this approach to implementations such as TD learning and successor representation, and further show that trained recurrent neural networks on specific tasks involving planning develop structured subspaces resembling the ones postulated in STA.

      The idea of treating attracting trajectories unfolding in time as the computational substrate for planning is very interesting and potentially important. The explicit construction of a state x time representational space and its implementation via recurrent dynamics are appealing and convincing in the idealized tasks considered. I found the manuscript to be refreshingly explicit regarding several of the assumptions and limitations of the models, for example, the fact that certain advantages can be viewed as properties of the state space itself and not necessarily of a fundamentally new planning mechanism.

      Overall, the manuscript presents a cool attractor model that extends in time and explores its performance in a subset of illustrative tasks involving planning. My doubts concern mostly the interpretation and scope of the claims made in the manuscript. Here are a few comments where I detail my questions/concerns:

      We appreciate the enthusiasm about the manuscript and its potential importance. We address the remaining questions and concerns below.

      (2.1) The authors nicely discuss that much of the difference between STA and classical TD or SR agents is "in some sense a property of the state space rather than the decision making algorithm," and that TD and SR could in principle be implemented in a comparable space x time representation. This is fair, but it also suggests that the central contribution of the manuscript lies primarily in the representational factorization (state x time tiling) and its dynamical implementation via attractors, rather than in a fundamentally new planning algorithm or theory, mechanistic or not. I think theory should be distinguished from mechanism, and it would therefore help the reader to describe the conceptual advancement more as a novel mechanism or implementation than a novel (mechanistic) theory for decision/planning.

      We respectfully disagree that ‘theory’ has to be distinguished from ‘mechanism’. We do agree that ‘computation’ and ‘mechanism’ can often be distinguished. However, we think theories can live at either of these (and other) levels of explanation. What we propose is indeed a potential mechanism for planning that combines recently characterised prefrontal spacetime representations with attractor dynamics to infer desirable ‘plans’. As we show in the revised manuscript, this mechanism resembles the computation of ‘planning-as-inference’, which has previously been proposed in cognitive science (e.g. Botvinick & Toussaint, 2012). Our theory is therefore not about the computation – it is about the mechanism. The title “A mechanistic theory of planning…” is meant to clarify what level of description our paper addresses.

      As an example of the importance of mechanistic theories, the computation of angular velocity integration can be implemented in many different ways. Seminal work by Skaggs et al. (1994) and others in the 1990s showed how it can be implemented in neural networks, inspired by experimental data. These theories paved the way for detailed experimental characterisations of the fruit fly head direction circuit more than two decades later (Turner-Evans et al., 2017; Kim et al, 2017; and others). Inspired by this and other success stories, we think an important role of theoretical neuroscience is to develop theories about neural mechanisms that can be tested in future experiments!

      (2.2) Related to my previous point, I think it would be helpful to position STA more explicitly relative to computational/theoretical literature in which attractor networks encode temporally ordered patterns (so effectively including future times). For example, classical extensions of Hopfield networks with asymmetric connectivity implement retrieval of sequences and ordered transitions between patterns (Sompolinsky & Kanter, 1986). More recently, sequential attractors and limit-cycle dynamics have been constructed in structured recurrent networks by the Morrison group (Parmelee et al., 2021). These works do not implement an explicit discretized state x future-time tiling as in STA and do not specifically discuss the usage for planning. However, they do provide concrete precedents for attractor dynamics over temporally structured trajectories in terms of mechanism. It would be useful to discuss this literature and clarify a little what's new mechanistically in the view of the authors.

      We agree that this is not the first use of attractor networks to represent or compute sequences. Instead, we show that a combination of spacetime representations with attractor dynamics is sufficient to compute plans in dynamic problems known to depend on prefrontal cortex. As the reviewer points out, the primary difference from most previous work lies in the fact that the entire sequence is encoded in a single fixed point of the STA dynamics. This differs from e.g. Sompolinsky & Kanter, where the population encodes one element at a time and generates sequences as limit cycles. The instantaneous encoding of an entire sequence in the STA is what enables planning through parallel message passing rather than sequential search. This is highlighted in the main text of the revised manuscript, which also includes a supplementary discussion of the similarities and differences between the STA and related work on sequences in attractor networks.

      Updated main text:

      “Entorhinal grid cells are also embed a world model in their connectivity (McNaughton et al., 2006), but they only encode a single location at a time (Vollan et al., 2026). Such networks can generate sequences, but the individual elements are represented one by one (Sompolinsky and Kanter, 1986; Kleinfeld, 1986; Widloski et al., 2025). The spacetime attractor suggests that circuit principles in prefrontal cortex resemble other cortical areas that use structural knowledge to infer features of the world. The major difference is that PFC instantaneously represents many points in time, which generalises known circuit principles to complex planning.”

      (2.3) A central claim of the manuscript is that space-time trajectories are attractors of the STA dynamics. The manuscript does provide empirical evidence consistent with attractor-like behavior. However, it is not explicitly shown whether trajectory representations persist in the absence of sustained external inputs. So it's not clear to me whether the trajectories should be interpreted as intrinsic attractors of the recurrent system, which can be selected by delivering transient inputs, or whether they must be stabilized by a specific continuous external drive. It would be useful if the author could clarify/discuss this point.

      We show in the revised paper that the fixed points of the STA dynamics take the form r <sub>δ</sub> = e<sup>R<sub>δ</sub></sup> ◦ (Ar<sub>δ−1</sub>) ◦ (A<sup>T</sup> r<sub>δ+ 1</sub>) (Methods). Here, r δ is the activity of neurons representing expected locations in δ actions; R δ is the reward function in δ actions; and A is the environment adjacency matrix. These fixed points depend on the reward inputs through the first term. In the absence of reward inputs, the fixed points are ‘diffusive’, while still respecting the transition structure of the environment. In the presence of reward inputs, they concentrate probability mass on trajectories with high expected reward. We have clarified these properties in the main text and introduced a new Figure 3 that characterises the fixed points of the STA in more detail. See also RE1 and RE2.

      (2.4) As far as I understand it, reward information is provided as input to specific populations encoding future time steps, and that's essential for rapid adaptation without rewiring connectivity. How such future-time-specific reward inputs would be generated and routed to distinct neural populations isn't entirely clear to me. Since this seems to be an essential component of the model, I think it would be important to discuss more deeply the source and plausibility of these reward signals related to different timesteps.

      See RE3 for a comprehensive response to this question. Briefly, ‘planning’ is often defined as the problem of computing a trajectory that maximises future reward, given a reward function, initial state, and transition function. The reward function provided to the agent indicates which future states it would be desirable to reach, but not how to reach them (see Author response image 1). We agree that the challenge of estimating future reward is an interesting question, but it is beyond the scope of this paper. We have clarified this distinction in the main text and added a supplementary discussion that speculates about where reward information could originate in biological circuits.

      (2.5) The authors note that vanilla STA scales linearly with planning horizon, and discuss potentially hierarchical extensions for longer horizons. They acknowledge that learning abstractions remains an open challenge, yet the examples of planning in the manuscript are restricted to very short temporal horizons and limited branching complexity. It is not obvious to me in what cases the current implementation and interpretation of STA remains viable (for example, in terms of relaxation iterations) as the horizon and branching factor increase. Relatively simple planning can be managed by simpler, less costly models/algorithms, whereas complex planning is a lot harder to deal with, and it's something that a mechanistic "theory" should address. In the context of the claims of the paper in its present form, I think this is possibly the most important conceptual and practical limitation in the manuscript.

      It is correct that planning gets increasingly challenging with planning depth. In the absence of noise, the STA scales to sequences of up to 12-13 actions – and even longer if the minimum path length is known a priori. Performance gets progressively worse when recurrent activity and parameter noise increase. The revised paper includes a new Supplementary Figure S1A-C that shows how the STA planning ability depends on planning depth for different levels of noise.

      We do not consider planning depth to be a major limitation of the work, since humans are rarely thought to plan much more than 6 steps into the future at a single level of abstraction (e.g. van Opheusden et al., 2023). Instead, we believe that hierarchical planning is used to infer trajectories to distant goals (Eckstein & Collins, 2020). To illustrate this point, we have now implemented a proof-of-principle hierarchical STA in Supplementary Figure S1D-E. This simulation shows how an ‘abstract plan’ inferred by one STA can be treated as a goal to infer a more ‘detailed plan’ in a second STA. In principle, this enables the system to compute plans that are arbitrarily long, provided they can be broken down into chunks smaller than the limits imposed by the analyses in Supplementary Figure S1A-C.

      Finally, RNNs learn an STA-like algorithm when trained on dynamic planning problems with a planning depth of 6. It is therefore not clear to us whether simpler and less costly algorithms can be easily implemented in the dynamics of recurrent networks.

      (2.6) The RNN analyses show that trained networks develop structured subspaces aligned with future time indices and exhibit perturbation behavior consistent with attractor-like dynamics. The manuscript also explicitly notes differences between the trained RNN and the handcrafted STA (e.g., long-range couplings between subspaces and differences in behavior of lower-value trajectories under perturbation), which I much appreciated. My doubt is on the specificity of this result, as trained RNNs on fixed-horizon tasks can develop latent dimensions correlated with temporal progress within a trial or time-to-goal. I think it would help the reader to clarify whether the results demonstrate that STA-like computations emerge in RNNs trained on planning tasks, or that RNNs generally develop some kind of structured spacetime representations when tasks involve future timesteps and some degree of flexibility in the decisions.

      An important point to note is that the subspaces we identify do not encode time-to-goal, since they are all active at the very beginning of the trial. We also show that RNNs trained on simpler static tasks do not learn the same algorithm (Supplementary Figure S7) and do not generalise to dynamic problems (Figure 5F). Finally, other algorithms are capable of solving the dynamic problems we study (e.g. the ‘value agent’ in Figure 5B-E). We therefore do not think it is trivial that RNNs learn an STA-like algorithm.

      We do think that ‘structured spacetime representations’ generally emerge in RNNs trained on tasks that involve flexible behaviour in changing environments – in some sense that is the claim we are trying to make. It is known that spacetime representations are optimal for structured sequence memory tasks (e.g. Whittington et al., 2025; Dorrell et al., 2026), and we think this is for exactly the same reason. In sequence working memory, the reward function changes in time – for each action, the reward is only non-zero at the corresponding sequence element. However, the adjacency matrix is uniform for sequence memory – any sequence element can follow any other sequence element – so there is no need for planning. We are therefore not claiming that spacetime representations only emerge in the specific planning task we consider here. Instead, we expand the set of problems solvable by such representations to also include adaptive planning known to depend on prefrontal cortex. We have made this more explicit in the revised manuscript.

      Updated main text:

      “Together, our analyses show that RNNs trained on a dynamic planning task learn to approximate a spacetime attractor. This was also true across variations in model architecture (Methods; Figure S10; Figure S11). These results extend previous findings that explicit spacetime representations are optimal for sequence memory (Supplementary Note; Whittington et al., 2023; Dorrell et al., 2026; Wang et al., 2025). Additionally, RNNs with too few hidden units to learn a spacetime attractor failed to solve the task (Figure S12), suggesting that other solutions are not readily learned by gradient descent.”

      A few more minor points, mainly concerning clarity:

      (2.7) The main dynamical equation combines a log-domain recurrent term, a floor operation, and a log-sum-exp normalization step, followed by exponentiation. The intuition/logic behind this specific formulation could be clarified for the reader. For example it would be helpful to explain why the recurrent input appears inside a log, and also whether/how these operations relate to any multiplicative constraint.

      The specific form of these equations comes from the intuition that the STA approximates planning as an inference process over future trajectories. We have clarified this in the revised manuscript, which explicitly shows how these equations relate to planning-as-inference as formulated previously (e.g. Botvinick & Toussaint, 2012; Levine, 2017).

      (2.8) While the computational cost of successor representation in an expanded NT x NT representation is discussed, the corresponding scaling of STA in terms of number of units and connections (as a function, for example, of the planning horizon) isn't clear to me. Perhaps the authors could compare costs more explicitly.

      The memory cost of a spacetime-SR would be (NT)^2 and the computational cost (NT)^3 (it is possible that both of these could be reduced by taking advantage of the structured nature of the spacetime successor matrix, but that is beyond the scope of this work). The memory cost of the STA is NT, and the computational cost is (NT)^2 (each iteration of the network dynamics requires the calculation of T matrix-vector products of size NxN, and the number of steps to convergence is approximately linear in T). We have included this comparison in the Supplementary Discussion of the revised paper.

      (2.9) In the RNN analyses, structured subspaces aligned with future time indices are shown. I couldn't find a quantification of how much variance is captured by the subspaces, relative to other latent dimensions. Adding it would help get a feeling for the strength of the alignment.

      We have added a new Supplementary Figure S6 to the revised manuscript, which quantifies the variance explained by the future-coding subspaces over the course of a trial. The variance explained by the K dimensions encoded by these subspaces is substantially higher than a random baseline, and it approaches the upper bound given by the top K PCs. Interestingly, the future-coding subspaces all explain a lot of variance early in the execution period. During later stages of execution, only the ‘immediate future’ subspaces explain substantial variance. This suggests that the RNN only maintains information in subspaces that represent times before the end of the trial.

      References

      Botvinick, Matthew, and Marc Toussaint. "Planning as inference." Trends in cognitive sciences 16.10 (2012): 485-488.

      Dorrell, William, et al. "An Efficient Computing Theory of Prefrontal Structured Working Memory Representations." bioRxiv (2026): 2026-02.

      Eckstein, Maria K., and Anne GE Collins. "Computational evidence for hierarchically structured reinforcement learning in humans." Proceedings of the National Academy of Sciences 117.47 (2020): 29381-29389.

      Kim, Sung Soo, et al. "Ring attractor dynamics in the Drosophila central brain." Science 356.6340 (2017): 849-853.

      Levine, Sergey. "Reinforcement learning and control as probabilistic inference: Tutorial and review." arXiv preprint arXiv:1805.00909 (2018).

      Mattar, Marcelo G., and Máté Lengyel. "Planning in the brain." Neuron 110.6 (2022): 914-934.

      Skaggs, William, et al. "A model of the neural basis of the rat's sense of direction." Advances in neural information processing systems 7 (1994).

      Turner-Evans, Daniel, et al. "Angular velocity integration in a fly heading circuit." Elife 6 (2017): e23496.

      Van Opheusden, Bas, et al. "Expertise increases planning depth in human gameplay." Nature 618.7967 (2023): 1000-1005.

      Whittington, James CR, et al. "A tale of two algorithms: Structured slots explain prefrontal sequence memory and are unified with hippocampal cognitive maps." Neuron 113.2 (2025): 321-333.

    1. eLife Assessment

      This is an important work, using preclinical models and reporting that a shared inflammatory and vascular leakage response following Treg depletion, UVB irradiation, and DNFB contact hypersensitivity facilitates outgrowth of premalignant melanocytes. However, some of the work to support the claims is incomplete; this pertains particularly to the quantification of melanocyte expansion and sample size inconsistencies. With some additional supporting evidence to strengthen the claims, the study will be of broad interest to cancer biologists and immunologists.

    2. Reviewer #1 (Public review):

      Summary:

      The immune system represents a source for melanocyte-extrinsic determinants of melanoma. Multiple immune cell types, including natural killer cells and CD8+ cytotoxic T lymphocytes, destroy cancer cells directly, and this anti-tumor activity is widely thought to eliminate many nascent tumors before they become clinically detectable. Regulatory T (Treg) cells, which are defined by expression of the transcription factor Foxp3, function as critical suppressors of lymphocyte activation in both homeostatic and disease contexts. Intratumoral Treg cell accumulation has been associated with disease progression in the clinic, and Treg cell depletion inhibits melanoma outgrowth in transplantable models of the disease.

      However, the role of Treg cells during the early, premalignant stage of melanocyte expansion has not been examined. To study the interplay between incipient melanoma and cutaneous inflammation, the authors subjected an autochthonous murine model of melanoma [LSL-BrafV600E;Ptenfl/fl;Tyr::CreERT2 (BPT)mice] to three distinct inflammatory immune perturbations. Each of these perturbations accelerated premalignant melanocyte outgrowth, which was unexpected given that both Treg cell depletion and DNFB treatment markedly enhanced conventional T (Tconv) cell infiltration into the skin. Detailed analysis of each inflammatory response revealed a shared cellular and molecular signature comprising myeloid infiltration, characteristic cytokines and tissue remodeling factors, and vascular permeability. Altogether, the results support the hypothesis that oncogenic mutations in melanocytes along with altered immune response and inflammation synergistically drive melanomagenesis.

      Strengths:

      (1) The use of the three distinct inflammatory immune perturbations, such as transient Treg cell depletion, acute UV- B irradiation, and 2,4-dinitrofluorobenzene (DNFB)-induced contact hypersensitivity.

      (2) The re-examination of the role of Treg cells in transplantable models of tumor growth by Subcutaneous (s.c.) implantation of syngeneic B16F10 melanoma cells as widely used to study anti-tumor immunity and to interrogate the effects of Treg cells on tumor suppression.

      (3) The use of the LSL-TdTomato mice for measuring the TdTomato fluorescence in each immune cell type to assess uptake of melanocyte antigen.

      (4) scRNA-seq data document a broad and rapid myeloid inflammatory response in Treg cell-deficient skin.

      Weaknesses:

      (1) The expression of inflammatory mediators Il1b, Il6, and TNFa, and angiogenic mediators Hif1a and Ang2 in all of the three models of immune perturbation has been verified at the transcript level by qRT-PCR, and it remains to be determined whether it correlates with the same at the protein level.

      (2) scRNA-seq data to profile the diversity of immune cells (CD45+) in the ear skin of the DNFB-treated contact hypersensitivity model are currently missing.

      (3) The authors indicate that at least 2 prior studies directly implicated inflammatory macrophages in the melanocyte proliferation response. However, no attempts were made for the identification of the UVB-driven factors underlying myeloid recruitment by which these cells activate melanocytes.

      (4) It is very surprising to observe that altered immune responses, such as enhanced Tconv cell priming and activation in Treg cell-deficient skin, failed to antagonize mutant melanocyte outgrowth in the BPT model. While UVB-induced skin inflammation differs somewhat from the response to Treg cell depletion, both feature the infiltration of tissue remodeling macrophages and vascular instability. The findings that contact hypersensitivity can promote the expansion of non-malignant BRAF (V600E)Pten- Het melanocytes have interesting implications for benign hyperpigmentation conditions, such as post-inflammatory hyperpigmentation, Riehl's melanosis, and melasma.

    3. Reviewer #2 (Public review):

      Summary:

      Tran and colleagues investigate how inflammation alters the earliest stages of melanoma tumorigenesis in mice carrying LSL-BrafV600E, Ptenfl/fl, and Tyr-CreERT2 alleles. They compare transient regulatory T cell depletion, acute UVB irradiation, and DNFB-induced contact hypersensitivity. Each perturbation increases ear pigmentation and Tyrp1 expression after oncogene induction. The inflammatory settings also share recruitment of monocytes and macrophages, expression of inflammatory and tissue-remodeling programs, and increased vascular permeability. Dexamethasone attenuates the DNFB-associated phenotype. A secondary finding of particular interest is that regulatory T cell depletion accelerates the premalignant BPT phenotype but inhibits B16F10 tumor growth, suggesting that regulatory T cells can have different effects during tumor initiation and established transplantable disease.

      The study addresses an important question that is difficult to approach using transplantable tumor models. The data convincingly show that each perturbation produces substantial inflammation in the skin and that vascular leakage accompanies the response. At present, though, the central biological endpoint is not sufficiently separated from melanogenesis. Darkening of the ear and increased Tyrp1 RNA can reflect more pigment or altered differentiation within the existing oncogene-carrying melanocytes rather than an increase in their number, particularly given that pigment content is itself variable in transformed melanocytes, which range from heavily pigmented to nearly amelanotic. This issue is especially important in the UVB and DNFB experiments, where inflammatory signals can alter pigmentation directly.

      Strengths:

      The autochthonous BPT model is a major strength. It preserves the native relationship between melanocytes and the surrounding stromal and immune compartments during lesion initiation. Including three distinct inflammatory perturbations makes the recurring association with melanocyte-associated readouts more persuasive than any single model would be. The paired-ear DNFB design is efficient and controls for inter-animal variability. The combination of flow cytometry, single-cell RNA sequencing, intravital imaging, and Evans Blue assays provides useful complementary evidence that the inflammatory interventions remodel the local tissue environment. The B16F10 experiments help establish that the unexpected effect of regulatory T cell depletion is specific to the early autochthonous setting rather than a general failure of the depletion model. The BT-Het experiment is also thoughtful in asking whether inflammation can enhance the phenotype of oncogene-carrying melanocytes in a nevus-stage context that does not proceed to full malignant progression after oncogene induction alone.

      Weaknesses:

      The strongest caveat concerns the central claim. The outgrowth readouts are ear darkening and bulk Tyrp1 expression, but both may report pigment or differentiation state rather than the number of oncogene-carrying melanocytes. Pigment content is not a reliable proxy for cell number here, since the same population can darken or lighten without any change in cell number. No direct count or lineage-reporter measurement is provided for the regulatory T cell, UVB, or DNFB comparisons. Until that gap is filled, the data support increased pigmentation of oncogene-carrying melanocytes more firmly than the premalignant expansion named in the title, and this concern is most pronounced in the UVB and DNFB settings, where inflammation can change pigmentation on its own.

      Secondly, the proposed shared mechanism is largely associative. Dexamethasone appropriately shows that inflammation as a whole is required for the DNFB phenotype, but as a broad anti-inflammatory it cannot isolate any single component. The manuscript singles out blood vessel remodeling as particularly important, and that specific attribution exceeds what a non-selective drug can show, especially as no individual pathway is selectively blocked in a tumor-initiation experiment and Il6 is reduced only modestly. The authors acknowledge that the precise chain of causation is unresolved, so the vascular claim should be softened to match or tested directly.

      Also, several of the mechanistic conclusions rest on thin or single cohorts and on single-cell data whose replication is not fully reported, making them less convincing than the inflammatory phenotypes themselves. The systemic regulatory T cell model shows the consequences of body-wide depletion rather than a skin-specific regulatory T cell function, and the inferred monocyte-to-macrophage trajectory reflects transcriptional similarity rather than a demonstrated lineage path. The interpretation of dendritic-cell TdTomato uptake as evidence of antigen presentation or T cell priming is not supported by a direct measure of reactivity.

      Finally, the nevus-stage framing should be corrected. The manuscript frames the BT-Het experiment as testing non-oncogenic conditions, but those melanocytes carry BrafV600E, so it is better read as inflammation-enhanced behavior of oncogene-carrying melanocytes at the nevus stage.

    4. Reviewer #3 (Public review):

      Summary:

      Tran et. al. investigate how inflammation in the skin influences the early stages of melanomagenesis. They use an autochthonous, tamoxifen-inducible mouse melanoma model (LSL-BrafV600E;Ptenfl/fl;Tyr::CreERT2, "BPT") to examine three inflammatory perturbations: transient depletion of regulatory T cells, acute ultraviolet-B irradiation, and contact hypersensitivity induced by 2,4-dinitrofluorobenzene (DNFB). They report that each perturbation promotes the recruitment of immune cells, especially inflammatory monocytes and macrophages, increased expression of inflammatory and tissue-remodeling factors, and enhanced vascular permeability, which ultimately increases the outgrowth of premalignant melanocytes measured by local pigmentation and expression of the melanocyte-associated gene Tyrp1. In the DNFB model, the authors showed that treatment with dexamethasone reduces the effects of contact hypersensitivity on pigmentation, inflammatory gene expression, and vascular leakage, potentially providing a translational angle.

      Strengths:

      An interesting observation is that transient Treg depletion promotes premalignant melanocyte outgrowth in the autochthonous BPT model while inhibiting the growth of transplantable B16 F10 tumors. This contrast is consistent with a role for Tregs in limiting inflammatory disruption of the skin during early tumorigenesis and highlights the value of autochthonous models. These findings may also have broader implications for understanding the stage- and context-dependent functions of Tregs in cancer.

      Another strength of this manuscript is the comparison of three mechanistically distinct inflammatory perturbations. Treg depletion, UVB irradiation, and DNFB-induced contact hypersensitivity engage different inflammatory pathways but converge on myeloid-cell recruitment, inflammatory gene expression, and increased vascular permeability. This convergence strengthens the conclusion that an acute inflammatory microenvironment is associated with enhanced melanocyte outgrowth during the early premalignant phase.

      Weaknesses:

      The paper convincingly establishes a correlation between the inflammatory signature and melanocyte outgrowth across three distinct perturbations. However, the mechanistic claim that myeloid cells and/or vascular remodeling drive melanocyte expansion rests primarily on the dexamethasone experiments in the DNFB model. Because dexamethasone broadly affects immune, stromal, endothelial, and melanocytic compartments, these experiments do not establish that inflammatory monocytes/macrophages or vascular destabilization are specifically required for the melanocyte response.

      A related limitation is that the proposed monocytic origin of the inflammatory macrophage population following perturbation remains inferred. Although the scRNA-seq data and pseudotime analysis in Figure 4 - Supplement 2 are consistent with a trajectory from monocytes to macrophages, they do not exclude local reprogramming of resident macrophages into an inflammatory state. This alternative is particularly relevant because resident macrophage populations have been implicated in vascular remodeling and tumor outgrowth (PMIDs: 36493773 and 40216154).

      Finally, the assessment of melanocyte outgrowth is largely through increased pigmentation and whole-ear Tyrp1 expression. Although these measurements may reflect increased melanocyte abundance, they may also be influenced by melanogenic activity or increased Tyrp1 expression per cell. More direct evidence of melanocyte proliferation, such as Ki67 or EdU/BrdU staining specifically within TdTomato-positive melanocytes, would support the use of "expansion" and "proliferation" throughout the manuscript. Histopathological characterization of lesion architecture, atypia, proliferation, and invasion would also help establish the premalignant nature of the lesions.

    1. eLife Assessment

      This useful study reports the application of targeted RNA-probe hybridization-capture metagenomic sequencing to characterize Klebsiella pneumoniae detected in post-mortem lung tissue from infants and young children who died from respiratory illness in Lusaka, Zambia. The findings provide a convincing proof-of-concept for the use of targeted sequencing to recover clinically relevant genomic information from archived post-mortem tissue specimens. However, the evidence supporting broader epidemiological and public health conclusions should be interpreted with caution given the small number of cases, the limitations inherent to culture-independent sequencing, and the fact that the study was not designed to determine the source or setting of acquisition.

    2. Reviewer #1 (Public review):

      Summary:

      This study described the genomic features of K. pneumoniae in tissue samples from children in Zambia who had died from pneumonia, where K. pneumoniae was believed to be the causative organism. This was a substudy of a previous study which described more broadly the aetiologies of community-acquired pneumonia in this population, using minimally invasive tissue sampling.

      In this study, the authors used culture-independent molecular tools to characterise the K. pneumoniae genomes from post-mortem tissue samples from 7 children who had died of pneumonia. They described a diversity of strain lineages in the 7 children, with different capsular types and virulence profiles. Notably, all strains had a wide array of antimicrobial resistance genes. This is probably not surprising given the propensity of this organism to acquire AMR genes, but it is concerning in isolates from community-acquired infections.

      The study also highlights the value of using advanced culture-independent molecular techniques, and the scope of data that can be obtained. In settings where culture is not always available, storage and subsequent remote analysis of these samples is an option to better understand the microbiology (although expense is still a barrier, and culture-based microbiology should not be completely neglected).

      Strengths:

      Ascribing aetiologies in respiratory infections is not (yet) an exact science, and there is likely to always be some doubt about whether an organism (whether identified by culture or nucleic acid detection) is a true pathogen. However, the authors have used a robust combination of histology, microbiology, radiology, clinical assessment and verbal autopsy, and this is probably as good as we can get it at present.

      The molecular techniques employed and the analysis were robust and technically sound. There are some areas where there is inconsistency between the results from two samples from the same patient, and this is likely due to the lower number of reads - important information for future similar studies. It highlights the potential limitations of this technique.

      The data obtained (albeit from a small sample set) are consistent with the other data from Africa, which provides support for the value and reliability of the technique itself.

      Weaknesses:

      The major weakness is the small sample size. The 7 patients analysed in this study are a subset of the children in the larger study who had been identified as having died from K. pneumoniae respiratory infection. So while the findings of this study highlight the potential role of this organism as a respiratory pathogen and provide some insight into the distribution of lineages in the community, the results don't really allow for major changes to clinical or diagnostic practice as yet. The data highlight a potentially under-recognised problem and would be useful to inform future studies.

      A minor weakness is more related to the journal layout, where methods are presented last. The results are not as easy to follow without reviewing the methods - in particular the description of histopathological findings and results of PCR on the biopsies. Until one realises that 6 biopsies had been taken from each of the deceased children, and a subset of these biopsies used for the study, the results seemed confusing.

    3. Reviewer #2 (Public review):

      Summary:

      This manuscript applies a culture-independent hybridization-capture metagenomic sequencing approach to characterize Klebsiella pneumoniae detected in post-mortem lung tissue from fatal pediatric pneumonia cases in Lusaka, Zambia. The study addresses an important challenge in retrospective genomic investigations where cultured isolates are unavailable and demonstrates the potential of targeted sequencing to recover clinically relevant genomic information directly from archived tissue specimens. The authors report sequence types, capsular loci, antimicrobial resistance determinants, virulence-associated genes, and evidence of closely related isolates in two cases. The work is valuable as a proof-of-concept application of targeted sequencing in challenging post-mortem specimens and provides useful descriptive genomic data from a setting where such information remains limited. However, several epidemiological and public health interpretations extend beyond what can be supported by the available data. The study includes only seven successfully sequenced children from a single setting and was not designed to determine the source of acquisition, transmission pathways, or population-level distributions of antimicrobial resistance or capsular types. The manuscript would therefore be strengthened by more consistently framing the findings as a descriptive genomic investigation of K. pneumoniae detected in children who died outside hospital settings, rather than as evidence of community-acquired infection or broader epidemiological shifts.

      Strengths:

      The principal strength of the manuscript is its methodological contribution. The authors demonstrate that hybridization-capture metagenomic sequencing can recover informative genomic data from post-mortem lung tissue in cases where conventional culture-based sequencing is not available. This is an important technical advance for retrospective studies, minimally invasive tissue sampling platforms, and settings where sample degradation, prior antibiotic exposure, or lack of routine culture limits genomic surveillance.

      The study also addresses an important public health problem. K. pneumoniae is a major cause of severe infection and antimicrobial resistance globally, yet its role in fatal pediatric pneumonia outside hospital settings remains difficult to define. The generation of sequence type, capsular locus, antimicrobial resistance, and virulence-associated gene data from post-mortem specimens is therefore useful and may inform future study designs. The identification of closely related isolates in two infants is also potentially important and raises hypotheses about shared sources or transmission that could be explored in larger studies.

      Another strength is that the authors appropriately acknowledge several technical challenges, including low numbers of K. pneumoniae-assigned reads in some specimens and unresolved or discordant capsular locus calls. These issues are important for readers considering the utility of this approach in low-input or mixed-specimen contexts.

      Weaknesses:

      The main weakness is that the epidemiological framing is stronger than the data allow. The manuscript repeatedly refers to community-acquired K. pneumoniae pneumonia and broader community epidemiology. However, the available data do not establish community acquisition, community transmission, or an epidemiological shift from nosocomial to community disease. Several children appear to have had prior healthcare contact or other potential healthcare-associated exposures, and the study design cannot determine where acquisition occurred. The findings would be more accurately framed as K. pneumoniae detected in post-mortem lung tissue from children who died outside hospital settings.

      Causal attribution also requires more careful wording. Detection of K. pneumoniae in post-mortem lung tissue, together with histopathology and DeCoDe findings, provides important supportive evidence that the organism may have been in the causal chain leading to death. However, this does not necessarily establish that K. pneumoniae was the sole or direct cause of fatal pneumonia, particularly where multiple pathogens were detected.

      The small sample size and case selection strategy limit the generalizability of the findings. Only seven children were successfully sequenced, and specimens appear to have been selected partly based on molecular signal. This is technically understandable, but it may introduce selection bias by enriching for cases with higher bacterial burden, better DNA preservation, or other specimen characteristics. As a result, the observed lineage diversity, resistance gene profiles, virulence-associated loci, and capsular locus distribution should not be interpreted as representative of community-acquired infections or broader population epidemiology.

      The validation of the hybridization-capture approach also requires strengthening. Comparing outputs from different genomic analysis tools applied to the same sequencing data may assess bioinformatic concordance, but it does not independently validate the method. Ideally, the approach should be benchmarked against clinical K. pneumoniae isolates or matched specimens with conventional whole-genome sequencing data. Without this, it is difficult to assess the accuracy of sequence type, capsular locus, antimicrobial resistance determinant, virulence locus, and plasmid marker recovery, especially in low-read or mixed-specimen contexts.

      Species-level attribution of antimicrobial resistance, virulence-associated genes, and plasmid replicons is another important limitation. In a culture-independent metagenomic study, these features cannot automatically be assigned to the identified K. pneumoniae lineage because many such elements are shared across Enterobacterales and may originate from co-detected organisms. This affects interpretation of antimicrobial resistance, hypervirulence, and MDR-hypervirulence convergence.

      Overall, the authors achieved their methodological aim of demonstrating that targeted sequencing can recover useful genomic information from challenging post-mortem specimens. However, the epidemiological, transmission, antimicrobial resistance, and vaccine-related conclusions should be tempered.

    1. eLife Assessment

      Using the clownfish model, this study examines how growth, feeding, and agonistic behavior result in socially dominant or subordinate states in size- and age-matched individuals of the clownfish, Amphiprion percula. The authors complement this work with whole-body transcriptomics and find significant variation in genes and gene co-expression modules related to growth and satiety-related pathways, as well as ossification-related genes. They provide solid evidence that emerging dominants grow more, eat more, and behave more aggressively than subordinate or solitary individuals; these phenotypic differences are accompanied by distinct gene expression profiles, including variation in growth- and satiety-related pathways. The work is valuable in advancing our understanding of how the social environment regulates phenotypic change; however, claims regarding the mechanistic role of gene expression are only partially supported by the current analyses.

    2. Reviewer #1 (Public review):

      Summary:

      Overall, this is an interesting and well-written manuscript on a fascinating question in a "charismatic" model system.

      Strengths:

      (1) The Introduction is concise, though it might be helpful to the non-specialist reader to learn a bit more about what is known about the social control of somatic growth across diverse species (including humans), which would help to make this work more generally interesting.

      (2) The experiment is well designed.

      (3) The data collected are comprehensive.

      (4) The complementary analysis of both feeding and aggression/submission data with and without known social roles is a neat idea and compelling!

      Weaknesses:

      The authors have addressed my concerns quite well. They now discuss the HPA/stress axis in some detail and also examined the extent to which growth, food intake, agonistic behavior, and/or gene expression patterns are coordinated across P1 vs P2 pairs. While still not ideal, they now provide a reasonable rationale for using whole bodies for the transcriptome analysis. Finally, the Discussion has been streamlined and is easy to follow.

    3. Reviewer #2 (Public review):

      In this manuscript, the authors test growth, behavior, and gene expression in pairs of clownfish as they establish social dominance hierarchies, examining patterns of gene expression in these pairs after dominance has been established. The authors show solid evidence that emerging dominant clownfish show increased growth, aggression, and food consumption compared to their submissive or solitary counterparts, eventually adopting distinct gene expression profiles.

      Major Comments:

      (1) The Introduction is comprehensive, but it could be condensed. Likewise, the discussion could be condensed. There is considerable redundancy between the methods, the results, and the legend in Figure 1. The authors should consolidate and remove the redundancy.

      (2) For Figure 3, the authors are showing PC2 and PC3; why is PC1 not shown? There is so much overlap between the three groups in PC2 vs PC3; it seems unlikely that researchers could conclusively identify any individual as belonging to a group based on the expression profile. The ovals shown do not capture all the points within each of the groups, and particularly the grey S oval seems misaligned with the datapoints shown.

      (3) The authors indicate that the 15 replicates exhibiting the greatest size difference between P1 and P2 were selected for gene profiling. Does this mean that each of the P1 and P2 were pairs with each other? Have the authors tried examining the gene expression patterns in a paired manner? E.g., for the pairs that showed the greatest size differences, do they also show the greatest differences in gene expression? Do the P1s show the most extreme differences from P2s that also show the most extreme P2 differences? Perhaps lines on Figure 3A connecting datapoints from the P1 and P2 pairs would be informative.

      (4) For the specific target pathways that are up- and downregulated in the different backgrounds, I recommend that the authors include boxplots (or heatmaps) showing the actual expression values for these targets. Figure 6 shows a heatmap for appetite-related genes, and it would be great to see a similar graph for the metabolism and glycolysis genes; it would also be informative to see similar graphs for hormonal and sexual maturation pathways as well.

      (5) Particularly given that there is a relatively small number of genes enriched in the different rank conditions, I did not understand the need to do the WGCNA module analysis. I thought that an analysis of GO terms across the dataset would have been more meaningful than the GO term analysis shown in Figure 4, which considers only genes assigned to the "brown WGCNA module". This should be simplified or clarified.

      (6) The authors say that they have identified coordinated changes in behaviors and the "underlying gene expression, leading to the emergence" of social roles. This is a little bit misleading, since the gene expression analysis occurred well after the behavioral and phenotypic differences emerged. Presumably, the hormonal and genetic shifts that actually caused the behavioral and phenotypic difference occurred during the weeks during which the experiment was underway, and earlier capture of the transcriptome would presumably reveal different patterns, and ones that would be considered more causative. The authors acknowledge this in 434-435, but it could be emphasized further.

      (7) The authors have measured a number of differences between the different dominance classes of fish. All these differences were measured relative to the other classes, but in my view, the Solitary group was the closest to a baseline control. So, I'm not sure that it is fair to say that "P2 and S individuals showed consistent downregulation of these genes and pathways" (line 401). I encourage the authors to emphasize the differences in gene expression from the "perspective" of the P1 individuals compared to the baseline of P2 and S individuals. Line 474 says that "P2 fish showed significant upregulation" of a number of pathways. It should be very clear what that is compared to (compared to P1, presumably?)

      (8) Along the same lines, the authors say in line 514 that subordinates and solitaries strategically downregulate their growth. I'm not convinced that this is the case: I would consider this growth trajectory to be the default and the baseline. I would interpret that under certain social conditions, a P1 dominant pattern of growth, behavior, and gene expression is allowed to emerge.

      Comments on revised version:

      The manuscript has been carefully revised. The authors have also responded adequately to all of my previous comments.

    4. Reviewer #3 (Public review):

      Summary:

      The authors tested the hypothesis that interactions among size- and age-matched rivals will lead to the emergence of social roles, accompanied by divergence in four aspects of individual phenotypes: growth, feeding behavior, fighting behaviors, and gene expression in clownfish.

      Strengths:

      The data of growth, feeding rate, and fighting behaviors support the authors claim.

      Weaknesses:

      The results obtained solely from the whole-body transcriptome are limited in supporting the authors' research question. However, the revised manuscript explicitly states this as a limitation.

    5. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      Overall, this is an interesting and well-written manuscript on a fascinating question in a "charismatic" model system.

      Strengths:

      (1) The Introduction is concise, though it might be helpful to the non-specialist reader to learn a bit more about what is known about the social control of somatic growth across diverse species (including humans), which would help to make this work more generally interesting.

      (2) The experiment is well-designed.

      (3) The data collected are comprehensive.

      (4) The complementary analysis of both feeding and aggression/submission data with and without known social roles is a neat idea and compelling!

      Thank you for the positive feedback!

      Here, we investigate phenotypic plasticity associated with the adoption of social roles in the clown anemonefish, with strategic growth being just one aspect of that plasticity. Strategic growth, also known as social control of growth, is a fascinating form of adaptive phenotypic plasticity, whereby individuals modify their growth and size in response to fine-scale changes in social conditions (Buston & Clutton-Brock, 2022). In cooperative breeding systems with high reproductive skew, particularly fishes and mammals (possibly including humans), individuals have been shown to i) increase growth/size on the acquisition of dominant status (Dengler-Crish & Catania, 2007; Johnston et al., 2021; Thorley et al., 2018; Van Schaik & Van Hooff, 1996; Walker & McCormick, 2009), ii) increase growth/size when paired with size matched reproductive rivals (Huchard et al., 2016; Reed et al., 2019; this study), and iii) decrease growth/size to avoid conflict (Buston, 2003; Heg et al., 2004; Wong et al., 2007). While strategic growth is fascinating and clearly occurring in this study, we show coordinated changes of multiple aspects of the phenotype as fish adopt social roles. Therefore, we deliberately framed the Introduction broadly to avoid biasing the reader toward viewing growth as the sole or main driver.

      Weaknesses:

      (1) I was surprised that the HPA/stress axis was not considered here at all. Wouldn't we expect that subordinates have increased stress axis activation, which in turn could inhibit their growth and aggressive behavior?

      We also expected to see the HPA/stress axis activated in subordinates, which is why we carried out a targeted exploration of genes known to play a role in this axis. We did not find any genes that were significantly differentially expressed. We believe that there could be two explanations for this. First, from a methodological perspective, it could be due to our use of a whole-body RNA-seq, which may have masked this signal. Alternatively, the stress axis might play a more complex role than just acting as a simple on/off switch for reduced growth. Its activation may peak when competition over size is at its highest (during week one) or, conversely, it may peak later and help maintain reduced growth once hierarchies are firmly established (particularly after the dominant individual reaches its maximum size). To understand the role of the stress axis, future studies should observe how its activation varies over time. We acknowledge that the absence of a stress‑axis signal and its potential explanations were not clearly discussed in the original manuscript. In the revised version, we have addressed this in the Discussion at lines 564-567 and Methods at lines 795-797 and 802-805.

      Discussion lines 564-567 and 577-580:

      “These include appetite regulation (orexigenic and anorexigenic signaling), metabolic pathways (e.g., glycolysis, lactic fermentation, TCA cycle, fatty acid β-oxidation), growth-regulating pathways (GH/IGF, insulin/PI3K-AKT, mTOR, and Hippo), and potential molecular signatures to varying levels of social stress.”

      Methods lines 795-797 and 802-805:

      “To further investigate GE differences, we explored genes associated with growth (including thyroid signaling), appetite regulation, metabolism, and stress (corticoid) pathways across social positions.”

      “A complete list of retrieved A. percula gene IDs were then filtered against the whole-body GE dataset, and pathways containing significant genes associated with social position were reported (Supplementary Table S2; appetite and metabolic genes shown).”

      (2) To what extent are growth, food intake, agonistic behavior, and/or gene expression patterns coordinated across P1 vs P2 pairs? The lack of such an analysis seems like a missed opportunity.

      We had a similar thought. Specifically, we were interested in testing the hypothesis that the final size ratio of pairs, which is indicative of the amount of conflict remaining, would predict gene expression. We examined gene expression within pairs to test for coordinated changes and repeated the analysis, accounting for the pair size ratio. In both cases, we found no clear or consistent pattern within pairs. We have included these analyses in the revised manuscript, Methods lines 781-7946-804 and 807-814, and Supplementary Materials lines 144-155 and 164-172 along with three new Supplementary Figures: Fig. S8, S9, S10.

      Supplementary Materials lines 144-155 and 164-172:

      “We have identified genes associated with growth and ossification that showed strong positive correlations with body size. For this set of genes, we tested the hypothesis that the final size ratio of pairs (P2 SL / P1 SL), which is indicative of the amount of remaining conflict (Wong et al., 2007), would predict variation observed in gene expression within social position (Fig. 5B). Visual inspection of the size ratio annotation in the heatmap of these genes (Supplementary Fig. S8), together with a PCA of paired individuals (P1 and P2) based on the same row Z-score values, was used to assess whether remaining conflict explained expression variation (PERMANOVA: p = 0.858; Supplementary Fig. S10A). These results suggest that size ratio between pairs was not associated with gene expression variation within social position.”

      “We have identified genes associated with appetite regulation and metabolism, which were significantly downregulated in P1 individuals compared to P2 and S fish. As above, we tested the hypothesis that the final size ratio of pairs (P2 SL / P1 SL), which is indicative of the amount of remaining conflict (Wong et al. 2016), would predict variation observed in gene expression within social position (Fig. 6). We found no significant effect (PERMANOVA: p = 0.844; Supplementary Fig. S9 and S10B), indicating that size ratio does not explain gene expression differences within social positions.”

      Methods lines 781-794 and 807-814:

      “In the heatmap, samples (columns) were a priori grouped by social position (P1, P2, S) and subsequently ordered within each group by gene expression similarity, reflected by the column dendrogram.

      Additionally, we tested the hypothesis that size ratio of pairs (P2 SL / P1 SL), which is indicative of the amount of conflict remaining (Wong et al., 2016), would predict variation in gene expression. This was assessed by visual inspection of the same heatmap with size ratio annotation using Complex Heatmaps (Gu et al., 2016), together with a PCA of the expression of these genes of paired individuals (P1, P2) only. We also performed a PERMANOVA analysis using the same row Z-score values (adonis2 function in vegan package: Oksanen et al., 2017) accounting for social position and genetic background (clutch ID), using restricted permutations to account for the non-independence of P1 and P2 within pairs was carried out. Solitary individuals were excluded from this analysis because they lack size ratio data. We found that size ratio between pairs was not associated with gene expression variation within social position (see Supplementary Materials; Supplementary Fig. S8, S10A).”

      “In the heatmap, samples (columns) and genes (rows) are grouped a priori by social position (P1, P2, S) and pathway (APT, TCA, GLY), respectively, with dendrograms reflecting expression similarity within each group. For the candidate gene heatmap, we repeated the PCA and PERMANOVA exploration and associated heatmap visualization for pairs only (as described above), to test whether final size ratio (P2 SL / P1 SL) predicted gene expression variation within social positions. We found that size ratio between pairs was not associated with gene expression variation within social position (see Supplementary Materials; Supplementary Fig. S9, S10B).”

      (3) What was the rationale for using whole bodies for the transcriptome analysis? Given the hypotheses, the forebrain or hypothalamus and certain other organ systems (e.g.,liver, gonads, skin, etc.) would have been obvious candidate tissues here. I realize that cost is always a consideration, but maybe a focus on the fore-/midbrain could have been prioritized.

      We decided to use whole-body samples for this initial transcriptomic analysis to capture a broad view of gene-expression differences while keeping sequencing costs and sample requirements manageable. We agree with the reviewer that future work should explore specific tissues sampled from individuals at multiple time points to disentangle transcriptomic differences across tissue types. In our revised manuscript we explicitly state the limitations and outline future steps in the Discussion at lines 553-564, as well as in the Methods section we now explicitly state the rationale behind whole-body RNA-seq and acknowledge the limitations of this approach in lines 688-692.

      Discussion line 5536-564:

      “In this study, we used whole-body transcriptomics, which revealed overarching expression patterns, however, we acknowledge that this approach limited our ability to detect finer-scale signals. Obviously, this approach cannot resolve tissue-specific gene expression changes (Lu et al., 2020; Roux et al., 2023; Yin et al., 2023) which are critical for a complete understanding social role differentiation, such as adjustments of behavior, appetite, and growth (a list we consider non-exhaustive). To disentangle gene expression signatures associated with socially induced phenotypes such as the strategic up- and downregulation of growth, future studies should include sampling points aligned with the onset of these physiological shifts. Moreover, to understand underlying pathways involved, tissue-specific transcriptomics will be essential. Sampling multiple tissues across multiple time points would disentangle nuanced regulatory processes underlying coordinated functional shifts leading to different social phenotypes of strategic growth.”

      Methods lines 688-692:

      “To initially explore overall gene expression patterns associated with coordinated changes during the emergence of social roles, whole-body RNA-seq was performed to keep both sequencing cost and sample requirements manageable. We acknowledge the limitations of this approach and consider this study a hypothesis-generating tool that lays the foundation for future studies.”

      (4) Given the preceding point, why was a fold-change threshold used for assessing DEGs (supplementary Figure 3)? There is no biological justification to ever use a fold-change threshold, especially in bulk RNA-seq analysis. This is particularly true here, where wholebodies were used for RNA-seq analysis, which is a bit unusual. Relatively small cell populations (such as hypothalamic neurons that regulate growth or food intake) may show substantial gene expression variation across social types, yet will be masked by the masses of other cells in the whole body sample. However, gene expression may still vary significantly, albeit the fold-difference may be small. I therefore suggest a reanalysis that omits any fold-change threshold.

      We thank the reviewer for this important point, and agree that an arbitrary fold‑change cutoff is inappropriate/unnecessary. It should be noted that this fold-change cutoff was only used in this single figure, and all other analyses used p-values from the entire dataset. We have removed the fold‑change threshold cutoff and corrected the Figure (previously Supplementary Figure 3) now Supplementary Fig. S5, and all corresponding text.

      (5) Why is the analysis of color (hue, saturation) buried in the supplementary materials? Based on the hypotheses that motivated the study, color seems just as relevant as food intake, growth, and agonistic behavior, so even if the results are negative, they should be presented in the main paper.

      We agree that color can be an important social signal, so we included color measurements in our experimental design. However, after careful consideration of the color results, we decided that our experimental timing and husbandry changes introduced multiple confounding factors, preventing us from drawing confident conclusions. Specifically, our fish were ≈1 month old at the transfer from larval to experimental tanks and had already begun to deepen their orange hue, before our experiment. (In the wild, they would settle at one to two weeks of age, prior to the deepening of the orange hue). Once individuals attain a certain hue, it seems that color development can be halted, but not reversed. The transfer also involved changes in lighting, tank background, and diet, factors known to strongly affect coloration (Maytin et. al., 2018). Our results show a uniform shift in orange hue and saturation across social groups, suggesting that these confounding factors might have dominated changes in hue.

      For transparency, we report the color data in the Supplementary Materials, but we caution against drawing any strong conclusions. In the revised Supplementary Materials document, we added lines 231-236 recommending that future work should involve a targeted experiment to robustly test for the effect of the adoption of social roles on coloration or the effect of coloration on the adoption of social roles.

      Supplementary Materials lines 231-236:

      “Together, these considerations suggest that timing, environmental uniformity, and future reproductive potential may limit the expression or detection of socially mediated color plasticity under laboratory conditions. Future studies should carefully consider experimental design to minimize confounding factors and robustly test the effects of the adoption of social roles on coloration and the effects of coloration on the adoption of social roles.”

      (6) The Discussion is sometimes difficult to follow. The authors may want to consider including a conceptual graphic that integrates the different aspects of growth and satiety regulation, etc., into a work-in-progress model of sorts, which would also facilitate clearer hypotheses for future research.

      Thank you for flagging that parts of the Discussion are a bit difficult to follow. In the revised manuscript, we worked to improve readability of the Discussion. We also appreciate the suggestion of including a conceptual schematic. For this manuscript, we refrained from adding such a “work-in-progress” schematic, as we felt that our whole-body RNA-seq approach and single gene expression sampling time point substantially limited our ability to make predictions.

      Reviewer #2 (Public review):

      In this manuscript, the authors test growth, behavior, and gene expression in pairs of clownfish as they establish social dominance hierarchies, examining patterns of gene expression in these pairs after dominance has been established. The authors show solid evidence that emerging dominant clownfish show increased growth, aggression, and food consumption compared to their submissive or solitary counterparts, eventually adopting distinct gene expression profiles.

      Major Comments:

      (1) The Introduction is comprehensive, but it could be condensed. Likewise, the discussion could be condensed. There is considerable redundancy between the methods, the results,and the legend in Figure 1. The authors should consolidate and remove the redundancy.

      Thank you for flagging that parts of the manuscript could be condensed, we will work on this as we revise the manuscript.

      (2) For Figure 3, the authors are showing PC2 and PC3; why is PC1 not shown? There is so much overlap between the three groups in PC2 vs PC3; it seems unlikely that researchers could conclusively identify any individual as belonging to a group based on the expression profile. The ovals shown do not capture all the points within each of the groups, and particularly the grey S oval seems misaligned with the datapoints shown.

      We understand the concern raised by the reviewer about the overlap among points in the PCA. We have explored PC1-PC3 and found that PC2 and PC3 showed the clearest, statistically significant clustering by social position, while PC1 did not capture any variation due to social position. We have explored whether other factors might be masking differences, such as genetic relatedness, tank effects, total read count per sample, and found that none of these factors explained sample clustering. Regarding the ellipses shown around the points, they were not intended to capture all points, but rather they show the estimated 95% multivariate t-distribution for that given social group. We revised the figure legend (Fig. 3) to clearly reflect this. . In addition, for transparency we have revised the Results section (lines 279-285) and Methods section (lines 754-765) to clarify that we have performed PCAs on (1) all genes, (2) top 50%, (3) top 25%, (4) top 5% most variable genes, and all three pairwise PC comparisons (PC1 and PC2, and PC1 and PC3) which we show in the added Supplementary Fig. 3S.

      Results lines 279-285:

      “PCAs were performed on all genes, and the top 50%, top 10%, and top 5% of the most variable genes, all of which showed consistent clustering across all pairwise PC comparisons (PC1 vs PC2; PC2 vs PC3; PC1 vs PC3; Supplementary Fig. S3). Of all the examined PCAs, PC2 and PC3 with all genes showed the clearest clustering by social position (p = 0.003), and revealed overall no significant difference in gene expression between P2 and S individuals, with P1 individuals exhibiting more distinct clustering (Fig. 3A; replicate‑annotated version of Fig. 3A in Supplementary Fig. S4; for all other PCAs see Supplementary Fig. S3).”

      Methods lines 754-765:

      “To test for overall GE patterns between social positions and genotypes, data were rlog-normalized and the effect of genotype and social position were compared through a Principal Component Analysis (PCA), followed by PERMANOVA analysis using Euclidean distances in the vegan package (Oksanen et al., 2017). To assess whether the strength of clustering by social position differed depending on the genes included, additional PCAs were performed on four gene subsets: (1) all genes, (2) the top 50%, (3) the top 10%, and (4) the top 5% of the most variable genes. For each subset, the first three principal components (PC1 vs PC2, PC1 vs PC3, PC2 vs PC3) were visualized to identify axes capturing variation associated with social position. To test for overall gene expression differences between social position and genotype, separate PERMANOVAs were performed for each PCA using the adonis2 in the vegan package (Oksanen et al., 2017).”

      (3) The authors indicate that the 15 replicates exhibiting the greatest size difference between P1 and P2 were selected for gene profiling. Does this mean that each of the P1and P2 were pairs with each other? Have the authors tried examining the gene expression patterns in a paired manner? E.g., for the pairs that showed the greatest size differences,do they also show the greatest differences in gene expression? Do the P1s show the most extreme differences from P2s that also show the most extreme P2 differences? Perhaps lines on Figure 3A connecting datapoints from the P1 and P2 pairs would be informative.

      Yes, “15 replicates exhibiting the greatest size difference between P1 and P2 were selected for gene profiling” refers to pairs of P1 and P2, we made sure this is clearly stated in the revised Results (lines 266-269) and Methods (lines 686-688). Yes, we have explored gene expression data considering the size difference between pairs, and found that it showed no clear differences in gene expression patterns (see our response and manuscript edits above under Reviewer #1 point 2). We also added a version of the main text Fig. 3A which clearly labels replicates within the PCA as Supplementary Fig. S4.

      Results lines 266-269:

      “To test the prediction that whole-body gene expression patterns will vary with social roles once clear social positions have emerged (P1, P2 and S), we performed gene expression profiling on 45 individual fish (experimental groups containing paired individuals and corresponding solitaries; N = 15 samples per social position; Fig. 1D).”

      Methods line 686-688:

      “Fish from these 15 replicates (n=45 samples; 15 paired individuals and their corresponding solitaries) were individually homogenized to allow equal RNA extraction from all tissues.”

      (4) For the specific target pathways that are up- and downregulated in the different backgrounds, I recommend that the authors include boxplots (or heatmaps) showing the actual expression values for these targets. Figure 6 shows a heatmap for appetite-related genes, and it would be great to see a similar graph for the metabolism and glycolytic genes; it would also be informative to see similar graphs for hormonal and sexual maturation pathways as well.

      We have explored genes across a broad set of metabolic pathways (glycolysis, TCA cycle, lactic fermentation, PDH complex, cholesterol biosynthesis, fatty-acid synthesis, and beta-oxidation) and show all metabolic genes that showed significant differential expression between P1, P2, and S in Figure 6. Overall, very few metabolism-associated genes were significantly differentially expressed, which is why we decided to combine appetite-regulation and metabolism-associated genes into a single figure (Figure 6). In our revised version of the manuscript, we have modified Fig. 6 to clearly indicate which pathways the genes belong to and modified Methods section lines 795-797 and 802-805.

      Methods lines 795-797 and 802-805:

      “To further investigate GE differences, we explored genes associated with growth (including thyroid signaling), appetite regulation, metabolism, and stress (corticoid) pathways across social positions.”

      “A complete list of retrieved A. percula gene IDs were then filtered against the whole-body GE dataset, and pathways containing significant genes associated with social position were reported (Supplementary Table S2; appetite and metabolic genes shown).”

      We also examined hormonal pathways (glucocorticoid and thyroid signaling), but did not find genes in these pathways that were significantly differentially expressed. Finally, we would like to clarify that our samples consist of two-month-old juvenile individuals that are sexually immature —under ideal conditions, clown anemonefish can mature in one to two years, but they can also remain sexually immature for a decade or more (Buston 2004; Buston & García, 2007) — which is why we did not observe distinct molecular signatures of sexual maturation. We recognize that the sentence at line 520 was misleading, as we did not identify any gene expression signature that we could confidently associate with signs of sexual maturation. We have revised the sentence in the Discussion (used to be line 520, now line 543-544) as well as in the Methods section we added lines 601-606.

      Discussion line 543-544:

      “In our current study, individuals within pairs were ultimately progressing towards rank 1 dominant female and rank 2 subordinate male roles.”

      Methods lines 601-606:

      “All fish used in this experiment were sexually immature juveniles (one-month-old at the beginning, two-month-old at the end), as clown anemonefish reach sexual maturity between the ages of one and two years old. Also, under ideal conditions, individuals can remain sexually immature for a decade or more (Buston 2004b; Buston & García, 2007). Therefore, in this study, the emergence of social roles and associated phenotypes do not reflect changes associated with sexual maturation.”

      (5) Particularly given that there is a relatively small number of genes enriched in the different rank conditions, I did not understand the need to do the WGCNA module analysis. I thought that an analysis of GO terms across the dataset would have been more meaningful than the GO term analysis shown in Figure 4, which considers only genes assigned to the "brown WGCNA module". This should be simplified or clarified.

      To clarify, GO enrichment analysis does not establish correlations with traits, it only describes which functions or pathways are over-represented in a given gene set. That is why we began by using WGCNA to define gene sets (modules) that are correlated to phenotypes. Our primary rationale for WGCNA was to identify modules of co-expressed genes that show significant statistical correlation with the phenotypes of interest (social role: P1, P2, S; growth; and food intake). Pairwise differential expression analysis (Figure 3B) identified a few hundred significantly differentially expressed genes, but those tests treat genes independently and are not able to help us link coordinated changes of co-expressed genes to phenotypes of interest. Because WGCNA is blind to traits, it first identifies groups of co-expressed genes, which can help resolve gene expression patterns.

      We therefore ran WGCNA on the rlog-transformed dataset to identify modules of co-expressed genes that show significant correlation with phenotypes of interest. For every module that showed such a correlation, we performed GO enrichment and carefully evaluated the resulting GO enrichment trees (see Supplementary Figs. S6, S7). The brown module was highlighted in the main text because it was one of the modules with a significant correlation to growth, and its associated GO enrichment showed clear growth-related signals that were not identified in the pairwise differential expression analysis results.

      In our revised manuscript we have clarified the rationale for the analytical approaches in the Results section at lines 297-303 and Methods section lines 767-772 and 775-779.

      Results lines 297-303:

      “...a weighted gene co-expression network analysis (WGCNA; Langfelder & Horvath, 2008) was conducted on the rlog-transformed gene expression (GE) dataset. Unlike pairwise DEG analysis, which treats genes independently, WGCNA identifies groups of co-expressed genes (modules) and tests whether modules of eigengene expression are correlated with variation in phenotypes of interest. For modules showing significant correlations with phenotypes, gene ontology (GO) analyses were performed using Fisher`s exact tests to identify over-represented pathways within each module.”

      Methods lines 767-772 and 775-779:

      “To test whether gene expression patterns are correlated with observed phenotypic changes in growth and appetite across social positions, a Weighted Gene Correlation Network Analysis (WGCNA) was conducted on the rlog-transformed GE dataset (Langfelder & Horvath, 2008). WGCNA identifies groups of co-expressed genes (modules) and correlates eigengene expression of each module with phenotypes of interest to determine modules of genes whose expression correlates with trait variation.”

      “For modules whose eigengene expression correlated with phenotypes of interest, gene ontology (GO) enrichment analysis was subsequently performed using Fisher's exact tests (presence/absence in a module) to identify over-represented pathways within each module across GO divisions of Biological Processes (BP), Molecular Functions (MF), and Cellular Components (CC) (Wright et al., 2015).”

      (6) The authors say that they have identified coordinated changes in behaviors and the"underlying gene expression, leading to the emergence" of social roles. This is a little bit misleading, since the gene expression analysis occurred well after the behavioral and phenotypic differences emerged. Presumably, the hormonal and genetic shifts that actually caused the behavioral and phenotypic difference occurred during the weeks during which the experiment was underway, and earlier capture of the transcriptome would presumably reveal different patterns, and ones that would be considered more causative.The authors acknowledge this in 434-435, but it could be emphasized further.

      We appreciate the reviewer raising this point. In the updated version of the manuscript, we have revised wording to convey that food intake, agonistic behavior, size and growth, and gene expression are all changing continuously, in response to each other and in response to social feedback (Introduction lines 133-136). An underappreciated aspect of this system (and likely many other systems) is that phenotype (including transcriptome) influences the outcome of social interactions, and the outcome of social interactions influences the phenotype (including the transcriptome). Earlier capture of the transcriptome would reveal different levels of gene expression, reflecting the state of the system at that moment in time.

      Introduction lines 134-137:

      “Underlying this cascade is an underappreciated aspect of this system (and likely many other systems) where phenotype (including gene expression) influences the outcome of social interactions, and the outcome of social interactions influences the phenotype (including gene expression).”

      (7) The authors have measured a number of differences between the different dominance classes of fish. All these differences were measured relative to the other classes, but in my view, the Solitary group was the closest to a baseline control. So I'm not sure that it is fair to say that "P2 and S individuals showed consistent downregulation of these genes and pathways" (line 401). I encourage the authors to emphasize the differences in gene expression from the "perspective" of the P1 individuals compared to the baseline of P2and S individuals. Line 474 says that "P2 fish showed significant upregulation" of a number of pathways. It should be very clear what that is compared to (compared to P1, presumably?)

      We agree with the reviewer that solitary individuals are the most intuitive baseline. Indeed, the experimental design included solitary fish because we expected they would serve as a useful control. Without social restraint, we anticipated they would show unrestricted growth, feeding, behavior, and associated gene‑expression patterns, similar to dominants.

      We initially ran analyses using solitaries as the baseline, but after examining the results, which showed subordinate‑like characteristics for the solitary individuals, we concluded that solitary individuals are not an ecologically appropriate control for this context. Removing juveniles from a social context and housing them in isolation may be stressful and can affect physiology and behavior in ways that do not reflect a natural baseline. From a life‑history standpoint, solitary living is not the typical state for A. percula.

      For these reasons, we reanalysed the dataset using the dominant (P1) as the reference to enable more ecologically meaningful comparisons (this choice was somewhat arbitrary, subordinates could also have been used as the reference). Given that gene expression is relative, we interpret results from both the dominant (P1) and subordinate (P2) perspectives in the Discussion to provide a complete view. We have clarified wording throughout the manuscript to make it clear that everything is relative (Introduction lines: 159-162 and 167-169; Methods lines: 622-625 and 726-728) as well as revised language throughout to make sure comparisons are clear.

      Introduction lines 159-162 and 167-169:

      “This experimental design allowed individuals equal opportunity to attempt taking on the dominant social role, however, through continuous social interactions (or in the case of solitaries, due to lack of social interactions), individuals took on varying social roles as dominant and subordinate members.”

      “Given that all phenotypes are entirely dependent on social context with no intrinsic baseline in this species, we arbitrarily chose P1 (dominant) as the statistical reference group for all downstream analyses.”

      Methods lines 622-625 and 726-728:

      “This experimental design allowed individuals equal opportunity to attempt taking on the dominant social role, however, through continuous social interactions (or in the case of solitaries, due to the lack of it), individuals took on varying social roles as dominant and subordinate members.”

      “Given that all phenotypes are entirely social context-dependent with no intrinsic baseline, in this species, we arbitrarily chose P1 (dominant) as the statistical reference group for all downstream analyses.”

      (8) Along the same lines, the authors say in line 514 that subordinates and solitaries strategically downregulate their growth. I'm not convinced that this is the case: I would consider this growth trajectory to be the default and the baseline. I would interpret that under certain social conditions, a P1 dominant pattern of growth, behavior, and gene expression is allowed to emerge.

      We respectfully disagree with the idea that a single baseline/reference growth trajectory exists for any individual of this species. Growth of individuals is entirely social context-dependent: neither fast nor slow growth represents an inherent baseline. When two size‑matched juveniles meet and compete to establish dominance, accelerated growth is the expected trajectory. By contrast, juveniles joining an existing hierarchy are expected to exhibit reduced growth, which minimizes conflict and facilitates their social integration. Unlike species that show non socially mediated growth trajectories, clown anemonefish do not have a context‑independent growth rate, rather, individuals constantly readjust their growth according to their immediate social environment (Buston 2003).

      Therefore, growth trajectories must be considered from the perspective of all group members, because they emerge from interactions among individuals rather than reflecting an intrinsic baseline. In this study, we were interested in the establishment of dominance hierarchy and how individuals adjust their phenotypes during this process. By experimentally pairing size‑matched rivals, both individuals are initially expected to pursue the dominant trajectory, and thus neither individual represents a default state. Instead, the outcome reflects a social decision, after which both individuals reinforce their emerging social roles through coordinated changes. See our response and manuscript edits above under Reviewer #2, point 7 (public reviews).

      Reviewer #3 (Public review):

      Summary:

      The authors tested the hypothesis that interactions among size- and age-matched rivals will lead to the emergence of social roles, accompanied by divergence in four aspects of individual phenotypes: growth, feeding behavior, fighting behaviors, and gene expression in clownfish.

      Strengths:

      The data on growth, feeding rate, and fighting behaviors support the authors' claims.

      Thank you for the positive feedback!

      Weaknesses:

      Gene analysis conducted in this study is not sufficient to clarify how the relevant genes actually regulate growth and behavior.

      The information obtained from whole-body gene expression analysis is very limited.Various gene expression is associated with the regulation of fighting behaviors, food intake, growth, and metabolism, and these genes are regulated differently across tissues, even within a single individual. Gene expression analysis should be performed separately for each tissue.

      We understand the reviewer’s concern about whole‑body transcriptomes and agree that tissue‑specific sampling would provide greater resolution of the mechanisms linking gene expression to growth, agonistic behaviors, and food intake. For this initial study, however, we deliberately chose whole‑body samples to capture a broad, unbiased view of gene expression differences while keeping sequencing costs and sample requirements manageable. We explicitly acknowledge the resulting interpretational limits in the Discussion (lines 497; 553–5647), and suggest in the last paragraph that the patterns reported here should be used to build on in future studies exploring targeted, tissue‑specific hypotheses. See our response and manuscript edits above under Reviewer #1, point 3 (public reviews).

      Clownfish undergo sex change depending on social status and body size, as the authors mention in the manuscript. Numerous gene expressions are affected by sex change. It is unclear how this issue was addressed.

      We thank the reviewer for raising this point. Sex change and sexual maturation can indeed drive major transcriptional shifts in clown anemonefish, but our experiment did not encompass such a life‑history transition. All individuals in this experiment were juveniles (≈1 month old at the start, ≈2 months old at the end) and were sexually immature at these ages. Clown anemonefish reach sexual maturation around one to two years under ideal conditions, can delay sexual maturation for years under normal conditions (Buston 2004; Buston & García, 2007), and sex change in the genus Amphiprion is known to take over ~5 months (Moyer & Nakazono, 1978). Accordingly, individuals in this study were not sexually mature, and sex change was not biologically plausible over the five-week experimental period of our study. We recognize that the sentence at line 520 may be misleading, as we did not identify any gene expression signature that we could confidently associate with signs of sexual maturation. During the revisions we made sure that it is clearly stated that the fish in this study were sexually immature. See our response and manuscript edits above under Reviewer #2, point 4 (public reviews).

      Recommendations for the authors:

      Reviewing Editor Comments:

      While appreciating the work presented in the manuscript, we note a few common concerns that the authors could prioritise in their revisions:

      (1) The transcriptomics

      The authors have used whole-body RNA-seq and so are constrained in the kind of mechanistic or tissue-specific insight it can really provide. There are also questions about how far the authors can go in linking these data to causation, given that sampling happened after the phenotypes had already diverged.

      We therefore recommend tempering the causal language, clarifying the rationale for their analytical choices (WGCNA, fold-change threshold...), and more explicitly discussing the limitations of the whole-body RNA-seq and what can (and cannot) be concluded from it.

      We thank the editors for these recommendations. We have addressed each point as follows:

      Tempering causal language: We have revised our manuscript and tempered with language throughout to remove wording of “influence” or “cause” and instead used “associated with”, “correlated with”, or “suggesting”, as appropriate. We made changes to our Abstract (lines 39-41), Introduction (lines 115-118, 121-123), and Discussion (lines 431-434, 571-574), and Methods sections (lines 795-797). Whereas other parts of the manuscript already used non-causal language such as Discussion lines 333, 355, 358, 381, 437, etc.

      Abstract lines 39-41:

      “Here we identify associations between changes in gene expression, growth, and feeding behavior regulation that reinforce social role differentiation during dominance hierarchy formation in clownfish.”

      Introduction lines 115-118, 121-123:

      “Gene expression profiling offers a powerful approach to uncover coordinated changes in gene expression across multiple pathways, providing insight into associations between the development of social role-specific phenotypes and their underlying proximate mechanisms.”

      “Its well-annotated genome (Lehmann et al., 2019) enables transcriptomic analyses that can help characterize the molecular patterns associated with these processes.”

      Discussion lines 431-434, 571-574:

      “To explore underlying GE differences, we examined genes known to be associated with appetite regulation and metabolism in A. ocellaris (Herrera et al., 2025), and found downregulation of these genes in P1 individuals compared to P2 and S.”

      “Together, these approaches provide a powerful framework for uncovering the dynamic, context-dependent mechanisms that regulate strategic growth and coordinated changes associated with the establishment and maintenance of dominance hierarchies in social vertebrates.”

      Methods lines 795-797:

      “To further investigate GE differences, we explored genes associated with growth (including thyroid signaling), appetite regulation, metabolism, and stress (corticoid) pathways across social positions.”

      WGCNA rationale: We have revised our Results and Methods sections to explicitly define why we used WGCNA and GO enrichment analysis, and how these analyses provided more information than simple pairwise differential gene expression analysis. See our response and manuscript edits above under Reviewer #2, point 5, (public reviews).

      Fold-change threshold: In our revised manuscript, we have removed the arbitrary imposed fold-change threshold cutoff, which was only used in that single figure. See our response and manuscript edits above under to Reviewer #1, point 4, (public reviews).

      Whole-body RNA-seq limitations: We have revised the Methods and Discussion sections, and see our response and manuscript edits above under Reviewer #1, point 3 (public reviews).

      (2) Framing and interpretation

      Several of the reviewers' comments flag that the manuscript overstates coordination or "strategic" regulation, or where what's being treated as the baseline/derived isn't clear (for example, whether P2/S are actively downregulating or whether P1 represents the divergent trajectory). We recommend revisiting any wording that implies stronger mechanistic inferences than the data actually support (and defining more clearly what is meant by baseline and socially induced state).

      We thank the editors for these comments and feedback. We have revised the manuscript to clearly state that our system does not have the traditional baseline/socially induced states, rather it is all social context-dependent. In particular, we have adjusted the wording throughout the manuscript to make it clear that everything is relative, and we clarified that our choice of P1 as the reference group was arbitrary. We have made revisions in the Introduction (lines 159-162 and 167-169) and Methods (lines 622-625 and 726-728). See our response and manuscript edits above under Reviewer #2, points 7 and 8, (public reviews).

      We have also revised the manuscript to temper with wording that could be interpreted as implying stronger mechanistic or directional conclusions than our data allow. We adjusted language throughout to remove wording of “influence” or “cause” and instead used “associated with”, “correlated with”, or “suggesting”, as appropriate. See our response and manuscript edits above under Reviewing Editor Comments, (1) the transcriptomics (recommendations to authors).

      Reviewer #1 (Recommendations for the authors):

      (1) The reader would benefit from a brief overview of the physiology and molecular basis of somatic growth, regulation of food intake, and aggressive behavior, especially in teleost fishes. This would also help the authors with formulating hypotheses that are a bit more explicit when it comes to the transcriptomic part of the study.

      We thank the reviewer for this suggestion, however, as another reviewer raised concerns about the length of the Introduction as is and we felt that adding detailed paragraphs on the physiology and molecular basis of somatic growth, food intake regulation, and aggressive behavior would significantly increase the length of our Introduction.

      As a compromise, we have revised the relevant sections of the Introduction (lines 109-115) to point readers to key review papers and primary literature where the molecular basis of these traits is covered in detail. We believe this approach balances the need for mechanistic context with manuscript length constraints, while also allowing readers with specific interests to follow up with the relevant literature.

      Introduction lines 109-115:

      “Most studies, often conducted without relevant social context of an individual or in species entirely lacking dominance hierarchies, have examined single pathways to uncover variation in coloration (Salis et al., 2019, 2021), appetite regulation (see review for teleosts: Volkoff, 2019), behavior (Bender et al., 2006; Renn et al., 2008; Santema et al., 2013; Solomon-Lane et al., 2022; see review for teleosts: St-Cyr & Aubin-Horth, 2009), and growth (Beckman, 2011; Lu et al., 2020; see reviews for teleosts: Reinecke, 2010; Zhou et al., 2024), leaving the broader molecular shifts associated with social role adoption unresolved.”

      (2) I certainly agree with that statement that "Understanding the proximate mechanisms that facilitate phenotypic adjustments is key to disentangling whether phenotypes are the cause or consequence of social rank" (line 90f.), though maybe the authors can elaborate a bit, as this relationship is quite dynamic and obviously goes both ways.

      We agree that this relationship is dynamic and bidirectional, and we have revised the Introduction to reflect this more explicitly. In particular, we added text noting that, in social vertebrates, phenotype, including gene expression, can both influence and be influenced by social interactions, often through continuous feedback loops rather than a clear directional relationship (Introduction lines 93-96). We also revised the experimental framing in the Introduction and Methods (Introduction lines 159-162 and 167-16972; Methods lines 622-625 and 726-728) to emphasize that phenotypes are socially context-dependent in A. percula and that the statistical reference group was chosen for analytical clarity rather than as a true biological baseline.

      Introduction lines 93-96, 159-162 and 167-169:

      “This is further complicated by the fact that, in social vertebrates, phenotype (including gene expression) can both influence and be influenced by social interactions, often through continuous feedback loops rather than a clear directional relationship.”

      “This experimental design allowed individuals equal opportunity to attempt taking on the dominant social role, however, through continuous social interactions (or in the case of solitaries, due to lack of social interactions), individuals took on varying social roles as dominant and subordinate members.

      “Given that all phenotypes are entirely dependent on social context with no intrinsic baseline in this species, we arbitrarily chose P1 (dominant) as the statistical reference group for all downstream analyses. Given that all phenotypes are entirely social context-dependent with no intrinsic baseline in this species, we arbitrarily chose P1 (dominant) as the statistical reference group for all downstream analysis.”

      Methods lines 622-625 and 726-728:

      “This experimental design allowed individuals equal opportunity to attempt taking on the dominant social role, however, through continuous social interactions (or in the case of solitaries, due to the lack of it), individuals took on varying social roles as dominant and subordinate members. This experimental design allowed individuals equal opportunity to attempt taking on the dominant social role; however, through continuous social interactions (or in the case of solitaries, due to the lack of it), individuals took on varying social roles as dominant and subordinate members.”

      “Given that all phenotypes are entirely social context-dependent with no intrinsic baseline, in this species, we arbitrarily chose P1 (dominant) as the statistical reference group for all downstream analyses. Given that all phenotypes are entirely social context-dependent with no intrinsic baseline in this species, we arbitrarily chose P1 (dominant) as the statistical reference group for all downstream analysis.”

      (3) Much of the information in the last paragraph of the Introduction (lines 151ff.) is best presented in the Methods section.

      We respectfully disagree with this suggestion. The eLife journal format does not include a standalone Methods section preceding the Results, so we expect most readers not to consult the Methods before reading the Results. We believe it is important to briefly orient the reader to the experimental approach and the ideas tested, thus providing some context before they encounter the Results and Discussion.

      (4) What count (TPM or similar) and abundance (above count threshold in fraction of samples) thresholds were used for the transcriptome analysis? Maybe I missed it, but how many genes were in the analysis?

      We thank the reviewer for pointing this out. We have updated our Methods section (lines 714-719) to explicitly include filtering steps used as well as included the total number of genes that were used in downstream analysis.

      Methods lines 714-719:

      “The read count file was then imported into R version 4.3.1 (R Core Team, 2021), size factors were estimated for each sample using the median ratio method in DESeq2 (Love et al., 2014) to account for differences in sequencing depth across samples. No outlier samples were identified, and all samples (n=45) were retained for downstream analyses. Genes with a mean raw count <10 were then removed, retaining 24,840 genes for downstream analyses.”

      (5) Were all these genes used in PCA, and if so, why? Would it not make more sense to only use the 50% or 25% most variable genes (which would likely enhance the separation of social types)? Also, did the authors inspect higher-order PCs to see whether any of them separate the social types or separate samples according to some other variable(e.g., size, hue, feeding, any technical factors, etc.)?

      We explored PCAs using four gene subsets: all genes, the top 50%, top 10%, and top 5% most variable genes, and all pairwise PC comparisons (PC1 vs PC2, PC1 vs PC3, PC2 vs PC3; Supplementary Fig. S3). Of all these examined PCAs, PC2 and PC3, with all genes, showed the clearest, statistically significant clustering by social position, and more stringent filtering did not strengthen this signal. We therefore retained all genes in the primary analysis to avoid imposing arbitrary filtering thresholds that could exclude biologically relevant low-variance genes. See our response and manuscript edits above under Reviewer #2, point 2, (public reviews).

      Regarding the inspection of higher-order PCs for potential confounding variables, we examined whether genetic background (clutch ID), tank identity, and sequencing depth explained clustering patterns across PC1–PC3 and found no evidence that any of these variables (PCA for sequencing depth not shown). See our response and manuscript edits above under Reviewer #2, point 2, (public reviews).

      We did not explicitly explore whether body size at week 5 drove any of the observed separation among social positions. Rather, we investigated whether size ratio, an indicator of the amount of remaining conflict within a pair, could drive gene expression variation within social positions. See our response and manuscript edits above under Reviewer #1, point 2, (public review).

      Regarding orange hue, we do not believe that orange hue at week 5 (the time point at which whole-body samples were collected for gene expression profiling) represents a meaningful socially mediated result. See our response and manuscript edits above under Reviewer #1, point 5, (public reviews).

      Considering food intake, we chose not to explore this as a potential driver of PC variation because food intake could not be reliably assigned to all individuals at week 5. For a subset of pairs, individuals could not be confidently distinguished in videos, and we did not want to make assumptions that could introduce biases into the analysis.

      (6) Figures 5/6: Based on the gene expression shown in the heatmaps, I cannot see how the samples would cluster so cleanly by social type. Consider a bootstrapping analysis and provide bootstrap values and "confident" nodes. Also, what does "matrix" in the legend refer to? I assume some measure of gene expression level, maybe z-scored?

      We thank the reviewer for flagging this point. We acknowledge that the samples do not cluster freely by social position in the heatmaps and we have revised our Methods (see our response and manuscript edits above under Reviewer #1, point 2, public reviews) and Figure captions to clearly reflect this. To clarify, samples in Figures 5 and 6 were not ordered by unsupervised hierarchical clustering; rather, they were first grouped a priori by social position and then ordered within each group by similarity. We chose this presentation because it best illustrates the gene expression patterns associated with social position, which is the primary focus of the manuscript. We have updated the figure legends of Figures 5 and 6, to clarify that the color scale previously denoted as matrix in the heatmap represents row-scaled Z-scores of normalized gene expression values.

      To assess whether social position explains a gene expression variation in an unsupervised framework, we performed a PCA followed by PERMANOVA (using the adonis2 function in R, 999 permutations, Euclidean distance; see Author response image 1). Both analyses used the same row Z-score-scaled expression values as shown in Figures 5 and 6. Results showed a significant effect of social position on gene expression (PERMANOVA: A: social position p <0.001, B: social position p < 0.001; marginal clutch ID significance p=0.015), which was primarily driven by dominant individuals (P1) being significantly different from both subordinate (P2) and solitary (S) individuals. To include these into our manuscript.

      Author response image 1.

      Principal component analysis (PCA) of gene expression based on heatmaps of A) growth, B) appetite, and metabolism genes. PCA was performed on gene sets of the main text (growth: Fig 5B and appetite and metabolism: Fig. 6), using row Z-score scaled expression values, consistent with the heatmap scaling. Each point represents one individual. Ellipses represent 95% confidence intervals around each social position group. PERMANOVA results (adonis2; shown in the bottom right corners).

      (7) I applaud the authors for considering genetic/relatedness effects in their experimental design and analysis, but I am confused by the microsatellite vs. SNP analyses: why even use microsatellites in this day and age? Were only the SNP data used as the source of genetic information? This should be clarified.

      In our experiment, paired individuals were initially assigned a temporary rank based on their size at the start of the experiment; however, the initially larger individual (often only by a tenth of a mm) does not necessarily emerge as the dominant (P1). To avoid any assumptions regarding the identity of individuals in given social positions at the end of the experiment, we needed to verify that individuals within pairs were correctly identified. For the 45 individuals included in the gene expression dataset, this was done by calling SNPs directly from the TagSeq data. For the remaining 36 individuals not included in the gene expression dataset, we opted for microsatellite genotyping, as a validated panel with established markers was already available from a previous study (Rueger et al., 2025), making it a reliable and cost-effective solution. We have clarified the text in the Methods (lines 630-636) and Supplementary Materials (lines 42-49 and 65-67).

      Methods lines 630-636:

      “This step was necessary to avoid assumptions regarding social roles and fish identity. At the beginning of the experiment, individuals within pairs were provisionally assigned a social rank based on initial body size; as size differences were negligible (often less than 0.1mm), the initially larger individual does not necessarily emerge as the dominant (P1). To correct this assumption, identities were verified at the end of the experiment using either microsatellite genotyping or SNPs called from TagSeq data, depending on whether individuals were included in the gene expression dataset (see Supplementary Materials).”

      Supplementary Materials lines 42-49 and 65-67:

      “This step was necessary to avoid assumptions regarding social roles and fish identity, and the microsatellite method allowed for a reliable and cost-effective solution as a panel with established markers was already available from a previous study (Rueger et al., 2025). At the beginning of the experiment, individuals within pairs were provisionally assigned a social rank based on initial body size; as size differences were negligible (often less than 0.1 mm), the initially larger individual does not necessarily emerge as the dominant (P1). To correct this assumption, identities were verified at the end of the experiment.”

      “Similarly, for the remaining 45 individuals, we applied a necessary correction step to avoid assumptions regarding social role and fish identity. For these 45 individuals, we called SNPs from our TagSeq data to assign clutch identity.”

      Reviewer #2 (Recommendations for the authors):

      (1) Line 520 indicates that individuals showed early gene expression signatures of sexual maturation, but I did not see where those results were presented.

      We have addressed this recommendation, the claim “individuals showed early gene expression signatures of sexual maturation” was incorrect and has been removed from the revised manuscript (Discussion lines 543-544). We have also updated the Methods section (lines 601-606) to explicitly clarify that all individuals were sexually immature juveniles throughout the experiment. See our response and manuscript edits above under Reviewer #2, point 4, (public reviews).

      (2) The paragraph starting at line 409 refers to GE profiles. I was confused about what that was.

      Do the authors mean GO profiles?

      We did not find any modules using GO enrichment analysis which showed strong enrichment of appetite- and metabolism-related GO terms. Therefore, we looked for differentially expressed genes in the rlog-normalized gene expression dataset that are known to be associated with appetite regulation and metabolism. We have revised the Discussion (lines 431-434) and Methods (lines 812-815) to explicitly state what we are referring to and avoid confusion.

      Discussion lines 431-434:

      “In our study, P1 individuals showed increased food intake compared to P2 individuals. To explore underlying GE differences, we examined genes known to be associated with appetite regulation and metabolism in A. ocellaris (Herrera et al., 2025), and found downregulation of these genes in P1 individuals compared to P2 and S.”

      Method lines 802-805:

      “A complete list of retrieved A. percula gene IDs were then filtered against the whole-body GE dataset, and pathways containing significant genes associated with social position were reported (Supplementary Table S2; appetite and metabolic genes shown).”

      (3) I saw several typos, e.g., Vulcano in Supplementary Figure 3.

      Thank you for flagging this. We have addressed typos such as Supplementary Fig. S4 (used to be Supplementary Fig S3) See our response and manuscript edits above under Reviewer #1, point 4 (public reviews).

      (4) The personal observations cited in 495 should be more explicit. Which author made these observations, over how long, in how many instances?

      We thank the reviewer for this comment. We have updated the text to explicitly name the authors who made these observations, to clarify that they were made during two independent long-term field studies, and to note that this pattern was observed in 8 or more instances across the two field studies (Discussion lines 516-520).

      Discussion lines 516-520:

      “Similar patterns have been anecdotally observed in the wild, where juvenile clownfish remained small for extended periods (over four months) following the loss of a dominant partner, only initiating changes in social role towards dominant characteristics upon the arrival of a new group member (personal observations of 8+ instances during two long-term independent studies by Pete Buston and Lili Vizer).”

      References:

      Buston, P. (2003). Forcible eviction and prevention of recruitment in the clown anemonefish. Behavioral Ecology, 14(4), 576–582. https://doi.org/10.1093/beheco/arg036

      Buston, Peter M. (2004). Territory inheritance in clownfish. Proceedings of the Royal Society B: Biological Sciences, 271(SUPPL. 4), 252–254.

      Buston, P. M., & García, M. B. (2007). An extraordinary life span estimate for the clown anemonefish Amphiprion percula. Journal of Fish Biology, 70(6), 1710–1719. https://doi.org/10.1111/j.1095-8649.2007.01445.x

      Buston, P., & Clutton-Brock, Tim. (2022). Strategic growth in social vertebrates (WITH REVIEWER COMMENTS). Trends in Ecology & Evolution, 37(8), 694–705. https://doi.org/10.1016/j.tree.2022.03.010

      Dengler-Crish, C. M., & Catania, K. C. (2007). Phenotypic plasticity in female naked mole-rats after removal from reproductive suppression. THE JOURNAL OF EXPERIMENTAL BIOLOGY.

      Heg, D, Bender, N, & Hamilton, I. (2004). Strategic growth decisions in helper cichlids. Proceedings of the Royal Society of London. Series B: Biological Sciences, 271(suppl_6). https://doi.org/10.1098/rsbl.2004.0232

      Huchard, E, English, S, Bell, M B. V., Thavarajah, N, & Clutton-Brock, T. (2016). Competitive growth in a cooperative mammal. Nature, 533(7604), 532–534. https://doi.org/10.1038/nature17986

      Johnston, R A., Vullioud, P, Thorley, J, Kirveslahti, H., Shen, L., Mukherjee, S., Karner, C. M., Clutton-Brock, T, & Tung, J (2021). Morphological and genomic shifts in mole-rat ‘queens’ increase fecundity but reduce skeletal integrity. eLife, 10, e65760. https://doi.org/10.7554/eLife.65760

      Maytin, Alexander K., Davies, Sarah W., Smith, Gabriella E., Mullen, Sean P., & Buston, Peter M. (2018). De novo transcriptome assembly of the clown anemonefish (Amphiprion percula): A new resource to study the evolution of fish color. Frontiers in Marine Science, 5(AUG), 1–11.

      Moyer, J. T., & Nakazono, A. (1978). Protandrous Hermaphroditism in Six Species of the Anemonefish Genus Amphiprion in Japan (No. 2). The Ichthyological Society of Japan. https://doi.org/10.11369/jji1950.25.101

      Reed, C., Branconi, R., Majoris, J., Johnson, C., & Buston, P. (2019). Competitive growth in a social fish. Biology Letters, 15(2), 20180737. https://doi.org/10.1098/rsbl.2018.0737

      Rueger, Theresa, Bhardwaj, Anjali Kristina, Turner, Emily, Barbasch, Tina Adria, Trumble, Isabela, Dent, Brianne, & Buston, Peter Michael. (2022). Vertebrate growth plasticity in response to variation in a mutualistic interaction. Scientific Reports, 12(1), 11238.

      Thorley, J, Katlein, N, Goddard, K, Zöttl, M, & Clutton-Brock, T. (2018). Reproduction triggers adaptive increases in body size in female mole-rats. Proceedings of the Royal Society B: Biological Sciences, 285(1880), 20180897. https://doi.org/10.1098/rspb.2018.0897

      Van Schaik, C P., & Van Hooff, J A. R. A. M. (1996). Toward an understanding of the orangutan’s social system. In Linda F. Marchant, Toshisada Nishida, & William C. McGrew (Eds.), Great Ape Societies (pp. 3–15). Cambridge University Press. https://doi.org/10.1017/CBO9780511752414.003

      Walker, S P. W., & McCormick, M I. (2009). Sexual selection explains sex-specific growth plasticity and positive allometry for sexual size dimorphism in a reef fish. Proceedings of the Royal Society B: Biological Sciences, 276(1671), 3335–3343. https://doi.org/10.1098/rspb.2009.0767

      Wong, M. Y. L., Buston, P. M., Munday, Philip L., & Jones, Geoffrey P. (2007). The threat of punishment enforces peaceful cooperation and stabilizes queues in a coral-reef fish. Proceedings of the Royal Society B: Biological Sciences, 274(1613), 1093–1099. https://doi.org/10.1098/rspb.2006.0284

    1. eLife Assessment

      This study presents valuable evidence that activating a lysosomal signaling pathway lowers blood glucose in diabetic mice, pointing to a potential new strategy for treating metabolic disease. However, the evidence for the proposed mechanism is inadequate: the glucose transporter's movement and glucose uptake were not measured in the physiologically relevant tissue, and the measurement that was performed used an unconventional technique. As a result, whether this transporter-dependent pathway explains the glucose-lowering effect has not yet been established.

    2. Reviewer #1 (Public review):

      Summary:

      This article purports to show that ML-SA8, a synthetic activator of the lysosomal TRPML1 channel, results in AMPK activation and glucose uptake in hepatocytes, and that this action has therapeutic potential for metabolic disease. The final figure shows that glucose levels are improved in db/db mice, although it is not entirely clear whether this is due to an effect on the liver, on other tissues, or on glucose production or uptake. The earlier figures try to make the case that SA8 causes activation and GLUT4 translocation and glucose uptake in liver cells; however, these data are not convincing. GLUT4 is expressed at such low levels in liver that it is likely not physiologically important. The authors use a fluorescent glucose analog to measure glucose uptake, and this molecule has been shown to enter cells largely by fluid phase endocytosis. Overall, this reviewer finds the premise misguided and the data unconvincing.

      Strengths and Weaknesses:

      The initial figures show phosphorylation of AMPK on Thr172, but no downstream effects are shown. Usually, to convincingly show that AMPK activity is increased, it would be appropriate to immunoblot phospho-ACC or some other substrate. This is minor.

      Lines 135-148: GLUT4 is not expressed at levels that are significant for physiology in liver cells, and its function in liver is not particularly relevant. The authors cite references 38-40 to support that it may be expressed at low levels in liver, but no knockout studies have been done to show that this expression is physiologically important.

      Figure 1e is not convincing. No controls are included to show the specificity of the antibody for immunofluorescent staining. No intracellular GLUT4 is visible in the unstimulated samples.

      In Figure 1f, again, the data are not convincing. The bands seem too sharp for GLUT4, which has 12 membrane-spanning domains as well as an N-linked glycosylation, so that it usually runs as a smear.

      Figure 1h. Data are not convincing. 2-NBDG is not a valid approach to measure glucose uptake. 2-NBDG enters cells largely via fluid phase endocytosis, and its accumulation is independent of known GLUT inhibitors such as cytochalasin B (Yazdani et al., MBoC 2022; PMID: 35921166; see also PMID: 42287154). The idea that such a bulky derivative of glucose could enter the transporter channel is not compatible with known structural data.

      Supplementary Figure 5 uses 2-NBDG glucose uptake again. This reviewer is not convinced that the data reflect transporter-mediated glucose uptake, as suggested by the authors. As well, although palmitate treatment of cells can cause an insulin-resistant-like phenotype in some cell types, this is not characterized in the present work. Finally, as noted, one would not expect hepatocytes to exhibit insulin-responsive glucose transport. Glycogen synthesis is the main insulin-regulated step that might be affected.

      The data in Figures 2b,c,f,g,k,l are not convincing. Again, 2-NBDG is used.

      For the glucose consumption measurements in other panels of Figure 2, the methods section states that cells were cultured in 10 mM glucose. What volume was used? It is difficult to believe that a monolayer of cells would consume very much of the glucose that is present in the culture medium. Data are shown as a percent of controls, and look reasonable, but it would be helpful to include absolute as well as relative units.

      In Figure 2, in experiments using the TRPML1 KO cells, no panel is shown to demonstrate knockout. The authors cite a previous paper for the construction of these cells, but the control immunoblot should still be shown here.

      In Figure 3, controls are missing in the BAPTA experiment in Figure 3a (only SA8-treated cells were treated with BAPTA and with EGTA). Again, it would be helpful to have p-ACC or some other readout of AMPK activity, and not just AMPK phosphorylation. 2NBDG is again used in this figure.

      Line 212-213 the text states "considering our finding that TRPML1-mediated Ca2+ release is essential for AMPK activation." This has not been shown. The work uses chelators and does not necessarily indicate a role for TRPML1. The drug may be specific, as suggested by the authors, but the way this phrase is worded is too strong. As well, AMPK was shown to be phosphorylated, but full activation towards its various substrates has not been shown.

      Figure 4cd suggests that GLUT4 expression is increased by 2 or 3-fold in the liver of DB+SA8-treated mice, compared to controls. This may be the case, but its abundance is still likely ~1000-fold less in liver compared to skeletal muscle or adipose tissue. This reviewer is still not convinced that this is physiologically relevant. The images in Supplementary Figure 8 suggest a larger increase, but it remains uncertain whether the staining really represents GLUT4.

      Data showing that blood glucose and HbA1c are reduced in SA8-treated mice are reasonable, and GTTs and ITTs are shown. Unfortunately, there are no insulin concentrations, and it remains uncertain whether glucose production is reduced or uptake is increased (or if both effects are present).

      In the discussion, the authors again state that GLUT4 is present in the liver and that it regulates hepatic glucose homeostasis, and they cite reference 63. This review article does not argue that GLUT4 acts in the liver to regulate hepatic glucose homeostasis, but that its actions in muscle and fat have secondary effects on the liver.

    3. Reviewer #2 (Public review):

      Summary:

      The manuscript contains interesting studies suggesting that pharmacological activation of TRPML1 could be useful to treat T2D by increasing glucose uptake via activation of AMPK. Preclinical studies suggest the inhibitor improved blood glucose in Db/Db mice. Ex vivo studies in cell lines examine both pharmacologic and genetic manipulations, both to activate and to inactivate TRPML1, and the results consistently suggest that TRPML1 activates AMPK and increases glucose uptake.

      Strengths:

      The manuscript is well written, and the studies are carefully performed.

      Weaknesses:

      All mechanistic studies were performed in transformed cell lines; conclusions would be stronger if performed in primary cells. The in vivo studies were only performed in male mice. Performing metabolic studies in both sexes is standard practice now. Whether the findings would extend to females was not tested and remains uncertain. Some controls are missing, such as plasma membrane loading controls for fractionation studies. The GLUT4 staining was performed after fixation and permeabilization, yet control cells appear to be devoid of intracellular (and all) staining, a confusing result that doesn't reflect the expected biology.

    4. Reviewer #3 (Public review):

      Summary:

      Zhu et al. present a proof-of-concept for targeting the lysosomal calcium channel MCOLN1/TRPML1endolysosomal ion channels to restore type 2 diabetes mellitus (T2DM). Using synthetic TRPML1 agonists (ML-SA8) and genetic manipulation, the authors demonstrate that TRPML1 stimulation triggers localized lysosomal calcium release. This calcium efflux sequentially activates CaMKKβ and phosphorylates AMPK at Thr172 in various cell models, including palmitic acid-induced insulin-resistant HepG2 cells. This signaling pathway promotes GLUT4 translocation to the plasma membrane and increases intracellular glucose uptake. When administered daily to diabetic db/db mice over six weeks, ML-SA8 lowers fasting and random blood glucose, improves oral glucose and insulin tolerance tests, reduces hepatic steatosis, and lowers serum ALT and AST levels.

      Strengths:

      Based on the TFEB-independent pathway activated by TRPML1 and the experimental approaches described by Medina's group (PMID: 31822666), the authors use a combination of pharmacological and genetic tools to dissect such an intracellular signaling pathway. Additionally, the animal experiments show consistent phenotypic improvements across independent metabolic parameters. The ability of ML-SA8 to restore glycogen deposition and clear hepatic lipid accumulation in db/db mice without causing weight loss or overt toxicity provides a strong rationale for exploring lysosomal targets in metabolic disease.

      Weaknesses:

      (1) The authors focus almost exclusively on hepatic GLUT4 to explain the observed glucose disposal. However, other glucose transporter isoforms such as GLUT2 dominate basal glucose transport. While the authors show increased AMPK phosphorylation in skeletal muscle and adipose tissue, they do not measure GLUT4 translocation or glucose uptake in these primary disposal organs. As a result, attributing systemic glycemic recovery primarily to hepatic GLUT4 translocation overlooks the major physiological roles of peripheral tissues.

      (2) In both HepG2 cells and mouse liver tissues, ML-SA8 treatment increases total GLUT4 protein expression in addition to plasma membrane localization. Because total protein pools expand, the enrichment of GLUT4 in plasma membrane fractions cannot be cleanly attributed to acute vesicular translocation alone. The manuscript does not explain the timescale or mechanism behind this rapid total protein upregulation, leaving a mechanistic gap between acute ion channel gating and protein expression.

      (3) While the in vitro specificity of ML-SA8 is well-controlled, the systemic animal experiments lack a specific rescue or knockout control. Small-molecule agonists administered intraperitoneally over six weeks can exert off-target effects. Without demonstrating that co-administering the TRPML1 inhibitor ML-SI5 blunts the therapeutic effect in vivo, or showing that ML-SA8 lacks efficacy in TRPML1-null mice, the definitive link between in vivo glycemic recovery and TRPML1 activation remains incomplete.

    5. Author response:

      Public Reviews:

      Reviewer #1 (Public review):

      (1) This article purports to show that ML-SA8, a synthetic activator of the lysosomal TRPML1 channel, results in AMPK activation and glucose uptake in hepatocytes, and that this action has therapeutic potential for metabolic disease. The final figure shows that glucose levels are improved in db/db mice, although it is not entirely clear whether this is due to an effect on the liver, on other tissues, or on glucose production or uptake. The earlier figures try to make the case that SA8 causes activation and GLUT4 translocation and glucose uptake in liver cells; however, these data are not convincing. GLUT4 is expressed at such low levels in liver that it is likely not physiologically important. The authors use a fluorescent glucose analog to measure glucose uptake, and this molecule has been shown to enter cells largely by fluid phase endocytosis. Overall, this reviewer finds the premise misguided and the data unconvincing.

      Thank you for the critical comments and constructive suggestions. We have carefully considered all the concerns raised and provide our point-by-point responses below.

      (2) The initial figures show phosphorylation of AMPK on Thr172, but no downstream effects are shown. Usually, to convincingly show that AMPK activity is increased, it would be appropriate to immunoblot phospho-ACC or some other substrate. This is minor.

      Suggestion was taken! We will investigate the effect of SA8 on AMPK downstream effectors i.e. ACC activation by Western blotting. Ie p-ACC (Ser79) / total ACC.

      (3) Lines 135-148: GLUT4 is not expressed at levels that are significant for physiology in liver cells, and its function in liver is not particularly relevant. The authors cite references 38-40 to support that it may be expressed at low levels in liver, but no knockout studies have been done to show that this expression is physiologically important.

      We thank the reviewer for this critical comment. We agree that GLUT4 is not the predominant hepatic glucose transporter, but it is expressed at a relatively low level in the liver compared to other tissues.

      Regarding the physiological role of GLUT4, it has been shown to mediate glucose uptake in hepatic stellate and sinusoidal endothelial cells (Tang and Chen 2010, Karim, Liaskou et al. 2014). Furthermore, ischemia‑reperfusion (IR) significantly upregulated the expression of GLUT4 in the liver, rather than GLUT2, and the increased GLUT4 localized to the membrane peripheries of hepatocytes and enhanced glucose uptake, which in turn led to marked glycogen deposition (Kim, Jung et al. 2014, Kurabayashi, Furihata et al. 2022). These studies establish a clear physiological role for GLUT4 in the liver.

      Notably, Ranalletta, et. al. (2005) showed that GLUT4-null mice exhibit compensatory alterations in hepatic glucose and lipid metabolism, including increased hepatic glucose uptake and triglycerides conversion (Ranalletta, Jiang et al. 2005), indicating that GLUT4 ablation influences liver metabolism. Nevertheless, these existing evidences including the GLUT4 expression data and the functional changes observed in GLUT4 null mice—supports the relevance of GLUT4 in hepatic glucose metabolism. However, it is necessary to perform the liver-specific GLUT4 KO studies to clarify the role of GLUT4 in the liver. (We will incorporate these points and limitations into the revised Discussion section).

      (4) Figure 1e is not convincing. No controls are included to show the specificity of the antibody for immunofluorescent staining. No intracellular GLUT4 is visible in the unstimulated samples.

      We understood the reviewer’s concern about the specificity of GLUT4 antibody for immunofluorescent staining. The GLUT4 antibody (Abcam, ab33780) employed in our study has been extensively validated in previous studies for both Western blotting (Xie, Liu et al. 2024, Amanollahi, Holman et al. 2025, Ando, Takeda et al. 2025) and immunofluorescence (see also Johansson, Mannerås-Holm et al. 2013, Xiao, Zhang et al. 2025). in addition, we confirmed its specificity in our system by Western blot (Suppl. Fig.4), which showed a single band at ~45–55 kDa. Collectively, the combination of published validations, our own data supports the specificity of GLUT4 immunofluorescence detection.

      About the intracellular GLUT4 signal in Fig 1e. We apologize for the unclear GLUT4 signal in our original Fig. 1e. This was due to an inadvertently short exposure for the control condition. To improve this, we have now acquired new images with uniformly increased exposure time for all groups. As shown in the new Fig. 1e, GLUT4 is now clearly detected in control cells and mainly in cytosol, and ML-SA8 treatment obviously increases its accumulation at the plasma membrane.

      (5) In Figure 1f, again, the data are not convincing. The bands seem too sharp for GLUT4, which has 12 membrane-spanning domains as well as an N-linked glycosylation, so that it usually runs as a smear.

      We understood the reviewer’s concern. As an N-glycosylated membrane protein, GLUT4 may exhibit broader or diffuse migration patterns on immunoblotting. Nevertheless, the final band pattern is influenced by multiple factors, e.g. antibody specificity, sample preparation, electrophoresis conditions, and detection condition. Of note, multiple independent studies have shown endogenous GLUT4 as a relatively “sharp” immunoreactive band around 55–60 kDa (Gurley, Ilkayeva et al. 2016, Habtemichael, Li et al. 2021, Wu, Yu et al. 2024) (see also Ando et al., 2025; Amanollahi et al., 2025; Xie et al., 2024), which is very consistent with our results. Thus, we are confident that our GLUT4 band is specific and reliable.

      (6) Figure 1h. Data are not convincing. 2-NBDG is not a valid approach to measure glucose uptake. 2-NBDG enters cells largely via fluid phase endocytosis, and its accumulation is independent of known GLUT inhibitors such as cytochalasin B (Yazdani et al., MBoC 2022; PMID: 35921166; see also PMID: 42287154). The idea that such a bulky derivative of glucose could enter the transporter channel is not compatible with known structural data.

      We thank the reviewer for raising this point. We acknowledge that the uptake mechanism of 2-NBDG is controversial. As the reviewer noted, some studies have reported that it enters cells largely by endocytosis in certain cell types, and this remains a subject of ongoing discussion. Nevertheless, 2-NBDG continues to be widely employed as a glucose uptake tracer in this field, including several recent high-profile studies (Nobs, Kolodziejczyk et al. 2023, Xiong, Helm et al. 2023, Wu, Lv et al. 2025) as following.

      (1) Nobs, S.P., et al., Lung dendritic-cell metabolism underlies susceptibility to viral infection in diabetes. Nature, 2023. 624(7992): p. 645-652.

      (2) Xiong, L., et al., Nutrition impact on ILC3 maintenance and function centers on a cell-intrinsic CD71-iron axis. Nat Immunol, 2023. 24(10): p. 1671-1684.

      (3) Wu, Y., et al., Dalbergia odorifera T.C. Chen leaf extract promotes microglial energy expenditure to phagocytize neutrophils after cerebral ischemia-reperfusion. Phytomedicine, 2025. 149: p. 157508.

      In the current study, given the consistency of our results with parallel functional assays, the use of 2‑NBDG is justified in this context.

      (7) Supplementary Figure 5 uses 2-NBDG glucose uptake again. This reviewer is not convinced that the data reflect transporter-mediated glucose uptake, as suggested by the authors.

      Please see response to #6.

      (8) As well, although palmitate treatment of cells can cause an insulin-resistant-like phenotype in some cell types, this is not characterized in the present work.

      We understand the reviewer’s concern regarding the characterization of PA-induced insulin resistance model.

      First, this PA-induced insulin resistant hepatic model is well-established and validated in several literatures (Lee, Cho et al. 2010, Zhang, Cai et al. 2020, Malik, Inamdar et al. 2024), and it has been used for T2DM natural and synthetic drug screening (Faria, Calixto et al. 2025).

      Second, we have characterized this model in our system. As shown in Suppl. Fig. 5a, b, insulin (100 nM, 0.5 h) induced an increase of glucose uptake in HepG2 cells measured by 2-NBDG (Yamada, Nakata et al. 2000). In contrast, in PA-treated HepG2 cells, this effect was almost completely blocked, suggesting that PA-treated HepG2 cells are less sensitive to insulin. Overall, this model has been well validated and is suitable for the purposes of our study.

      (9) Finally, as noted, one would not expect hepatocytes to exhibit insulin-responsive glucose transport. Glycogen synthesis is the main insulin-regulated step that might be affected.

      As stated in the Responses#9, PA-induced hepatic insulin-resistance has become a widely accepted in vitro model for investigating therapeutic strategies (Lee, Cho et al. 2010, Zhang, Cai et al. 2020, Malik, Inamdar et al. 2024, Faria, Calixto et al. 2025).

      (10) The data in Figures 2b,c,f,g,k,l are not convincing. Again, 2-NBDG is used.

      Please see response to #6.

      (11) For the glucose consumption measurements in other panels of Figure 2, the methods section states that cells were cultured in 10 mM glucose. What volume was used? It is difficult to believe that a monolayer of cells would consume very much of the glucose that is present in the culture medium. Data are shown as a percent of controls, and look reasonable, but it would be helpful to include absolute as well as relative units.

      We appreciate the reviewer’s comments and apologize for any confusion about the method.

      First, we have revised the methods section to clarify the assay procedure “Following the manufacturer’s protocol, 2.5 μL of sample (medium/standard) was mixed with 250 μL of working solution a 96-well plate. The mixture was then incubated at 37 °C for 10 min and the absorbance was measured…” (line 421-423).

      Second, in response to the suggestion to include both absolute and relative units, we will provide the data with absolute value for reviewer’s reference. In the main figures, we have retained the normalized data as this format allows direct comparison of treatment effects across independent experiments.

      (12) In Figure 2, in experiments using the TRPML1 KO cells, no panel is shown to demonstrate knockout. The authors cite a previous paper for the construction of these cells, but the control immunoblot should still be shown here.

      We thank the reviewer for raising this point. The TRPML1 knockout cell line used in our study was originally generated and provided by Prof. Haoxing Xu’s laboratory, this cell line has been validated in several published literature (Wang, Gao et al. 2015, Zhang, Cheng et al. 2016). We understand the reviewer’s concern, so we will further validate this TRPML1 KO cell line.

      (13) In Figure 3, controls are missing in the BAPTA experiment in Figure 3a (only SA8-treated cells were treated with BAPTA and with EGTA). Again, it would be helpful to have p-ACC or some other readout of AMPK activity, and not just AMPK phosphorylation. 2NBDG is again used in this figure.

      Suggestion taken! We will add the controls including BAPTA-AM and EGTA only data. p-ACC/total ACC will also be measured. About the 2-NBDG, please see Responses#6.

      (14) Line 212-213 the text states "considering our finding that TRPML1-mediated Ca2+ release is essential for AMPK activation." This has not been shown. The work uses chelators and does not necessarily indicate a role for TRPML1. The drug may be specific, as suggested by the authors, but the way this phrase is worded is too strong. As well, AMPK was shown to be phosphorylated, but full activation towards its various substrates has not been shown.

      We thank the reviewer for raising this concern. We fully agree that the Ca<sup>2+</sup> chelators experiment alone could not specially attribute the effect to TRPML1.

      In fact, we have performed experiment to address the TRPML1-dependent mechanism in the original submission. As shown in Fig. 2i-l, In TRPML1 KO HAP1 cells (Qi, Xing et al. 2021), ML-SA8-induced AMPK phosphorylation and cellular glucose uptake were almost completely abolished compared to wild-type (WT) HAP1 cells. Moreover, pharmacological inhibition of TRPML1 with a TRPML1 specific synthetic inhibitor-ML-SI5 completely abolished ML-SA8-triggered AMPK phosphorylation (Fig. 2d, e). In addition, in IR-HepG2 model, ML-SI5 could substantially inhibited ML-SA8-induced cellular glucose uptake (Fig. 2f-h).

      Accordingly, we have also revised the statement to” considering our finding that TRPML1-mediated Ca<sup>2+</sup> release is necessary for AMPK activation” to make the sentence more rigorous.

      (15) Figure 4cd suggests that GLUT4 expression is increased by 2 or 3-fold in the liver of DB+SA8-treated mice, compared to controls. This may be the case, but its abundance is still likely ~1000-fold less in liver compared to skeletal muscle or adipose tissue. This reviewer is still not convinced that this is physiologically relevant. The images in Supplementary Figure 8 suggest a larger increase, but it remains uncertain whether the staining really represents GLUT4.

      We understood the reviewer’s concern. About the physiological importance of GLUT4, please see Responses #3. About the specificity of GLUT4 antibody, please see Responses#4.

      (16) Data showing that blood glucose and HbA1c are reduced in SA8-treated mice are reasonable, and GTTs and ITTs are shown. Unfortunately, there are no insulin concentrations, and it remains uncertain whether glucose production is reduced or uptake is increased (or if both effects are present).

      We will measure the insulin concentrations.

      (17) In the discussion, the authors again state that GLUT4 is present in the liver and that it regulates hepatic glucose homeostasis, and they cite reference 63. This review article does not argue that GLUT4 acts in the liver to regulate hepatic glucose homeostasis, but that its actions in muscle and fat have secondary effects on the liver.

      Sorry for the oversights. We have supplementary more precise references (Rossetti, Stenbit et al. 1997, Kurabayashi, Furihata et al. 2022, Fan, Jiao et al. 2023, Jiang, Luo et al. 2024).

      Reviewer #2 (Public review):

      The manuscript contains interesting studies suggesting that pharmacological activation of TRPML1 could be useful to treat T2D by increasing glucose uptake via activation of AMPK. Preclinical studies suggest the inhibitor improved blood glucose in Db/Db mice. Ex vivo studies in cell lines examine both pharmacologic and genetic manipulations, both to activate and to inactivate TRPML1, and the results consistently suggest that TRPML1 activates AMPK and increases glucose uptake.

      Strengths:

      The manuscript is well written, and the studies are carefully performed.

      (18) All mechanistic studies were performed in transformed cell lines; conclusions would be stronger if performed in primary cells. The in vivo studies were only performed in male mice. Performing metabolic studies in both sexes is standard practice now. Whether the findings would extend to females was not tested and remains uncertain. Some controls are missing, such as plasma membrane loading controls for fractionation studies. The GLUT4 staining was performed after fixation and permeabilization, yet control cells appear to be devoid of intracellular (and all) staining, a confusing result that doesn't reflect the expected biology.

      Thank you for your support and the constructive suggestions! We will answer these questions in the following point-to-point responses.

      Reviewer #3 (Public review):

      (19) Zhu et al. present a proof-of-concept for targeting the lysosomal calcium channel MCOLN1/TRPML1endolysosomal ion channels to restore type 2 diabetes mellitus (T2DM). Using synthetic TRPML1 agonists (ML-SA8) and genetic manipulation, the authors demonstrate that TRPML1 stimulation triggers localized lysosomal calcium release. This calcium efflux sequentially activates CaMKKβ and phosphorylates AMPK at Thr172 in various cell models, including palmitic acid-induced insulin-resistant HepG2 cells. This signaling pathway promotes GLUT4 translocation to the plasma membrane and increases intracellular glucose uptake. When administered daily to diabetic db/db mice over six weeks, ML-SA8 lowers fasting and random blood glucose, improves oral glucose and insulin tolerance tests, reduces hepatic steatosis, and lowers serum ALT and AST levels.

      Strengths:

      Based on the TFEB-independent pathway activated by TRPML1 and the experimental approaches described by Medina's group (PMID: 31822666), the authors use a combination of pharmacological and genetic tools to dissect such an intracellular signaling pathway. Additionally, the animal experiments show consistent phenotypic improvements across independent metabolic parameters. The ability of ML-SA8 to restore glycogen deposition and clear hepatic lipid accumulation in db/db mice without causing weight loss or overt toxicity provides a strong rationale for exploring lysosomal targets in metabolic disease.

      Thank you for the support!

      (21) The authors focus almost exclusively on hepatic GLUT4 to explain the observed glucose disposal. However, other glucose transporter isoforms such as GLUT2 dominate basal glucose transport. While the authors show increased AMPK phosphorylation in skeletal muscle and adipose tissue, they do not measure GLUT4 translocation or glucose uptake in these primary disposal organs. As a result, attributing systemic glycemic recovery primarily to hepatic GLUT4 translocation overlooks the major physiological roles of peripheral tissues.

      We thank the reviewer for this critical comment and we agree with the reviewer. Our initial focus on hepatic GLUT4 was driven by our primary interest in liver metabolism and the fact that we observed a consistent and robust effect of ML‑SA8 on hepatic GLUT4 translocation, AMPK activation and glucose uptake. We acknowledge that the possible contribution of GLUT2 to these effects cannot be excluded. Also, ML‑SA8 may exert similar effects in other tissues, such as skeletal muscle and adipose tissue as we found out that GLUT4 levels were upregulated in these tissues (Suppl.Fig.8b, c). We therefore agree that the systemic glycemic recovery should be attributed to the multi‑tissue effects of ML‑SA8, rather than a liver-restricted phenomenon. Accordingly, we will revise the Discussion section to incorporate this. Note that the precise mechanisms of ML-SAs on GLUT2 and in peripheral tissues warrant further investigation.

      (22) In both HepG2 cells and mouse liver tissues, ML-SA8 treatment increases total GLUT4 protein expression in addition to plasma membrane localization. Because total protein pools expand, the enrichment of GLUT4 in plasma membrane fractions cannot be cleanly attributed to acute vesicular translocation alone. The manuscript does not explain the timescale or mechanism behind this rapid total protein upregulation, leaving a mechanistic gap between acute ion channel gating and protein expression.

      Great suggestion. we will re-calculated the enrichment of GLUT4 in PM with total GLUT4 protein to clarify.

      (23) While the in vitro specificity of ML-SA8 is well-controlled, the systemic animal experiments lack a specific rescue or knockout control. Small-molecule agonists administered intraperitoneally over six weeks can exert off-target effects. Without demonstrating that co-administering the TRPML1 inhibitor ML-SI5 blunts the therapeutic effect in vivo, or showing that ML-SA8 lacks efficacy in TRPML1-null mice, the definitive link between in vivo glycemic recovery and TRPML1 activation remains incomplete.

      We understand the reviewer’s concern about the specificity of ML-SAs for TRPML1. In fact, the specificities of ML-SAs and ML-SIs have been rigorously validated in previous studies using TRPML1 knockout (KO) cells (Sahoo, Gu et al. 2017, Yu, Zhang et al. 2020) and further confirmed in the atomic-resolution co-structures (Schmiege, Fine et al. 2017, Schmiege, Fine et al. 2021).

      In the current study we provide additional evidence supporting this specificity. We show that ML-SA8-induced AMPK phosphorylation and glucose uptake were abolished in TRPML1 knockout cells (Fig. 2i-l) and blocked by specific inhibitor ML-SI5 (Fig. 2d, e, f, g). In addition, ML‑SA5 has been show to lack efficacy in TRPML1‑null mice in several studies (Yu, Zhang et al. 2020, Zhang, Wang et al. 2024, Xing, Wang et al. 2025), further reinforcing the target specificity of this class of compounds. Hence, multiple lines of evidence support the on-target effect of ML-SA8: (1) in vitro blockade by ML-SI5 and TRPML1 KO; (2) consistent effects across structurally distinct TRPML1 agonists; (3) dose‑dependent responses in vivo; and (4) absence of efficacy in TRPML1‑null mice.

      We agree with the reviewer that rescue experiment or in vivo knockout validation would provide the most definitive proof of target specificity. However, breeding TRPML1‑null mice or performing extensive dose-finding studies for ML-SI5 in vivo would require substantial time and resources which probably fall beyond the scope of the current study. We have therefore acknowledge this limitation in the Discussion section and noted that future validation with ML-SI5 co-administration or TRPML1-null mice is warranted. We hope the reviewer finds our response acceptable.

      Amanollahi, R., S. L. Holman, A. S. Meakin, M. Padhee, K. J. Botting-Lawford, S. Zhang, S. M. MacLaughlin, D. O. Kleemann, S. K. Walker, J. M. Kelly, S. R. Rudiger, I. C. McMillen, M. D. Wiese, M. C. Lock and J. L. Morrison (2025). "In Vitro Embryo Culture Impacts Heart Mitochondria in Male Adolescent Sheep." J Dev Biol 13(2).

      Ando, T., R. Takeda, R. Kano, T. Kusano, Y. Nonaka, Y. Kano and D. Hoshino (2025). "Effects of pyruvate administration on mRNA expression of inflammatory cytokines in adipose tissue and whole-body glucose metabolism in male mice." Physiol Rep 13(15): e70362.

      Casimiro, I., N. D. Stull, S. A. Tersey and R. G. Mirmira (2021). "Phenotypic sexual dimorphism in response to dietary fat manipulation in C57BL/6J mice." J Diabetes Complications 35(2): 107795.

      Fan, X., G. Jiao, T. Pang, T. Wen, Z. He, J. Han, F. Zhang and W. Chen (2023). "Ameliorative effects of mangiferin derivative TPX on insulin resistance via PI3K/AKT and AMPK signaling pathways in human HepG2 and HL-7702 hepatocytes." Phytomedicine 114: 154740.

      Faria, B. Q., P. S. Calixto, G. Picheth, L. M. Ferreira, F. G. M. Rego, J. F. C. Guerra and M. H. M. Sari (2025). "Palmitate-induced hepatic insulin resistance as an in vitro model for natural and synthetic drug screening: A scoping review of therapeutic candidates and mechanisms." Chem Biol Interact 420: 111717.

      Gurley, J. M., O. Ilkayeva, R. M. Jackson, B. A. Griesel, P. White, S. Matsuzaki, R. Qaisar, H. Van Remmen, K. M. Humphries, C. B. Newgard and A. L. Olson (2016). "Enhanced GLUT4-Dependent Glucose Transport Relieves Nutrient Stress in Obese Mice Through Changes in Lipid and Amino Acid Metabolism." Diabetes 65(12): 3585-3597.

      Habtemichael, E. N., D. T. Li, J. P. Camporez, X. O. Westergaard, C. I. Sales, X. Liu, F. López-Giráldez, S. G. DeVries, H. Li, D. M. Ruiz, K. Y. Wang, B. S. Sayal, S. González Zapata, P. Dann, S. N. Brown, S. Hirabara, D. F. Vatner, L. Goedeke, W. Philbrick, G. I. Shulman and J. S. Bogan (2021). "Insulin-stimulated endoproteolytic TUG cleavage links energy expenditure with glucose uptake." Nat Metab 3(3): 378-393.

      Jiang, Y., P. Luo, Y. Cao, D. Peng, S. Huo, J. Guo, M. Wang, W. Shi, C. Zhang, S. Li, L. Lin and J. Lv (2024). "The role of STAT3/VAV3 in glucolipid metabolism during the development of HFD-induced MAFLD." Int J Biol Sci 20(6): 2027-2043.

      Johansson, J., L. Mannerås-Holm, R. Shao, A. Olsson, M. Lönn, H. Billig and E. Stener-Victorin (2013). "Electrical vs manual acupuncture stimulation in a rat model of polycystic ovary syndrome: different effects on muscle and fat tissue insulin signaling." PLoS One 8(1): e54357.

      Karim, S., E. Liaskou, J. Fear, A. Garg, G. Reynolds, L. Claridge, D. H. Adams, P. N. Newsome and P. F. Lalor (2014). "Dysregulated hepatic expression of glucose transporters in chronic disease: contribution of semicarbazide-sensitive amine oxidase to hepatic glucose uptake." Am J Physiol Gastrointest Liver Physiol 307(12): G1180-1190.

      Kim, S., J. Jung, H. Kim, R. W. Heo, C. O. Yi, J. E. Lee, B. T. Jeon, W. H. Kim, J. R. Hahm and G. S. Roh (2014). "Exendin-4 Improves Nonalcoholic Fatty Liver Disease by Regulating Glucose Transporter 4 Expression in ob/ob Mice." Korean J Physiol Pharmacol 18(4): 333-339.

      Kim, S. J., A. Gajbhiye, A. R. Lyu, T. H. Kim, S. A. Shin, H. C. Kwon, Y. H. Park and M. J. Park (2023). "Sex differences in hearing impairment due to diet-induced obesity in CBA/Ca mice." Biol Sex Differ 14(1): 10.

      Kurabayashi, A., K. Furihata, W. Iwashita, C. Tanaka, H. Fukuhara, K. Inoue, M. Furihata and Y. Kakinuma (2022). "Murine remote ischemic preconditioning upregulates preferentially hepatic glucose transporter-4 via its plasma membrane translocation, leading to accumulating glycogen in the liver." Life Sci 290: 120261.

      Lee, J. Y., H. K. Cho and Y. H. Kwon (2010). "Palmitate induces insulin resistance without significant intracellular triglyceride accumulation in HepG2 cells." Metabolism 59(7): 927-934.

      Luo, J., H. Alkhalidy, Z. Jia and D. Liu (2024). "Sulforaphane Ameliorates High-Fat-Diet-Induced Metabolic Abnormalities in Young and Middle-Aged Obese Male Mice." Foods 13(7).

      Malik, S., S. Inamdar, J. Acharya, P. Goel and S. Ghaskadbi (2024). "Characterization of palmitic acid toxicity induced insulin resistance in HepG2 cells." Toxicol In Vitro 97: 105802.

      Nguyen-Phuong, T., S. Seo, B. K. Cho, J. H. Lee, J. Jang and C. G. Park (2023). "Determination of progressive stages of type 2 diabetes in a 45% high-fat diet-fed C57BL/6J mouse model is achieved by utilizing both fasting blood glucose levels and a 2-hour oral glucose tolerance test." PLoS One 18(11): e0293888.

      Nobs, S. P., A. A. Kolodziejczyk, L. Adler, N. Horesh, C. Botscharnikow, E. Herzog, G. Mohapatra, S. Hejndorf, R. J. Hodgetts, I. Spivak, L. Schorr, L. Fluhr, D. Kviatcovsky, A. Zacharia, S. Njuki, D. Barasch, N. Stettner, M. Dori-Bachash, A. Harmelin, A. Brandis, T. Mehlman, A. Erez, Y. He, S. Ferrini, J. Puschhof, H. Shapiro, M. Kopf, A. Moussaieff, S. K. Abdeen and E. Elinav (2023). "Lung dendritic-cell metabolism underlies susceptibility to viral infection in diabetes." Nature 624(7992): 645-652.

      Qi, J., Y. Xing, Y. Liu, M. M. Wang, X. Wei, Z. Sui, L. Ding, Y. Zhang, C. Lu, Y. H. Fei, N. Liu, R. Chen, M. Wu, L. Wang, Z. Zhong, T. Wang, Y. Liu, Y. Wang, J. Liu, H. Xu, F. Guo and W. Wang (2021). "MCOLN1/TRPML1 finely controls oncogenic autophagy in cancer by mediating zinc influx." Autophagy 17(12): 4401-4422.

      Racine, K. C., L. Iglesias-Carres, J. A. Herring, K. L. Wieland, P. N. Ellsworth, J. S. Tessem, M. G. Ferruzzi, C. D. Kay and A. P. Neilson (2024). "The high-fat diet and low-dose streptozotocin type-2 diabetes model induces hyperinsulinemia and insulin resistance in male but not female C57BL/6J mice." Nutr Res 131: 135-146.

      Ranalletta, M., H. Jiang, J. Li, T. S. Tsao, A. E. Stenbit, M. Yokoyama, E. B. Katz and M. J. Charron (2005). "Altered hepatic and muscle substrate utilization provoked by GLUT4 ablation." Diabetes 54(4): 935-943.

      Rossetti, L., A. E. Stenbit, W. Chen, M. Hu, N. Barzilai, E. B. Katz and M. J. Charron (1997). "Peripheral but not hepatic insulin resistance in mice with one disrupted allele of the glucose transporter type 4 (GLUT4) gene." J Clin Invest 100(7): 1831-1839.

      Sahoo, N., M. Gu, X. Zhang, N. Raval, J. Yang, M. Bekier, R. Calvo, S. Patnaik, W. Wang, G. King, M. Samie, Q. Gao, S. Sahoo, S. Sundaresan, T. M. Keeley, Y. Wang, J. Marugan, M. Ferrer, L. C. Samuelson, J. L. Merchant and H. Xu (2017). "Gastric Acid Secretion from Parietal Cells Is Mediated by a Ca2+ Efflux Channel in the Tubulovesicle." Developmental Cell 41(3): 262-273.e266.

      Schmiege, P., M. Fine, G. Blobel and X. Li (2017). "Human TRPML1 channel structures in open and closed conformations." Nature 550(7676): 366-370.

      Schmiege, P., M. Fine and X. Li (2021). "Atomic insights into ML-SI3 mediated human TRPML1 inhibition." Structure 29(11): 1295-1302 e1293.

      Tang, Y. and A. Chen (2010). "Curcumin prevents leptin raising glucose levels in hepatic stellate cells by blocking translocation of glucose transporter-4 and increasing glucokinase." Br J Pharmacol 161(5): 1137-1149.

      Tukhovskaya, E. A., E. R. Shaykhutdinova, I. A. Pakhomova, G. A. Slashcheva, N. A. Goryacheva, E. S. Sadovnikova, E. A. Rasskazova, V. A. Kazakov, I. A. Dyachenko, A. A. Frolova, A. N. Brovkin, V. E. Kaluzhsky, M. Y. Beburov and A. N. Murashev (2022). "AICAR Improves Outcomes of Metabolic Syndrome and Type 2 Diabetes Induced by High-Fat Diet in C57Bl/6 Male Mice." Int J Mol Sci 23(24).

      Wang, W., Q. Gao, M. Yang, X. Zhang, L. Yu, M. Lawas, X. Li, M. Bryant-Genevier, N. T. Southall, J. Marugan, M. Ferrer and H. Xu (2015). "Up-regulation of lysosomal TRPML1 channels is essential for lysosomal adaptation to nutrient starvation." Proc Natl Acad Sci U S A 112(11): E1373-1381.

      Wu, D., H. C. Yu, H. N. Cha, S. Park, Y. Lee, S. J. Yoon, S. Y. Park, B. H. Park and E. J. Bae (2024). "PAK4 phosphorylates and inhibits AMPKα to control glucose uptake." Nat Commun 15(1): 6858.

      Wu, Y., W. Lv, S. Xiong, G. Cao, L. Fu, W. Liu, F. Shao, Y. Mei and Y. Lv (2025). "Dalbergia odorifera T.C. Chen leaf extract promotes microglial energy expenditure to phagocytize neutrophils after cerebral ischemia-reperfusion." Phytomedicine 149: 157508.

      Xiao, B., W. Zhang, N. Ji and Q. Chen (2025). "Knockdown of CCNB1 alleviates high glucose-triggered trophoblast dysfunction during gestational diabetes via Wnt/β-catenin signaling pathway." Open Med (Wars) 20(1): 20241119.

      Xie, Y., X. Liu, W. Liu, L. R. Carr, L. P. Lee, N. Imai, E. A. Ortlund and D. E. Cohen (2024). "Activity and phosphatidylcholine transfer protein interactions of skeletal muscle thioesterase Them2 enable hepatic steatosis and insulin resistance." J Biol Chem 300(11): 107855.

      Xing, Y., M. M. Wang, F. Zhang, T. Xin, X. Wang, R. Chen, Z. Sui, Y. Dong, D. Xu, X. Qian, Q. Lu, Q. Li, W. Cai, M. Hu, Y. Wang, J. L. Cao, D. Cui, J. Qi and W. Wang (2025). "Lysosomes finely control macrophage inflammatory function via regulating the release of lysosomal Fe(2+) through TRPML1 channel." Nat Commun 16(1): 985.

      Xiong, L., E. Y. Helm, J. W. Dean, N. Sun, F. R. Jimenez-Rondan and L. Zhou (2023). "Nutrition impact on ILC3 maintenance and function centers on a cell-intrinsic CD71-iron axis." Nat Immunol 24(10): 1671-1684.

      Yamada, K., M. Nakata, N. Horimoto, M. Saito, H. Matsuoka and N. Inagaki (2000). "Measurement of glucose uptake and intracellular calcium concentration in single, living pancreatic beta-cells." J Biol Chem 275(29): 22278-22283.

      Yu, L., X. Zhang, Y. Yang, D. Li, K. Tang, Z. Zhao, W. He, C. Wang, N. Sahoo, K. Converso-Baran, C. S. Davis, S. V. Brooks, A. Bigot, R. Calvo, N. J. Martinez, N. Southall, X. Hu, J. Marugan, M. Ferrer and H. Xu (2020). "Small-molecule activation of lysosomal TRP channels ameliorates Duchenne muscular dystrophy in mouse models." Sci Adv 6(6): eaaz2736.

      Zhang, G., X. Cai, L. He, D. Qin, H. Li and X. Fan (2020). "Skimmin Improves Insulin Resistance via Regulating the Metabolism of Glucose: In Vitro and In Vivo Models." Front Pharmacol 11: 540.

      Zhang, H., Y. Wang, R. Wang, X. Zhang and H. Chen (2024). "TRPML1 agonist ML-SA5 mitigates uranium-induced nephrotoxicity via promoting lysosomal exocytosis." Biomed Pharmacother 181: 117728.

      Zhang, X., X. Cheng, L. Yu, J. Yang, R. Calvo, S. Patnaik, X. Hu, Q. Gao, M. Yang, M. Lawas, M. Delling, J. Marugan, M. Ferrer and H. Xu (2016). "MCOLN1 is a ROS sensor in lysosomes that regulates autophagy." Nat Commun 7: 12109.

    1. eLife Assessment

      In their important study, Beaudet, Berger and Hendricks provide a mechanistic link between disease-associated tau hyperphosphorylation, loss of cooperative tau envelope formation on microtubules, and dysregulation of axonal transport prior to aggregation. Using complementary in vitro reconstitution and human iPSC-derived neuronal assays with phosphodeficient and phosphomimetic tau constructs targeting 14 disease-relevant sites, the authors convincingly show that phosphorylation state alters tau organization on microtubules and differentially impacts kinesin- and lysosome-based transport. The evidence is solid and well aligned with the conclusions and will be of interest to the biophysical, cell biology and neurodegenerative communities.

    2. Reviewer #1 (Public review):

      Summary:

      This work by Beaudet and colleagues aims at exploring the effect of phosphorylation on the formation of tau envelopes and consequently on axonal transport both in vitro on reconstituted microtubules and in human excitatory neurons derived from IPSCs.

      The authors found that a relatively widely used construct in which 14 serine or threonine residues often hyperphosphorylated in Alzheimer's disease are mutated to alanines (phosphodeficient) increases the density of tau envelopes compared to wildtype tau whereas a phosphomimetic (same residues mutated to glutamic acid) reduces envelopes density both in vitro and in human excitatory neurons derived from IPSCs.

      By analysing the trafficking of different kinesins (KIF1a and KIF5C), they observed different effects of tau phosphorylation status on the movement of these two motors.

      They then analyse transport of lysosomes by employing live imaging of lysotracker in human excitatory neurons derived from IPSCs transfected with wildtype, phosphodeficient or phosphomimetic tau observing that phosphodeficient tau seems to reduce transport of lysosomes while phosphomimetic increases transport compared to wildtype tau.

      Strengths:

      (1) The work aims to study a novel and underexplored topic in the tau field, tau envelopes, and investigate their relevance to Alzheimer's disease pathology.

      (2) Experiments are well conducted and of high quality.

      Weaknesses:

      Relying only on in vitro reconstituted microtubules and human neurons derived from IPSCs leaves some doubts about the relevance of these results for Alzheimer's disease considering the embryonic state of IPSCs-derived neurons, but the authors clearly discuss this point.

    3. Reviewer #2 (Public review):

      This manuscript examines how disease-associated hyperphosphorylation disrupts tau's role as a cooperative microtubule-binding regulator of intracellular transport. Using in vitro reconstitution assays and live-cell imaging in iPSC-derived neurons, the authors employ phosphomutant tau constructs (E14 to mimic hyperphosphorylation, AP to prevent phosphorylation) at 14 disease-associated residues to isolate phosphorylation effects independent of expression system-dependent PTM heterogeneity. The results show that hyperphosphorylated tau fails to form cooperative envelope-like structures on microtubules, instead binding diffusely and dissociating rapidly. In contrast, wild-type and phospho-resistant tau form cohesive envelopes that regulate motor protein access. At the single-molecule level, hyperphosphorylation reduces KIF5C inhibition while maintaining or enhancing KIF1A inhibition through altered processivity and detachment rates. In live neurons, hyperphosphorylated tau phenocopies tau knockout conditions, weakening tau-mediated inhibition of lysosome transport and increasing processive motility. The authors quantify tau binding using Gaussian mixture model-based image analysis and measure tau kinetics via FRAP, demonstrating that hyperphosphorylation-induced loss of cooperative binding correlates with dysregulated organelle transport. These findings establish a mechanism by which phosphorylation-driven disruption of tau's gatekeeper function on microtubules compromises axonal transport prior to aggregation in tauopathies.

      Comments on revised version.

      The authors did a good job responding to my comments and I support publication of the revised manuscript.

    4. Author response:

      The following is the authors’ response to the original reviews.

      We thank the reviewers for their thoughtful and constructive feedback. In response, we substantially revised the manuscript and clarified the rationale for using in vitro reconstitution assays and iPSC-derived neurons to determine how tau hyperphosphorylation alters its interaction with microtubules and its regulation of intracellular transport. We have also more clearly articulated how these findings relate to neurodegeneration and discussed the limitations of the model systems used.

      The primary concern of the reviewers was the justification for using COS7 cell lysates in reconstitution assays and iPSC-derived neurons as model systems. We have revised the manuscript to clarify that these experimental systems provided a means to isolate and examine how AD-related tau hyperphosphorylation alters tau-microtubule interactions and the regulation of intracellular transport. COS7 cells were selected because they are widely used for expression of mammalian proteins, including kinesins and tau. Human iPSC-derived neurons were chosen for their amenability to TIRF microscopy and transfection-based experiments, as well as the ability to compare CRISPR-generated tau knockout (MAPT-KO) neurons with their isogeneic control counterparts. Accordingly, we have revised the language throughout the manuscript to more clearly define the study’s objectives and emphasize that these systems were intentionally chosen as robust, well-controlled platforms for addressing specific mechanistic questions. We agree that they do not fully recapitulate AD pathology and that more representative models, such as mature, aged neurons or patient-derived neurons would be better suited to studying disease progression, and have included this limitation in the Discussion. However, because the central objective of this study was to dissect the mechanistic consequences of tau hyperphosphorylation on microtubule interactions and intracellular transport, we believe that these experimental approaches are well suited to address the questions asked.

      We also more explicitly addressed how background levels of phosphorylation may contribute to the effects observed with the pseudo-phosphorylation model of AD-related tau perturbations. We’ve addressed this by citing recent studies (Fan et al., 2025; Siahaan et al., 2026; Moretto et al., 2026) that quantitatively assess phosphorylation across expression systems and clarified how our experimental design, which directly compares WT, AP and E14 tau, effectively minimize uncertainty arising from background phosphorylation. While some degree of background phosphorylation is likely to be present, any resulting effects would be expected to occur consistently across all tau phospho-variants. We now discuss the limitation of our study that we did not directly quantify phosphorylation levels in cells.

      The reviewers also expressed concern about the potential influence of endogenous microtubule-associated proteins present in lysates and differences in tau occupancy on microtubules contributing to motility outcomes. To address this, we included additional analyses correlating tau intensity along microtubules with kinesin motility. We also expanded the Discussion to consider how tau competes with other MAPs for microtubule binding and how phosphorylation-dependent changes in tau–microtubule interactions may alter the MAP landscape. Consequently, the transport phenotypes observed with different tau phospho-variants may reflect both direct effects of tau and indirect effects arising from changes in MAP occupancy and competition on the microtubule lattice.

      We provide detailed, point-by-point responses to each reviewer comment below. We appreciate the thoughtful feedback from reviewers and are confident that the revisions, which include clearer language, strengthened justification of the experimental approaches, and additional supporting analyses, have substantially improved the clarity, rationale, and overall impact of the study.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This work by Beaudet and colleagues aims at exploring the effect of phosphorylation on the formation of tau envelopes and consequently on axonal transport, both in vitro on reconstituted microtubules and in human excitatory neurons derived from IPSCs.

      The authors found that a relatively widely used construct in which 14 serine or threonine residues, often hyperphosphorylated in Alzheimer's disease, are mutated to alanines (phosphodeficient), increases the density of tau envelopes compared to wildtype tau, whereas a phosphomimetic (same residues mutated to glutamic acid) reduces envelope density both in vitro and in human excitatory neurons derived from IPSCs.

      By analysing the trafficking of different kinesins (KIF1a and KIF5C), they observed different effects of tau phosphorylation status on the movement of these two motors.

      They then analyse transport of lysosomes by employing live imaging of lysotracker in human excitatory neurons derived from IPSCs transfected with wildtype, phosphodeficient or phosphomimetic tau, observing that phosphodeficient tau seems to reduce transport of lysosomes while phosphomimetic increases transport compared to wildtype tau.

      Strengths:

      (1) The work aims to study a novel and underexplored topic in the tau field, tau envelopes, and investigate their relevance to Alzheimer's disease pathology.

      (2) Experiments are well conducted and of high quality.

      Weaknesses:

      Relying only on in vitro reconstituted microtubules and human neurons derived from IPSCs leaves some doubts about the relevance of these results for Alzheimer's disease, considering the embryonic state of IPSCs-derived neurons.

      We agree with the reviewer that iPSC-derived neurons represent an immature state compared with the neurons most affected in Alzheimer’s disease. However, iPSC-derived neurons and in vitro reconstitution are robust experimental approaches that provide insight into (1) the effects of hyperphosphorylation on tau’s cooperative microtubules association and envelope formation, (2) how tau hyperphosphorylation affects the motility of kinesin motors that are sensitive to regulation by tau, and (3) how tau hyperphosphorylation alters the bi-directional transport of endogenous degradative organelles such as lysosomes. Our studies reveal the molecular effects of how hyperphosphorylation influences tau’s role in regulating intracellular transport and we believe that these findings will help to inform future studies examining how tau-related dysfunction first influences axonal transport, which would be expected to alter axonal health and homeostasis prior to the more severe pathological effects observed at later disease stages.

      We have included a paragraph under the subheading ‘Limitations of this study’ in the Discussion section to better contextualize our findings within the broader effort to understand tauopathies and Alzheimer’s disease. We clarify the limitations of using in vitro reconstitution and iPSC model systems on pages 20 and 21.

      Reviewer #2 (Public review):

      This manuscript examines how disease-associated hyperphosphorylation disrupts tau's role as a cooperative microtubule-binding regulator of intracellular transport. Using in vitro reconstitution assays and live-cell imaging in iPSC-derived neurons, the authors employ phosphomutant tau constructs (E14 to mimic hyperphosphorylation, AP to prevent phosphorylation) at 14 disease-associated residues to isolate phosphorylation effects independent of expression system-dependent PTM heterogeneity. The results show that hyperphosphorylated tau fails to form cooperative envelope-like structures on microtubules, instead binding diffusely and dissociating rapidly. In contrast, wild-type and phospho-resistant tau form cohesive envelopes that regulate motor protein access. At the single-molecule level, hyperphosphorylation reduces KIF5C inhibition while maintaining or enhancing KIF1A inhibition through altered processivity and detachment rates. In live neurons, hyperphosphorylated tau phenocopies tau knockout conditions, weakening tau-mediated inhibition of lysosome transport and increasing processive motility. The authors quantify tau binding using Gaussian mixture model-based image analysis and measure tau kinetics via FRAP, demonstrating that hyperphosphorylation-induced loss of cooperative binding correlates with dysregulated organelle transport. These findings establish a mechanism by which phosphorylation-driven disruption of tau's gatekeeper function on microtubules compromises axonal transport prior to aggregation in tauopathies. The paper provides interesting new knowledge for the field, but there are outstanding concerns that could be further addressed by the authors to strengthen and clarify the current manuscript:

      (1) Lack of Phosphatase-Treated Control and Explicit WT Phosphorylation Quantification

      Wild-type tau expressed in insect and mammalian cells is known to be phosphorylated by endogenous kinases (eg, GSK3, CDK5, MARK). The manuscript acknowledges this in the Discussion but provides no phosphatase-treated lysate control or quantification of endogenous phosphorylation on WT tau via phospho-specific Western blots. This leaves ambiguity about whether observed differences between WT and E14 reflect purely the introduced mutations or confounding baseline differences in phosphostate content.

      Tau contains ~85 putative phosphorylation sites and is modified by several kinases in cells. Studies by Siahaan et al. (2026) and Fan et al. (2025) provide detailed insight into tau phosphorylation heterogeneity, its role in protecting the microtubule lattice from severing enzymes, and the implications of phosphorylation patterns for aggregate formation. We reference these papers and include detailed description of these findings when initially establishing our justification for using pseudo-phosphorylation model.

      We used a pseudo-phosphorylation approach to test the effects of phosphorylation of specific residues in the proline-rich region and the pseudo-repeat domain in the C-terminus, which together with the microtubule-binding repeats, establish the minimal regions required for tau’s cooperative microtubule binding (Tan et al., 2019). This system enabled us to dissect the effects of tau phosphorylation without the added complexities of heterogeneity and multiple isoforms of tau that would otherwise be endogenously expressed. We’ve clarified these points in the revised manuscript (Pages 6, 7, 17, and 18).

      Background phosphorylation in the different phospho-variants used might contribute to the observed changes in tau’s MT interactions and regulation of transport. However, based on our results and the significance in the changes between the different phospho-variants, even if there is some basal level of phosphorylation, the results indicate that the effects of the pseudo-phosphorylation sites are strong enough to make observable changes above the basal levels of phosphorylation (see p. 6 of the revised manuscript).

      Disease-associated phosphorylation is likely more heterogeneous and dynamic than the pseudo-phosphorylation mutants used here, and phosphorylation at different sites may differentially regulate tau function (see p. 21 of the revised manuscript).

      (2) Limited Normalization of Motor Effects to Measured Tau Lattice Occupancy

      Although kinesin trajectories are classified inside vs. outside tau envelopes (inherently normalizing to local tau density), motor parameters are not systematically reported as functions of tau fluorescence intensity across all constructs. Co-purifying MAPs or microtubule-modifying enzymes in cell lysates is not quantified or excluded, leaving residual uncertainty about tau-specificity of observed motor inhibition. This should be at least acknowledged in the results section.

      As noted by the reviewer, it is challenging to compare conditions where the occupancy of tau on microtubules is dissimilar across conditions. To address this point, we performed a Spearman’s correlation analysis to compare how tau intensity affects kinesin dynamics along microtubules (Fig S3G). On page 12, our results show that kinesin dynamics are generally reduced in regions of high tau occupancy. However, in regions of comparable higher intensities, KIF5C is less inhibited by E14 tau, whereas KIF1A is less inhibited by AP tau.

      On pages 12 and 13, we acknowledge that while effects from other MAPs or motor proteins could potentially affect kinesin motility, we would expect that any effect from residual lysate components would be similar across tau phospho-variants.

      (3) Insufficient Citation of Prior Neuronal Tau Envelope Evidence

      In the Introduction, the authors state, "it was an open question if tau forms envelopes in neurons," but this understates existing evidence. Tan et al. (2019) report tau neuronal staining consistent with envelope formation, while Siahaan et al. (2021) provide more direct evidence in non-neuronal cells. The framing should acknowledge and integrate these prior findings.

      We agree with the reviewer that evidence from several studies using reconstitution systems, fixed neurons, and live cultured cells provides evidence of tau envelope formation in neurons. Specifically, tau envelopes have been observed along taxol-stabilized or GMPCPP-capped GDP microtubules in vitro (e.g., Dixit et al., 2008; Monroy et al., 2018; Tan et al., 2019; Siahaan et al., 2019), in 4% PFA-fixed and Triton X-100–extracted DIV7 mouse hippocampal neurons (Tan et al., 2019), and in live, non-neuronal U-2 OS cells following taxol treatment (Siahaan et al., 2022) or elevated pH (Siahaan et al., 2024). To our knowledge, our study is the first to demonstrate tau envelope formation in live neuronal cells under normal cell culture conditions. We revised the introduction (see pages 3 and 4) to more precisely position our findings within the context of prior studies.

      (4) Unclear Wording on Expression System-Dependent Phosphorylation

      The sentence "The phosphostate of tau is strongly dependent on the expression system" requires rewording. It is ambiguous whether this refers to the final phosphostate achieved after expression or the inherent phosphorylating capacity of each system. Clearer language would strengthen the methodological justification.

      On pages 6 and 7, we clarify the rationale for using COS7 cells to express GFP-tau and elaborate on recent papers demonstrating how different expression systems used to study tau (e.g., bacterial, insect, mammalian) produce tau with variable phosphorylation patterns (Siahaan et al., 2026; Fan et al., 2025).

      (5) Insufficient Quantification of Motor and Lysosome Transport Effect Magnitudes in Results Section

      The data on molecular motor motility and lysosome transport are densely described. The magnitude of effects (fold-changes, percentage differences) should be explicitly stated in the Results section when first presenting findings to orient readers to biological significance. For example, effect magnitudes for lysosome run lengths, velocities, and directional bias should be quantified in text, not left to figure inspection.

      We now incorporate the relevant quantifications in the text.

      (6) Incomplete Discussion of Projection Domain Necessity for Envelope Formation

      The Discussion states the projection domain is "a critical regulator of both tau-tau and tau-microtubule interactions," but does not engage with prior domain dissection work. Tan et al. (2019) found that the entire projection domain is not necessary for envelope formation in vitro. The authors should discuss which projection domain regions are specifically regulated by phosphorylation vs. required for cooperativity, providing a more nuanced interpretation than implied by their current framing.

      Tan et al. (2019) demonstrated that part of the proline-rich region (residues 198–244) within the N-terminal projection domain and the pseudo-repeat region within the C-terminus, together with the microtubule-binding repeats are the minimal region required to maintain tau’s ability to form cooperative envelopes along microtubules. We revised the text to better incorporate this previous work into the discussion and place our findings within this context. Our work demonstrates how phosphorylation within the proline-rich region and pseudo-repeats are important regulators of tau–tau cooperativity.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) It is unclear how the method implemented by the authors to identify tau envelopes works exactly (GMM and BIC) and how appropriate it is. It does not appear similar to what others have done in the literature on tau envelopes. Moreover, when checking the intensity plots along microtubules, both in Figures 1C and 2A, one is left to wonder if the observed differences are not simply caused by the different thresholds. Indeed, the intensity profiles do not seem greatly different between conditions in Figure 1C, whereas it is evident that the threshold is quite different, with it being lower for AP tau and higher for E14 tau, which explains the differences in how many envelopes are detected. The authors could also try to quantify in a different way (e.g. even just a threshold based on Average+SD) to see if the results remain the same?

      We used a GMM/BIC approach to avoid biased comparisons of tau envelope formation on microtubules across phospho-conditions. In TIRF assays, intensity signals are inherently inconsistent, making it difficult to directly compare fluorescence intensity signals on different microtubules across regions within the same field of view. Additionally, tau distribution between microtubules and in solution varied between conditions (e.g., background tau signal is elevated in E14 conditions compared to WT or AP tau). Given these challenges, we quantified and compared tau intensities on a per-microtubule basis, which produced more robust results. While we initially attempted the reviewer’s suggested approach of using average + SD, per-microtubule variability in minimum and maximum signal, along with differences in local background, prevented the ability to set a threshold that reliably captured intensity differences along microtubules across and within replicates. We now more clearly explain why we chose this approach (see page 7).

      (2) The authors should discuss the possibility that the presence of a GFP tag at the N terminus of tau could affect the formation of envelopes, given the importance of this region. Also, they refer to tau GFP in some points of the text and other times to GFP tau. As it would seem they have always used tau tagged at the N terminus, they should refer to GFP-tau in order to avoid confusion in the position of the tag.

      We agree with the reviewer that the position of the N-terminal GFP could influence the projection domain, However, all tau constructs carry the GFP tag at the same position and differ only in their phospho-site mutations. The correct nomenclature for “GFP-tau” is now consistently used throughout.

      (3) Previous work (Tan et al., 2019) has shown that different isoforms of tau have different propensities to form tau envelopes. The authors should specify in each figure which isoform of tau they are expressing.

      The tau isoform used throughout this study is 4R0N. The tau isoform is clearly identified in the revised text.

      (4) Figure 1 C-E: It would be interesting to see the size of envelopes quantified, also.

      We now include a comparison of the mean envelope width for each phospho-variant (Fig 1F).

      (5) Figure 2F: As the FRAP experiment is not on tau envelopes but generally on axonal tau, this needs to be clearly stated to highlight how this limits the link between the different FRAP dynamics and the behaviour of tau envelopes.

      We changed the text to indicate that we perform FRAP on axonal tau and not specifically tau envelopes.

      (6) The authors should discuss whether they expect Kif1a and Kif5c to be responsible for transporting lysotracker-positive vesicles in neurons? This does not seem to be the case based on a quick literature search. If these are not the motors responsible for the transport of lysosomes, why do the authors decide to look at the transport of these organelles and not others? Also, what is the rationale for studying the transport of lysosomes, an organelle that is mainly transported retrogradely, after identifying defects in kinesin transport? The authors could either study in vitro the effect of tau phosphorylation on the movement of a kinesin more directly linked to lysosomes (e.g. KIF5B, KIF1B) or study the transport of some other organelle which is mediated by KIF1A and KIF5C.

      We revised the text to clarify this point. Several studies have shown that kinesin-1 and -3 are strongly inhibited by tau, whereas kinesin-2 and dynein are less sensitive (Hoeprich et al., 2017; Chaudhary et al., 2018; and others). Within this context, we asked how phosphorylation alters tau’s inhibitory effects on motors that are most sensitive to tau. The in vitro reconstitution assays were not intended to isolate the effects of tau on lysosome-specific motors. Rather, they were used to determine how tau phosphorylation affects representative kinesin-1 and -3 motors that drive a substantial fraction of anterograde axonal transport and are among the most sensitive to tau-mediated regulation.

      We next examined lysosome transport using LysoTracker to investigate how tau phosphorylation influences bidirectional cargo transport. Lysosomes are transported by teams of kinesin-1, -2, -3, and dynein, making them a useful model for assessing the consequences of tau regulation in a more physiological context. Current models of bidirectional transport proposed that cargo movement emerges from tug-of-war, which is a result of a balance of forces generated by opposing motors. Under this assumption, strong inhibition of kinesin by tau would be expected to reduce anterograde transport and/or enhance retrograde transport by shifting this balance towards dynein. We have clarified throughout the manuscript that our goal was to determine how tau phosphorylation affects bidirectional transport and to interpret these findings within this context. Because defects in degradative pathways are thought to contribute to neurodegeneration, these experiments may also provide insight into how tau hyperphosphorylation disrupts lysosome function during disease.

      Although KIF5C and KIF1A are not the primary kinesin homologs responsible for lysosome transport, we expect that other kinesin-1 and kinesin-3 motors respond similarly to tau. The in vitro findings provide mechanistic insight into how tau phosphorylation could alter motor function and ultimately contribute to changes in lysosome trafficking and distribution within axons. On page 19, we further clarify that the magnitude of tau-mediated regulation is likely to vary among kinesin family members due to differences in their intrinsic motor properties, and that the effects on lysosome transport are therefore expected to be more nuanced than those observed for individual motors in vitro.

      (7) In the trafficking experiments with lysotracker in human excitatory neurons, there seems to be a large fraction of anterogradely transported lysotracker-positive organelles. Based on a quick search, it would appear this occurs frequently in IPSC-derived neurons, but it's not the case in primary neurons (see, for example, Kulkarni et al., 2022). Given that IPSCs-derived neurons maintain an immature embryonic maturation status (as correctly stated by the authors when mentioning that they express mainly 3R tau) and that neuronal maturation influences transport in primary neurons (e.g. Moutaux et al., 2018), the authors should discuss these aspects highlighting the possible limitations, especially considering the claim of importance of their results for Alzheimer's disease, a pathology that hits neurons at full maturation stages. Alternatively, they could perform a similar experiment in murine neurons at mature stages.

      See response to comments from reviewer 1 under ‘weaknesses’.

      (8) In the context of the previous point, the immature phenotype of IPSCs could explain the apparent discrepancy between the results obtained by these authors and previously published work (Hallinan et al., 2019), which found that mature hippocampal neurons expressing E14 tau had reduced transport of lysosomes. Moreover, these authors also described patches of higher intensity of tau along the axons formed by E14 tau compared to WT tau, which are closely reminiscent of tau envelopes. The authors should discuss these discrepancies.

      We agree that there are discrepancies between our findings and those reported by Hallinan et al. (2019). In that study, E14 tau was shown to misfold in cultured mouse hippocampal neurons, forming MC1-positive axonal aggregates that impair lysosome transport. In contrast, we observe nearly the opposite effect: E14 tau remains diffusely distributed in axons and produces a phenotype resembling tau knockout conditions, with enhanced lysosome transport.

      These differences may stem from methodological factors, including the neuronal models used (murine hippocampal cultures versus human iPSC-derived neurons), fixation and immunolabeling compared with live-cell imaging, differences in neuronal maturity (DIV), and the presence of endogenous tau versus our knockout-and-rescue approach. Importantly, tau aggregation may reflect later stages of disease progression, where aggregates physically clog axons leading to obstructed axonal transport rather than tau acting as a regulatory “roadblock” to specific motor proteins. We now cite this paper and compare our results with this study and discuss these discrepancies and their potential implications in the ‘Limitations of this study’.

      (9) Figures 4 and 5 are quite hard to read. Perhaps the distinction between proximal, mid and distal axon, although valuable, could be moved to the supplementary, maintaining an overall average, or the most significant of the 3 in the main figures to improve readability?

      We made substantial revisions to figures 4 and 5 and the associated analysis. Because the effects of tau on lysosomal transport were largely consistent across proximal, mid, and distal axonal regions, we combined these datasets and report overall transport trends within the axon (from ~50 µm distal to the AIS to ~50 µm proximal to the growth cone). The region-specific analyses and figures showing lysosome motility in each axonal segment have been moved to Supplementary Figure S4.

      (10) In the discussion, the authors write an entire paragraph on how their results are important to stress the importance of the N terminus of tau in the formation of tau envelopes. This is based on the fact that most of the residues mutated in the phosphomimetic and phosphodeficient constructs are located in the N-terminal projection domain. However, some of these residues are located in the C terminus of tau, which also appears to have a role in tau envelope formation (Tan et al., 2019). The experiments presented do not discriminate the phosphorylation of which of the 14 residues is important to mediate the effects. Hence, I feel this paragraph needs to be toned down or removed entirely.

      See response to comment 6 from reviewer 1.

      (11) The authors make a point of using tau produced in mammalian cells in the experiments performed in vitro, stressing the advancement compared to previous work that used tau produced in bacteria or insect cells. Although this is certainly closer to physiological conditions, the production is done in cancerous kidney cells, so I feel the author should highlight that neurons might drive a distinct phosphorylation pattern. Could recombinant tau be produced in neuroblastoma cells?

      In the revised manuscript, we’ve addressed this comment in the “Limitations of this study” page 21. We used COS-7 cells, which are not cancerous but immortalized fibroblast-like cells derived from African green monkey kidney obtained from ATCC. These cells were chosen because of their widespread use for protein expression and their high transfection efficiency. We agree that the physiology of COS-7 lysates is not directly comparable to that of neurons. Our intention was to convey that proteins expressed in mammalian systems undergo post-translational modifications and are produced by cellular machinery that more closely resembles neuronal systems than bacterial or insect expression platforms. Although neuronal cell lines such as neuroblastoma cells may appear more physiologically relevant, they are often difficult to transfect (Alabdullah et al., 2019). This can create practical challenges in equalizing protein concentrations and obtaining sufficient amounts of overexpressed tau from lysates. While methods exist to improve transfection efficiency, there is no literature that we found stating that these cells would yield protein expression characteristics more comparable to neurons than COS-7 cells. A more comprehensive evaluation of alternative expression systems would require a substantially deeper literature search or systematic characterization of multiple cell lines, which falls beyond the scope of this manuscript. Therefore, we relied on the robust COS-7 expression system and will clarify this rationale in the revised manuscript.

      Alabdullah AA, Al-Abdulaziz B, Alsalem H, et al. Estimating transfection efficiency in differentiated and undifferentiated neural cells. BMC Res Notes. 2019;12(1):225

    1. eLife Assessment

      This valuable study presents a promising framework for automated spike sorting in high-density extracellular electrophysiology. The core evidence is solid, but could be strengthened with extended simulations, broader benchmarking against existing sorters, and validation via expert curation. It will be of interest to the systems neuroscience community.

    2. Reviewer #1 (Public review):

      Summary:

      This work presents a flexible spike-sorting framework that allows users to run, swap, and benchmark individual modules commonly used in spike sorting. The paper argues that "opening the black box" is essential for understanding which components drive performance differences and for making progress toward more accurate and transparent spike sorting.

      Using this modular benchmarking pipeline, the work identifies electrode drift as a primary bottleneck for accurate sorting, and introduces an end-to-end sorter ("Lupin") that combines the best-performing modules and is reported to be on par or maybe even outperform existing spike-sorting packages on their benchmark. While the modules forming Lupin were chosen based on a single benchmark, end-to-end evaluation of different sorters now use multiple benchmarks, including different available datasets.

      Overall, this is a strong tool/resource contribution with clear potential to accelerate spike-sorting development and enable more rigorous comparisons. While not the main point of the paper and therefore less important, remaining claims regarding Lupin outperforming other sorters are not well supported. Lupin does not necessarily outperform all other sorters in real data if you consider 'finding most good units' more important than decreasing false positives - something that quality metrics post-sorting can take care off.

      Strengths:

      This work has high community value and practical utility. The effort to make benchmarking and spike sorting modules accessible and standardized is substantial, and likely to be broadly useful.

      Treating spike sorting as a set of interchangeable modules is a useful approach to some extent, and it enables targeted improvements rather than 'new sorters' popping up which are difficult to fully understand.

      Implementing this resource within SpikeInterface, an already widely used tool, will facility uptake and community contributions.<br /> Overall, I am positive about this manuscript as a resource paper. The core framework is compelling and timely.

      Weaknesses:

      (1) I appreciate the use of automatic curation tools to define the quality of units obtained from different spike sorters. However, a discussion on what is considered a good trade-off is missing. For example, one could argue that as long as you apply these curation tools post-sorting, finding more good-quality units is more important than keeping the number of false positives low, as these can be filtered out at a later stage by these curation methods. The manuscript currently seems to argue the balance between high number of good and low number of false positives is more important. This needs to be discussed and rationalized more explicitly, rather than using terms like 'clearly outperforms' and 'good trade-off'.

      (2) Although I agree that overfitting to specific data is a general problem with spike sorting (as you can also change parameters of individual sorters), and perhaps this is even required for the best results, I still miss an explicit discussion of this.

      (3) While the end-to-end evaluation is a very good addition, defining the best strategy for each module (which in turn led to choices what to use for Lupin) is still based on one benchmark only. This remains a weakness.

      Cmments on revised version:

      Previously identified weaknesses in limited support for Lupin being superior to other spike sorters, especially using only a single benchmark, have been largely addressed by 1) making superiority of Lupin less of a point in the manuscript and 2) adding more simulations and real datasets. Also, clarification on serial versus iterative spike sorting has been added to the discussion.

    3. Reviewer #2 (Public review):

      Summary

      Spike sorting, that is, assigning events detected in extracellular electrophysiology data to firing of individual neurons, is an inherently difficult computational problem involving multiple steps. The difficulty arises from low signal to noise, instability in signal due to relative motion of the tissue and recording sites, and large volumes of data. Experimental ground truth data - where the correct assignment of spikes in known - is not available in large enough quantities to test algorithms. This paper describes a tool for creating fully synthetic ground truth data and benchmarking the individual steps of spike sorting to dissect the impact of signal to noise, firing rate, and motion correction on each step. This information is used to construct an optimized algorithm for sorting these ground truth data. One result of particular interest is the dominant role of motion correction in degrading accuracy. Another important technical result is that motion correction via interpolation of the voltages traces yields similar accuracy to interpolation of the spike templates.

      Strengths

      The paper shows that useful insight can be gained through analyzing process step by step. While this analysis has also been done in papers presenting spike sorters (for example, Pachitariu (2024)) the tools presented here allow users and developers to do similar studies for their own work. This toolset will be useful to many labs, especially those working in less studied brain areas or model systems, cases where the tuning of standard spike sorting tools is not a good match to the data.

      Weaknesses/Limitations:

      The model ground truth data used in testing spike sorting and its components does not need to be a perfect match to experimental data to provide useful benchmarking. However, as with all measurements of spike sorting accuracy, extrapolation to experimental data can be complicated. Therefore, the insights gained concerning optimization of the individual steps should be interpreted as "correct for that model data. The comparison of the paper's new sorter to standard sorters on experimental recordings suggests that the benchmarking data is reasonable. Nevertheless, users of these tools will need to assess how well the simulated data matches their recordings.

    4. Reviewer #3 (Public review):

      Summary:

      In this manuscript, the authors describe two additions to an existing toolbox (SpikeInterface, Buccino et al., 2020, eLife). The first addition is an empirical simulator for extracellular recordings, in which spikes from predefined templates are added up with Gaussian noise. The second addition involves granting user-level access to intermediate processing steps along spike sorting algorithms. The authors demonstrate the toolbox by evaluating functions (e.g., event detection) or sets of functions (e.g., feature extraction + clustering) on their simulated data and suggest that a specific combination of function implementations provides performance improvement relative to kilosort4 (Pachitariu et al., 2024, Nature Methods).

      The validity of the work is poor. In particular, the simulator is unrealistic and the ground truth dataset is too short. Several spike sorters are used as straw men, and most sorters compared have never been described in peer-reviewed literature or in sufficient detail. The purely feedforward architecture of the modules is very limiting and irrelevant for modern sorters. Finally, the reporting of results is sporadic and does not follow scientific reporting standards.

      General comments:

      (1) Abstract, lines 14-16: "We then leverage these results to create a modular component-based spike sorter that can outperform Kilosort 4 on dense and large simulated recordings and produce similar quantitative results on real data." However, the artificial data are not "large" - they are very short. And "similar quantitative results" cannot be assessed on real data because those data do not have any ground truth. Because this is a revision and the authors have already received similar feedback from this Reviewer, I am not sure what to recommend.

      (2) In a previous comment, I indicated that the simulator itself is overly simplistic, and indicated that as far as I am concerned, the authors must improve it in one of two ways: (1) use a set of biophysical equations, with multi-compartmental modeling of currents and return currents; (2) use noised data from extracellular recordings; or (3) some combination thereof. The authors explained in their answer - but not in the MS - some of the shortcomings of biophysical simulators, but chose to do neither. My comment therefore remains unaddressed in the MS.

      (3) In a previous comment, I indicated that the duration of 10 minutes is too short. The authors chose to extend the duration to 30 minutes, which is insufficient for units that have low firing rates. Units that fire e.g., 0.1 spikes/s would have fewer than 200 spikes. If this MS is to be taken seriously, the simulation should be done on durations that (1) are similar to the potential applications - which may be many hours or even days; (2) allow identification of low-firing neurons. Therefore, about 3 hours is the bare minimum. Therefore, at present, my comment remains unaddressed.

      (4) In a previous comment, I indicated that some sorters have never been described in peer reviewed papers and therefore, they must either be removed from the present comparisons or be described in full. The authors chose to persist in relying on un-reviewed online documentation. Thus, my comment remains unaddressed, and the comparisons that involve those sorters (e.g., TDC, TDC2, SpyKING Circus 2) are simply invalid.

      (5) In a previous comment, I indicated that some sorters (SpyKING Circus 2 and TDC2) are effectively straw men and suggested to reorganize the MS to demonstrate its main goal. Specifically, I suggested to reorganize the manuscript so that after every module is evaluated separately based on a limited ground truth dataset, a single "best" sorter would be constructed, and then tested extensively (and compared to the de facto state of the art). Such reorganization would both demonstrate the utility of a modular approach and clarify the general usefulness of the outcome. The authors agreed that these are straw men and gave a historical account of why these were included. However, they did not explain these considerations in the MS itself, and chose not to reorganize the MS. Therefore, my comment remains unaddressed.

      (6) In a previous comment, I commented about the presentation, description, and interpretation of the results, and indicated that the choice to report point estimates makes any conclusions based on those results invalid. The authors did not make any changes to the reporting. Therefore, none of the results reported in the MS can be taken as an outcome of scientific inquiry. In other words, the MS does not deliver any solid findings - neither scientific nor methodological.

    5. Author response:

      The following is the authors’ response to the original reviews.

      As requested by all three reviewers we have added a new figure which applies our new end-to-end sorters on real openly available data. This demonstrates that our modular algorithms, optimized to work on simulated data, also performs well in real world cases. In addition, to demonstrate that our results are not due overfitting on simulated data on a specific probe geometry (Neuropixels 1.0), we added a Supplementary Figure demonstrating that our positive results for the end-to-end spike sorters are observed for numerous probe geometries (Neuropixels 2.0, SiNAPS, tetrode and Cambridge NeuroTech). We hope that the extended applications will convince the reviewers and readers of the robustness of our results.

      Reviewer #1 (Public review):

      Weaknesses:

      The reviewer identifies several weaknesses:

      (1) The main concern is the limited support for the claim that ’Lupin’ and individual modules’ outperform existing spike sorters.

      (2) Evidence is primarily from a single benchmark based on an intentionally simplified simulation. While the authors discuss the trade-offs between simulated and real data, the current evaluation does not provide enough diversity to justify claims of superiority.

      (3) While improving individual modules that run in a serial fashion could aid overall spike sorting performance, acknowledging that some end-to-end sorters work in an iterative fashion across multiple of these modules would be fair. Perhaps the optimal spike sorter is not a serial set of modules.

      (4) There is also a risk of benchmark overfitting. A modular approach makes it easy to select components that excel on specific benchmarks (or a specific project’s data characteristics) without generalizing.

      We would like to thank the reviewer for the comments and the valuable feedback. We revised our paper to answer the major concerns that were raised by the reviewer. Regarding the claims about Lupin (1,2), we modified the manuscript in two directions: (i) we attempted to stress that the goal of the paper is not the introduction of the new Lupin sorter per se, but rather presenting and highlighting the modularity of the sorting components framework. In doing so, we also toned down our claims of superiority; (ii) we added simulations on a diverse range of probes and included three real experimental datasets, obtained with different probes, in the results. Regarding the iterative aspect of some sorters – point (3) – we added a paragraph to the discussion highlighting that Kilosort4, unlike previous versions of KiloSort, [9] does not iterate over modules anymore, and thus it can be regarded as a serial algorithm. In addition, although all sorters we present are serial, our proposed framework does not prevent iterative schemes: for example, one could add a re-clustering step after a first template-matching pass. In fact, the modular framework makes this task even easier than before.

      Finally, related to point (4), we added some comments on the problem of overfitting with modular benchmarks in the Discussion, but we think that the risks have been mitigated with the addition of more end-to-end examples with various probe types and experimental data in the revised manuscript.

      The reviewer also points toward some possible ways to strengthen this work:

      (1) Evaluate on multiple simulation regimes, consider adding at least one biophysically detailed simulation, benchmark on multiple probe-geometries with neurons also clustered in different depth profiles (as this will affect drift solutions), and provide real-data validation. Even without full ground truth, real-data can be evaluated with expert curation, functional validation (e.g., refractory violations, quality metrics, unit waveform consistency), agreement across sorters, and consistency across time.

      (2) Related to real-data applicability, it is also important to acknowledge that modulatory approaches can enable overfitting to the needs of individual projects. Without real-data benchmarking (or benchmark diversity), it is unclear how the framework will guide users towards generalizable ’best practices’ rather than optimized configurations that work for their specific conditions.

      In response to these comments, the manuscript has been extended both with more simulated ground truth recordings and experimental data. For additional ground truth, we generated recordings with various probe geometries, to check that the results observed for our modular pipeline could generalize (see Supplementary Figure). Regarding real experiments, we chose three open dataset of various types (chronic and acute implantations, IMEC and Cambridge Neurotech devices) and compared the results of several sorters at a macroscopic level relying on high-level automatic curation tools [7, 3, 6] to quantify how many “good”, “oversplit” and “noise” units are found. This is now a new Figure 8 in our manuscript. We believe that the results from all these datasets demonstrate that Lupin is, as is claimed in the paper, on par with the most popular spike sorting algorithm, Kilosort4 [9].

      Regarding overfitting, we do not believe this is a direct consequence of the modular approach introduced in this article, but rather a general potential “risk” in the spike sorting field. Each spike sorter exposes a large array of parameters that users can tweak to attempt to optimize outcomes on specific datasets, but this end-to-end fine-tuning is hard to control and quantify. We believe that the modular benchmarks introduced in this paper may enable a finer, more controlled, and quantifiable parameter exploration. As an example, very high firing rates such as those observed in the cerebellum might require different parameters for peak detection than for neurons in the cortex. To test this, one could use our generation framework to mimic key macroscopic features of the system you are studying, and benchmark the peak detection step to find optimal parameters. Overall, we do not see this optimization strategy as a problem. Extracellular electrophysiology is so diverse based on brain regions, species, conditions, tasks, etc. that generalizable “best practices” might not exist.

      Reviewer #1 (Recommendations for the authors):

      (1) Tone down or further support the Lupin and specific modules’ superiority claims.

      This has been modified in the manuscript, and we added a final figure to discuss how Lupin is on par with Kilosort on real data, but with no claim of superiority.

      (2) Add benchmark diversity, this would test generalization and mitigate benchmark overfitting. Specifically: add more probe-geometries, and allow for different depth-profiles of clusters of neurons.

      We thank the reviewer for the suggestion, and indeed, we added some more benchmarks to convince the readers that the results observed can be generalized properly. More specifically, we extended the duration of the recordings to 30 min, and included additional benchmark datasets with four different probe geometries (Neuropixels 2.0, a Cambridge Neurotech probe, a SiNAPS probe and a tetrode).

      (3) Clarify how (x,y) positions of neurons are distributed around the shanks.

      This has been clarified in the Methods section. The (x,y) positions are generated uniformly within a rectangle covering the probe boundaries, plus a 20µm margin. Regarding the depth, z positions (distance from the probe) are drawn uniformly from the range [5,40] µm.

      (4) Add at least one real-data benchmark. Some suggestions for evaluation are: stability of firing rates over time, agreement across sorters, quality metrics, functional validation, expert curation.

      As suggested by almost all reviewers, we added some real-data benchmarks (see last figure). Since experimental data don’t have ground truth, we used automatic curation tools as a proxy for “goodness” of the results. We used automatic labels from Bombcell [3] and UnitRefine, which label units as good, multi-unit activity (MUA) and noise, and SLAy [7] for automatic merging, which correlates with the amount of putative oversplits. We felt it was fairer to use external curation tools instead of creating our own methods to assess quality.

      (5) Clarify recommendations for users facing drift. While your statement of ’not having drift is ideal’ is true, the reality is that many recordings have drift. Some practical solutions would be useful. For example, when to trust results and how to report drift sensitivity.

      The reviewer is right, and we added a sentence to clarify when motion correction methods should be used, in our opinions.

      (6) It would help to see what types of signals are most often missed, for example: low-SNR units, drifting units, bursty units, and show which modules affect which of these issues.

      This is already shown in Figures 4 of the manuscript, at the clustering level. These Figures show that cells with low firing rates and/or low SNR are most likely to be missed by all sorters. Our ground truth simulator does not include a bursting mechanism yet. We think that this would be an interesting aspect to simulate and we plan to include bursting units, with bursty spike trains and waveform modulation, in future releases. We thank the reviewer for the suggestion.

      (7) To overcome the issue of overfitting to specific datasets rather than generalization, it would help if a ’default’ or ’recommended starting point’ for users were described in more detail.

      We overcome the issue of overfitting by adding other artificial ground truth recordings (see Supplementary Figure S1), and also real world dataset (see Figure 8). In all these simulations, Lupin is used with default parameters, and this is, we believe, a good starting point. Of course, for very special needs (animal species, brain structures, ...), one might need to adapt parameters, but so as for any other sorters, and such an exploration of the parameter space is out of the scope of the current manuscript. This has been added in the Discussion.

      (8) A lot of the plots have ticks / labels too small to read in 100%, or show quite low-resolution. For example, Figure 3 and Figure 5. Please homogenize across all figures.

      The figures have been regenerated and homogenized as suggested.

      (9) Consider archiving the GitHub version used to generate the figures on Zenodo (DOI) for posterity.

      This has been done for the current state of the manuscript at https://zenodo.org/records/20407862 and the code is available at https://github.com/SpikeInterface/sorting_components_benchmark_paper

      (10) To make the manuscript more reader-friendly, I recommend adding graphics representing the different methods. For example, in Figure 3 one could add schematics of the two peak-detection methods.

      We thank the reviewer for this suggestion. However, this project and our manuscript is not introducing these methods, only re-implementing them in a modular framework. Hence we feel it is out of the scope of this manuscript to produce schematics of the many methods discussed in the paper readers should view the original sources to find out more information.

      (11) Discuss what is meant by KS-like clustering, as I was under the impression that KS4 is also iterative. This may be on the template-matching side, but it’s difficult to know where you draw the border between iterative clustering and (iterative) template-matching. Potentially, we would want to see these processes as one module together, as many sorters work in an iterative fashion across these steps.

      By KS-like clustering, we meant that the code is a direct port, in Python, of the clustering algorithm implemented in KiloSort4. However, we cannot guarantee that this is the exact same clustering method, because of the way the clustering of KiloSort is interleaved with some others steps part of the algorithm. To be more specific, Kilosort has its own special way of performing matched filtering, with a custom grid of templates generated internally with a higher resolution compared to the recording channels positions. These templates are then used to estimate a putative position of each spikes, and these positions are the ones that are used to start the clustering. Thus, the term KS-like comes from the fact that we do not reproduce this exact same mechanism. In SpikeInterface matched filtering is performed using an equivalent approach, but positions are not estimated in the exact same manner. Since the clustering code is equivalent, the results might differ. Regardless of these details, the clustering is not iterative. In former versions of KiloSort, the core algorithm used ideas from k-SVD algorithms, often used in the Machine Learning community. In such an algorithm, the goal is to learn a sparse dictionary of templates to reconstruct the signals, and indeed, there was some iterations between optimizing the templates and the spike times. However, this is no longer the case in KiloSort4. The algorithm works in a serial manner, following the global strategy mentioned in our paper.

      (12) Rather than showing results for one specific simulation, it would be more convincing if we saw the average result of multiple simulations. We don’t need to see individual neurons necessarily (e.g., Figure 3B and C, but this applies throughout the manuscript).

      While we tend to agree with the reviewer than averaging over multiple datasets might be more informative (as we did in the final figure, for end-to-end sorter comparison), we want to say that given the fact that we are using ground-truth recordings that are randomly generated, as long as we do not change the macroscopic properties of the recordings, results on various instances of the noise will be very similar, and averages might not be as informative as one could hope for. One option would be to vary the parameter space of these ground truth recordings, but then there are so many parameters (noise levels, firing rates, distributions of the cells, ...) than averaging everything, and/or even choosing what should be primarily studied is an open question on its own. However, to ease the redibility of the figures, individual neurons were removed from the plots.

      (13) Legends are often incomplete. For example, describe what individual dots are and what lines are in the different figures, even if it seems obvious.

      Legends have been updated.

      (14) Figure 7 would benefit from average thick curves for each model (optionally with individual thin lines for each instance, or error bars). Individual neurons can be left out. Figure 7E (and likewise Figure 5E) would benefit from having x labels to indicate precise labels, so the color attribution is reserved for specific models. It’s quite an intense ’search’ game to understand these figures.

      All the figures have been regenerated, and for the sake of readability, we removed the scatter plots for the individual neurons, focusing only on the averaged lines.

      (15) While I appreciate the effort is gigantic, it may be better for the reader to conclude that.

      This has been rephrased.

      Reviewer #2 (Public review):

      Major comments

      The model ground truth data used in the paper does not need to be a perfect match to experimental data to provide useful benchmarking. However, as with all measurements of spike sorting accuracy, extrapolation to experimental data can be complicated. Users of these tools will need to assess how well the simulated data matches their recordings.

      We agree with the reviewer that extrapolating our results to real data is difficult, and the same point was raised by the other referees. We have now extended the results to include three experimental datasets from different probes. Due to the lack of ground truth, we used automatic curation tools (Bombcell [3] and UnitRefine for labeling, SLAy for merging [7]) as a proxy for performance and showed that Lupin is on par with Kilosort4 on all datasets. We hope this gives users some idea of how well the new sorters will work on their data.

      Reviewer #2 (Recommendations for the authors):

      (1) Any comparison to experimental data would be welcome. Is it possible to add firing rate, amplitude, and template similarity distributions from measured recordings to the panels in Figure 2D?

      Experimental data is very diverse, depending on brain region, species, recording technology, etc. Instead of extending the comparison between simulated and experimental data, we rely on new Figure 8 to showcase the applicability and performance of the presented methods on real recordings.

      (2) Do yields of units passing basic quality metrics look different with the new sorter vs. established sorters? Another option that would help establish the range of applicability of the results would be a different model ground truth system, e.g., the hybrid ground truth data used by the authors in reference 7.

      As it can be seen on synthetic recordings (e.g Figure 7A-B), all sorters behave similarly with respect to firing rates, signal to noise ratios. To help establish the range of applicability, we added an extra Figure 8 in the manuscript, as described above.

      (3) Interpolation errors are likely larger for NP1.0 probes - which have 40 um vertical pitch - than NP2.0 probes, with 15 um pitch. If it’s possible to include even a small-scale comparison of the results from the 2.0 geometry, that would be very valuable for readers trying to decide what probe type to use.

      In the revised manuscript, we added new ground truth benchmarks with a NP2, and other, layouts to show that the key results do not depend on the geometry of the probe (see Supplementary Figure 1). But to address more specifically the point raised by the reviewer, we note that in the case of applying Lupin to NP2 layouts, there is only a loss of ≈ 7.5% of well-detected units between static and motion-corrected cases (see Supplementary Figure 1). Where when we apply Lupin to NP1 layouts (see Figure 7) there is a decrease of 20% of well-detected units. So clearly, a smaller pitch and the columnar arrangement of NP2 seems to help motion-correction methods, and thus reduce the failures due to motion. This has been added in the Discussion.

      About the text:

      (1) In panel E of Figures 5 and 7: Adding the name of the sorter algorithm under the bar charts would be helpful. It is encoded by the color of the outline of each box, but I found that cue rather subtle.

      The figures have been regenerated, with increased ticks and label fonts. We found that adding the names made the Figures too dense, and acronyms would not simplify the figures. In the end we decided not to add them.

      (2) Is the number of features (K) used for clustering already mentioned in the main text or methods? I couldn’t find it.

      We thanks the reviewer for pointing out this problem, and the answer (K = 5) has been added in the manuscript, in the methods section.

      (3) Equation 2 in Methods, defining accuracy, appears to be incorrect. I believe the correct equation is: accuracy = TPij/(Ni + Nj - TPij).

      The reviewer is right, and this has been corrected

      (4) In the abstract, line 18: Component based spike sorters => component based spike sorter.

      This has been corrected

      (5) Line 169 "we can artificially boost the signal-to-noise..." I don’t really see anything artificial about leveraging the extra information in neighboring sites. Maybe just remove that adjective?

      We removed “artificial” from the sentence.

      (6) Line 196-197, describing the result in Figure 3: Especially since 3A is a log plot, it would be helpful to add a percentage to the spikes missed on the high end. These are probably pretty unusual cases, under 0.1%?

      This has been commented in the text, but both because this number is only an approximation (the problem of pairing peaks between detected and ground truth is slightly ill-defined), and because it depends on the particular seeds and parameters of the artificial ground-truth, we avoided numerical values.

      (7) Line 567: "number of dimensions from M to K 5" => "number of dimensions from M to K[5]" That is, is the 5 meant to be a reference? Or is 5 the number of dimensions retained to describe the temporal waveform?

      5 is the value of K, and this has been corrected in the manuscript.

      (8) Line 585-590: Does this description of how to handle borders between groups of sites also apply to the KS-clustering method?

      For the KS-clustering methods, we used the original method implemented in KiloSort. In fact, KiloSort aggregates all the clusters found during the clustering steps, launched per bins, but duplicated templates (based on their shapes only) are removed before matching (and those are the cells at the borders found numerous times). After the template-matching step, once cells have been “populated” with theirs spikes, templates are once again removed and/or merged. However, as explained in the methods, currently this step has not been implemented in our framework. To be more explicit, during template matching, KiloSort stores the features of all discovered spikes, in a space where the effects of nearby spikes have been subtracted, to get a clearer picture of these features. In this “denoised” space, the clustering algorithms is launched again on all spikes (decimated), to assign labels.

      (9) Line 687: "In fact, a major main with Kilosort lies after the template matching step..." => "In fact, a major difference with Kilosort lies after the template matching step...".

      This has been corrected.

      Reviewer #3 (Public review):

      Major comments

      (1) The simulator itself has to be improved and extended. Right now, it simply generates, for every unit, a mother waveform from a sum of exponentials, scales that over channels, and then adds up multiple instantiations of every unit on every channel, along with noise. This is not a biophysical simulator: it is an ad hoc procedure, and the sentence "we firmly believe that.." (lines 482-483) does not make the procedure convincing. To make the simulator credible, the authors should: (1) use a set of biophysical equations, with multi-compartmental modeling of currents and return currents; (2) use noised data from extracellular recordings; or (3) some combination thereof.

      The reviewer is right when pointing out that the current ground truth generator is not “biophysical”, and this is why in the manuscript we used the terms “biophysically plausible”. However, we decided to changed this phrase to “phenomenological” in order to avoid confusion. We believe that our generation tool has the key ingredients to challenge (and also demonstrate limits of) modern spike sorters, which is the goal of the proposed simulator. Spike sorters performing well on such phenomenological simulated data should be a necessary, but not sufficient, condition to convince experimentalists that they will work well on real data. In previous papers [5, 4], we used MEArec [1], which relies on biophysical modeling of the reconstructed neurons to simulate extracellular potentials to generate ground-truth recordings. However, such biophysical simulations have two shortcomings: i) simulations are very slow and resource-hungry; ii) it is not guaranteed that the superior simulation environment (multicompartment modeling) translates to simulated data that are more similar to experimental data. In fact, the Kilosort4 paper [9] shows that the action potentials generated by MEArec [1] have an almost doubled duration compared to realistic data, which requires further ad-hoc parametrization. We believe this discrepancy can be due to the fact that virtually all multi-compartment models are built from in vitro slices, not in vivo recordings. Given these limitations, we decided to rely on a simpler but better controllable model for generating templates. Despite not being “biophysical”, the generation model can replicate, to some extent, the variability in waveforms (using different parameters for waveforms widths and spatial decays) and allows us to have a ground-truth model for drifting as well, that we can use to assess interpolation errors. Nevertheless, we agree with the reviewer that some aspects of the simulator can be improved, such as the structure of the noise in the data. We are currently working on the Spikeinterface side to improve the generation module so that it supports temporally correlated noise (spatial correlation is already supported and used in this manuscript).

      (2) The simulated dataset has to be extended in time. Maybe I missed something, but 500 units over 10 minutes, with some units having firing rates as low as 0.1 spikes/s, corresponds to some of the units firing an expected 60 spikes. This is clearly too short, and does not replicate the standard situation in extracellular experiments.

      We extended the simulated recording to an half an hour duration. However, as it can be seen in the Figures, this does not affect the main results of the paper.

      (3) The simulated dataset has to be extended in space. The choice of using NeuroPixels 1.0 geometry is a poor one. Many labs use other monolithic electrode arrays (MEAs, silicon probes, other rigid arrays); tetrodes remain a major tool, and flexible probes (polyimide, mesh) are evolving. Assessing algorithms over a single spatial architecture is likely to lead to local maxima in performance and potentially erroneous conclusions.

      We believe the that Neuropixels 1.0 geometry is a good choice: the NP1.0 paper it is the most highly cited paper about a high-density electrophysiology probe, suggested that it is currently the most widely used probe in the world. However, the reviewer is correct that there is a risk of overfitting. To demonstrate generalizability of our algorithm, we have added results obtained with NP2.0 layout, a Cambridge NeuroTech layout, a SINAPs probe layout and tetrodes. The results demonstrate that the Lupin sorter is not overfitted.

      (4) The existing spike sorters evaluated are not completely described. Some sorters (e.g., SpyKING Circus and KS4) were described in previous publications, but it is unclear whether the implementation that was used for the present tests is exactly the same as those previously published. More importantly, some of the sorters evaluated (e.g., TDC, TDC2, SpyKING Circus 2) were never described in a peerreviewed paper. This does not mean that they cannot be evaluated - but if they are, they must be described in full. Relying on the fact that the code is open source cannot replace a complete and accurate scientific description.

      The reviewer is rising a valid point, but we think it is beyond the scope of this paper to fully describe every component of every spike sorter mentioned in the paper. We made the deliberate choice to describe the sorters as chains of components in order to demonstrate the flexibility of our modular approach. Full details of every components can be found in the online documentation. In order to ensure the manuscript felt more complete and concrete, we revised the descriptions in our Methods section, in order to give some more details.

      (5) Related to the above, all relevant code should be made available online in permanent repositories, not only in author-controlled ones.

      We are not sure of what exactly is suggested by the reviewer. In order to clarify the situation, we pushed the notebooks and all the code needed to reproduce the Figures of the paper in a Zenodo archive https://zenodo.org/records/19695406

      (6) It is unclear why SpyKING Circus 2 and TDC2 are evaluated - these could potentially be described as straw men. I recommend reorganizing the manuscript so that after every module is evaluated separately based on a limited ground truth dataset, a single "best" sorter would be constructed, and then tested extensively (and compared to the de facto state of the art). Such reorganization would both demonstrate the utility of a modular approach and clarify the general usefulness of the outcome.

      Although we agree that these two sorters might be seen as straw men, we decided to keep them in the paper for various reasons. This has been clarified in the manuscript, but the primary reason is an historical one: the developers of these sorters decided to unite their efforts while designing new tools and algorithms, which led to the initial work on the modular framework described here. Because SpyKING CIRCUS and TriDesClous had to evolve for maintenance, it was decided to try to write a common “grammar” that would allow these two spike sorters to be described in the same framework. The reviewer is right in the fact that once the foundations were stabilized, most of the development efforts were put to Lupin, that was built as the best combination of all the expertise gained on the two aforementioned sorters. A second reason to keep them is that, once again, we decided that Lupin should not be the main focus of the paper. Of course, this is a nice illustration of what the modular framework can do, but we do not want to push it per se. What matters most is the methodology, especially since we can not claim, here, that Lupin would be the best spike sorter regardless of data types, probe geometries, .... We extended the paper with other probe geometries (see added Supplementary Figure 1) and real world data (see Figure 8), and we observed that Lupin was on par, and/or slightly better than KiloSort 4 with respect to number of False Positives for examples. But ultimately, the paper is really about the development of a common ecosystem such that all tools can be improved upon, at the community level.

      (7) The new algorithms developed, for example, clustering and template matching, have to be described in more detail, and demonstrated graphically on simple datasets. This can be done in supplementary material if the authors prefer not to extend the manuscript too much.

      As suggest by the reviewer, we tried to extend the methods of the clustering and the template matching steps, bearing in mind that some of them have already been published in detail elsewhere. We really want to underline that the central point of the paper is the modularity of the architecture, not so much the low-level details. To populate our framework and demonstrate its generalization, we implemented some key algorithms. But describing with schematics and in depth every individual methods is something that is not even done in papers focused on the spike sorters themselves. Later, the reviewer complains that the paper sounds like a technical report. Delving into more details would only amplify this problem.

      (8) This reviewer finds the description and interpretation of the results to be inadequate. As an example, focusing on Figure 5: The results in Figure 5A have to be supplemented and summarized as a scalar point estimate (e.g., median accuracy), an estimate of dispersion (e.g., using MAD, IQR, or SD), evaluated over multiple runs, and compared using statistical tests between tools and conditions (e.g., using a multi-dimensional analysis of variance, a mixed effect model, etc.). The results in Figure 5D must have an indication of dispersion. Any conclusions based on the numerical experiments must be based on these metrics and statistical evaluations.

      To try and simply the plots and their interpretation, we decided to remove the scatter plots of the individual neurons in all Figures, to really focus on the core trends. We believe that the results, such as the dependence on accuracy as a function of SNR, are difficult to capture with summary statistics, and that the results are best understood by looking at the plots we have made. If we wanted to make strong claims about one algorithm being more suitable for a specific task, then these summary statistics would be suitable. Instead, we are trying to demonstrate the general utility of the components framework, and we optimize Lupin simply to maximize the accuracy of each step.

      (9) The entire MS would benefit from expert proofreading; there are many language errors, mostly in indefinite articles and grammatical numbers.

      The manuscript has been intensively proofread by native english speakers.

      Reviewer #3 (Recommendations for the authors):

      (1) Lines 14-15: "...a... sorters...": either "...sorter..." or "...a... sorter...".

      This has been corrected

      (2) Line 30: "peak detection" or "event detection"?

      We prefer the term “peak detection", since this is exactly what the algorithm are looking for: spatio-temporal extrema in the signals

      (3) Line 30: the purely sequential structure of the modules is very limiting and generally incorrect. Template matching may replace peak detection, as is the case in many real-time hardware implementations. The feedforward process is limiting, and many sorters use feedback or multiple loops. It is unclear whether and how the proposed framework supports such structures.

      The reviewer raises a good point that we did not explain well in the manuscript. Since our framework is modular, with each component independent of the others, a developer has freedom to create “iterative" sorters. E.g. they could loop over a pair of steps until a criteria is met. In fact, this is one of the advantages of creating modular components. The examples we show in the manuscript are sequential feedforward structures, leading to this confusion. We have clarified the point in the text. We show sequential sorters because to our knowledge there are currently no iterative sorters in wide use. Older versions of Kilosort were iterative, but KiloSort4 is not. Modularity also allows us to isolate one component for a specific task. Hence, for a real-time implementation, we can pre-compute templates using the peak-detection and clustering components. Then these steps would be skipped for “online” sorting, which would only use the template matching component. The reviewer states that “template matching may replace peak detection”. Indeed, in Lupin, SC2 and TDC2, template matching does replace peak detection – the initial peak detection is only used to construct templates for downstream matching. We have clarified this point in the text.

      (4) Lines 62-63: Are all of these sorters supported by the spike Interface framework? Please include a table of which are and which are not. The same for every module.

      All the sorters listed are indeed supported by Spike Interface, and this has been added in the text.

      (5) Lines 84, 98, and elsewhere: TriDesClous 2 is mentioned. What about TriDesClous - is there such a sorter, and if yes, what is the scientific reference?

      TriDesClous (https://github.com/tridesclous/tridesclous) is a spike sorting pipeline that has not been properly published with a DOI, but that has been developed by the first author of the manuscript and has been used by many papers [8].

      (6) Line 84: Lupin - suggest reorganizing the manuscript around this sorter and evaluating it on multiple datasets.

      The paper is not about Lupin, and we tried to rewrite the manuscript in order to make this point more explicit. The paper is intended to be a proof of concept of the benefits that can be obtained thanks to a modular approach. This is why we do not want to reorganize the paper on Lupin itself.

      (7) Line 104, Figure 1, and elsewhere: What is the advantage of evaluating three sorters, if two are predicted to be worse than the third/state of the art? Suggest to reorganize the MS: (a) describe a proper simulator; (b) describe every individual modules: mention the existing algorithms and elaborate + demonstrate the new algorithms; (c) evaluate every individual module - on properly realistic dataset; (d) evaluate existing complete spike sorters + the proposed best combination - on the same ground truth dataset; (e) compare the best two sorters on multiple datasets. In addition, may identify the WEAKEST link in each sorter and demonstrate the improvement by replacing ("upgrading") that link alone.

      As clarified in the introduction and in the “End-to-end evaluation..." section, we decided to keep three sorters for various reasons. The first is historical: SpyKING CIRCUS 2 and TridesClous 2 were the first two sorters that motivated the creation of the sortingcomponents framework. Because both authors realized that they had so much code in common, they decided to unite their efforts while rewriting them and share some common building blocks. The second reason is that the paper is not about Lupin, per se, but more about the general philosophy of the modular architecture presented here. We want to push forward the idea, in the community, that we should share tools, ideas and algorithms in order to enhance the analysis pipelines. Keeping several sorters, even if sub-optimal, is a way to showcase the flexibility of the framework, and this is why we did not re-organize the manuscript as suggested by the reviewer.

      (8) Line 121: "can drastically cut the time.." - provide quantitative support. In general, avoid superlatives and unsupported statements.

      Spikeinterface has been primarily design to ease the comparison between spike sorting pipelines [2]. Thus all comparison metrics such as agreement matrices, false positives, false negatives, ... are available out of the box when using these Benchmark objects. It would be hard to provide a quantitative support for such a speedup since it will depend on the algorithm, but it allows developer to simply focus on the core implementation while benefiting for free of the whole ecosystem that will launch benchmarks and compare it, with appropriate metrics validated by a large community. Since this validation process, on its own, can be quite complex depending on the processing step, we truly believe the development gain is important, despite the fact that it might be hard to quantify. However, we rewrote the sentence in order to avoid superlatives.

      (9) Line 129: "powerful and fast way to generate artificial" - again, avoid statements with superlatives and lacking quantitative support.

      We rewrote the sentence to avoid superlatives.

      (10) Line 129: "powerful and fast way to generate artificial" - the assumptions made in constructing the artificial data critically and strongly affect the conclusions of the benchmark processes. For instance, it is well known that about 10% of the spikes in the cortex have a positive extrema, but the generator is limited to negative spikes. Also, the generator produces spatially-displaced and scaled versions of the waveform generated by the putative soma, but extracellular waveforms almost never behave that way. Even if the simulator is improved, it will always remain a simulator, and therefore the caveat should always be kept in mind.

      The reviewer is right: our simulator has some limitations compared to biological data. But the fact is that, even on synthetic data, there is still plenty of room for improvements of current spike sorting pipelines before even getting to real data, where ground truth are unknown. We believe that such ground truth simulator is a necessary, but not sufficient, condition to validate sorting algorithms. Other options such as hybrid recordings would also suffer from the same flaws, and might be even more questionable. In the revised version of the manuscript, we also added tests on real data to compare more qualitatively Lupin and KiloSort. The fact that results are in line with what is observed on synthetic data gives us confidence in our observations.

      (11) The authors should add a limitations section to the MS.

      Some limits have been more extensively discussed in the Discussion of the paper, with respect to the feedforward architecture, the validity of the ground truth data.

      (12) Line 129: Can the benchmark object be used with other data (e.g., existing)? The methods indicate that this is the case - please demonstrate.

      We think that a proper demonstration would be out of the scope of this paper, but indeed, as long as the user can provide a recording alongside with a sorting (exhaustive or not), then the Benchmark objects can be used. This allows the use of hybrid recordings, and/or manually curated datasets where users would be able to provide a ground truth. Of course, depending if the ground truth is exhaustive or not (if one knows the activity of all the neurons in the recordings), the metrics might not be the same. But everything is built-in in the object.

      (13) Line 144: 10 minutes is too short and nearly irrelevant. In particular, many units spike at rates much lower than 0.1 spikes/s, and in 10 minutes would emit less than a score of spikes (e.g., 0.01 spikes/s would accumulate, on average, 6 spikes..).

      We regenerated all the figures in the paper with 30 min long recordings, and the results remain similar to the 10 minute recordings.

      (14) Line 151: firing rates are not independent of the waveform as implicitly assumed; this should be accounted for, at least in the options in the simulator. Same for library waveforms - the firing rates should be a parameter that is optionally provided along with the waveform library.

      We are not sure what is meant by the reviewer. We believe that the point is that some cell types, with particular waveforms, have particular firing rates, such as fast-spiking interneurons. This could be dealt with in the current simulator, since users can provide, on a per cell basis, some particular values to generate the waveforms of the neurons. One could clearly imagine having some particular cell types with dedicated waveforms and firing rate parameters. However, for the sake of simplicity in the paper, we chose not to use such granularity. What the simulator can not do, at the moment, is to perform amplitude modulation of the templates as function of bursts for example. But this could easily be implemented, and should be part of future works to consolidate the generator.

      (15) Figure 2A, D: It is unclear why the specific zig-zag motion was simulated. Please rationalize, or use a motion from a real dataset. In particular, simulate (a) breathing-induced micro-motions, (b) gradual drift, and/or (c) a step jump.

      We used a zig-zag plus a Brownian motion to cover a rather broad range of continuous motion. Of course, as pointed out by the reviewer, the heterogeneity of real drifts is large, and again, we believe it is out of the scope of the manuscript to cover them all. In previous works, we explored how motion correction methods were working as function of the drifts [5], and the conclusion was that this simple continuous drift was already challenging enough to make the spike sorters fail. Adding discontinuities, as often encountered in experiment, would only make things worse. Finally, we added real world data (Figure 8) from three randomly picked dataset with heterogeneous probe geometries. They all come with some drift, that might be representative of what is typically dealt with.

      (16) Line 181: "more computationally demanding" - quantification?

      We forgot to add the reference to panel 3D, and this has been corrected in the manuscript.

      (17) Line 192: "(Figure 3A" - add ")"

      This has been corrected

      (18) Lines 195-6: "... not that missing a few spikes at peak detection will not have a large impact..." - this statement should be quantified.

      We reformulated the sentence, but the point here is simply that template-matching based algorithms use the template-matching step exactly for this purpose: to label spikes that would have been ignored/missed by the peak detection and clustering steps. The fact that template-matching based pipelines such as KiloSort, SpyKING-CIRCUS, ... outperforms clustering-based solution when detecting spikes [4] is a clear support for our sentence.

      (19) Lines 195-6: "... not that missing a few spikes at peak detection will not have a large impact..." - this statement, if correct, exposes a key weakness of the purely modular approach. For instance, assume that module 2.1 performance can be 0.5 and module 2.2 performance is 0.9, and that will have zero impact on the overall performance of a sorter that has module 4.1 as its fourth module, yielding an overall performance of 0.7. But when module 4.2 is used, module 2.1 performance of 0.5 results in an overall performance of 0.5, whereas module 2.2 performance of 0.9 translates to an overall performance of 0.9. The point should be clear now: evaluating every module in isolation cannot fully predict the behavior of the full system - even if a purely feedforward, single iteration (no loop) architecture is assumed. This requires algorithmic support in the proposed framework, and at the very least explicit discussion.

      The reviewer is right, benchmarking every module in isolation cannot fully predict the behavior of the full system. However, the fact that Lupin, built as an optimal combination of the components and can be on par, if not better in some situations, than KiloSort 4 on various artificial and real dataset makes a compelling argument in favor of assembling the best algorithmic pieces one after the other. In fully assembled spike sorting pipelines, the failures at one stage might be compensated by some algorithmic optimizations latter on. But still, being able to identify these failures, and eventually correct them at the appropriate level, i.e. as soon as they appear might be beneficial for the development of the tool, and to ease their readability/maintenance. This has been added in the discussion.

      (20) Line 203: "we hope that future efforts": This statement is strange for two reasons: (a) if the authors hope for it, why not simply do it? (b) if the authors cannot do it / it is out of scope, the proper place to mention their hopes is in the Discussion section.

      This has been moved to the Discussion section

      (21) Line 232: "motion is only partially compensated for" - so what is the utility of the compensation? This becomes clear later, but the statement is obscure at this point - please clarify.

      This has been clarified.

      (22) Figure 4A, right: The fact that the lines do not reach unity even for high FRs suggests that the FRs are NOT the key limiting factor. Please provide a scalar measure and a 2D analysis of the accuracy as a function of FRs and SNRs, separately for static and motion-corrected.

      We are not sure that we understand what is being requested by the reviewer. A scalar measure with a 2D analysis of the accuracy as a function of FRs and SNRs would mean, per case (static and motion corrected) at least 4 panels, thus a total of 8 panels. This seems like an overly dense figures. Further, a scalar measure is unlikely to provide more insight into the analysis.

      (23) Line 275: "kriging method" - describe.

      For a full description we refer readers to the original paper describing the method [9] and other work on motion correction [5]. We have added a brief description of the method in the text in the "Motion interpolation reduces spike sorting performance" section.

      (24) Lines 309-310: Where are the templates from in this case - true or estimated? If the latter, say it.

      This has been clarified.

      (25) Line 317: This is the point where I almost gave up - the MS up to this stage seemed very much like a technical report and is not organized in a clear, results-oriented manner. Even for a Methods-oriented paper, one expects to learn the key results, but those are lost in this MS.

      While we agree that this part of the manuscript is rather technical, this is something that we believe is at the core of some scientific questions that have not been yet properly addressed by most of the spike sorting pipelines. The idea of performing template-matching, i.e. to seek for spatio-temporal pattern in the signals while at the same time compensating only partially the motion has never been properly benchmarks. While template-matching is clearly a game-changer in static recordings, i.e. without motion, the extra-addition of motion might limit its performances. We firmly believe that our modular approach can be used to isolate such core questions from the whole integrated pipelines, and guide both users and developers.

      (26) Figure 6A, second and third panels: The fourth (rightmost) panel shows that there are differences between the "true" and "interpolated" templates, but this cannot be seen in the two middle panels - probably because too much information is overloaded on every channel. My suggestion is to show a reduced number of channels, but for each of those, show all five waveforms (corresponding to different spatial positions of the centers of mass) in a NON-overlapping display.

      The figure has been redrawn, as requested by the reviewer.

      (27) Figure 6 and the entire issue of the failure in motion correction + lines 331-332: It is not 100% convincing that the failure is not at the DREDGE level - please demonstrate that the motion is estimated perfectly.

      We checked that the DREDGE algorithm is correctly estimating the motion. To convince the reader this is the case, this has now been shown in Figure 2, panel D, where the simulated motion is displayed on top of the motion estimated by DREDGE. As shown, both are very similar.

      (28) Figure 6 and the entire issue of the failure in motion correction: If motion is estimated perfectly and interpolation is done properly, then is the failure an outcome of discrete spatial sampling? In other words, if the movement in space over time (in the simulator) were limited to discrete steps that correspond exactly to the positions of the electrodes, would the motion correction allow perfect performance? Stated differently: if the electrodes were not 15 or 20 micro-meters apart but rather 1 micro-meter apart, would the difference (e.g, between Figure 4A/4B, or between Figure 5A and Fig. 5B) be reduced? Please check and report.

      The reviewer is raising an interesting point, and indeed, this is something that we are planning to investigate. We felt it would have been too technical for the scope of this paper, which is aimed at showcasing the modular framework and its possibilities. But we are planning to explore such failures with ultra-dense probes as has been done in the DREDGE paper [10]. We also have the intuition that errors are originating from failures of discrete spatial sampling: hence the larger the sampling, the more pronounced the errors.

      (29) Figure 6B: It is impossible to discern which line is which - this underscores the general point of summarizing every CDF by a scalar with measures of dispersion (i.e., descriptive statistics) + statistical testing (i.e., quantitative statistics).

      While we agree with the reviewer that the plots are dense, they convey the global message of the paper and are the same that were used in [9]. To ease the comparison, we decided to stick to this representation.

      (30) Figure 6D: add error bars.

      There are no error bars, since there was only a single run in this figure, i.e. on a single dataset.

      (31) Line 318, line 328: So what is the result here - that the interpolation idea is not useful? If yes, what is the source of the failure? Is it discrete/poor spatial sampling as suggested above? Something else?

      The result is that interpolation, rather than motion estimation, is the bottleneck in recoding with motion. A careful analysis using data from dense probes, as already stated above, would be the focus of further work. But we suspect that errors results mostly from bad interpolation due to discrete spatial sampling.

      (32) Line 334: up to this point, there are no quantified results - and actually, no clear results at all.

      We rewrote the section accordingly, to make the key observations more striking. The goal of this section is not to propose quantified results and say which interpolation methods is the best (a full paper should be devoted to this question), but rather to show the possibilities offered by the modular framework, looking at questions that have been not yet properly addressed in spike sorting algorithms.

      (33) Line 345: "This is likely due" - provide support? Example?

      We rephrased the sentence accordingly.

      (34) Line 349: "this is mostly because both clustering..." - but feature detection was evaluated together with clustering, so a conclusion specifically about clustering is unwarranted. This is a general point - can the proposed software/modular framework evaluate feature extraction separately from clustering? If yes, please demonstrate.

      The reviewer is right about the fact that feature extraction was evaluated together with clustering, however, we want to stress that it is exactly the same method used by all the clustering algorithms developed in the paper. Thus, even if the feature detection method might not be optimal, it is not introducing any biases. Currently, almost all spike sorting algorithms these days are using PCA (or truncated SVD) as a feature detection method, on nearby channels where peaks are detected, and this is why we decided to only focus on this method.

      (35) Lines 358-374: If SpyKING-CIRCUS 2 and TriDesClous do not provide any advantage, why include them in the MS? Many other sorters could be included as well. In other words, what do we LEARN from the failure of these sorters to compete with KS4 and Lupin? If nothing, please remove them. If something, please state the conclusions clearly.

      Both these software have advantages, but we decided to keep them as they offer direct illustrations than implementation of fully integrated pipelines, other than Lupin, are possible within the proposed framework. TriDesClous2 is fast, and SpyKingCircus2 (in line with Spyking Circus [11]) is mostly tailored for in-vitro data, and thus can scale for more than thousands of channels while some other algorithms might not [9]. Of course, listing all the pros and cons of each software would be impossible within the scope of the paper, but we decided to keep them for illustrative purpose.

      (36) Line 369: "the main advantages of these sorters lie in their modularity" - modularity is an advantage only if it is useful for something - if the sorters are modular but perform poorly, the point is nullified.

      We agree with the reviewer, but we would like to point out that here, none of the modular sorters shown in the paper are performing poorly, and all have pros and cons as said in the previous point.

      We believe that this is important to show that modularity can offer some options both for users and developers with respect to probe geometries, animal species, data types, ...

      (37) Line 371: "on our dataset" - this is a very important limitation. See general comments, and discuss in the Limitations section of the Discussion.

      We extended the paper by adding more datasets, for different probe layouts, and also real datasets with meta comparison of state of the art spike sorters (see Supplementary Figure S1).

      (38) Figure 7C: the difference between the cyan (KS-like) and the orange (KS4) lines makes the KS-like irrelevant. Please improve or remove.

      The reviewer is right, and we removed the KS-like pipeline, since it at the stage of the manuscript it is not yet exactly like KiloSort 4.

      (39) Line 379: "each individual algorithmic steps" - should be "step".

      This has been corrected

      (40) Line 380: Is Lupin available for download and usage as a separate, standalone package - i.e., not as part of the spike interface framework? If not, please make it available. And if yes, indicate this clearly.

      Lupin is part of the SpikeInterface project, and thus can not be installed in a standalone mode, outside of the SpikeInterface framework

      (41) Line 385: "our... gigantic...effort" - avoid superlatives, especially to self.

      This has been removed

      References

      (1) A. P. Buccino and G. T. Einevoll. Mearec: a fast and customizable testbench simulator for ground-truth extracellular spiking activity. Neuroinformatics, pages 1–20, 2020.

      (2) A. P. Buccino, C. L. Hurwitz, S. Garcia, J. Magland, J. H. Siegle, R. Hurwitz, and M. H. Hennig. Spikeinterface, a unified framework for spike sorting. Elife, 9:e61834, 2020.

      (3) J. M. J. Fabre, E. H. v. Beest, A. J. Peters, M. Carandini, and K. D. Harris. Bombcell: automated curation and cell classification of spike-sorted electrophysiology data.

      (4) S. Garcia, A. P. Buccino, and P. Yger. How do spike collisions affect spike sorting performance? Eneuro, 9(5), 2022.

      (5) S. Garcia, C. Windolf, J. Boussard, B. Dichter, A. P. Buccino, and P. Yger. A Modular Implementation to Handle and Benchmark Drift Correction for High-Density Extracellular Recordings. eNeuro, 11(2):ENEURO.0229–23.2023, Feb. 2024.

      (6) A. Jain, R. Greene, C. Halcrow, J. A. Swann, A. Kleinjohann, F. Spurio, S. Graff, A. Pan-Vazquez, B. Kampa, J. Gall, S. Grün, O. Winter, A. Buccino, M. H. Hennig, and S. Musall. UnitRefine: A community toolbox for automated spike sorting curation.

      (7) S. Koukuntla, T. DeWeese, A. Cheng, R. Mildren, A. Lawrence, A. R. Graves, K. E. Cullen, J. Colonell, T. D. Harris, and A. S. Charles. SLAy-ing oversplitting errors in high-density electrophysiology spike sorting. bioRxiv: The Preprint Server for Biology, page 2025.06.20.660590, 2025.

      (8) J. Magland, J. J. Jun, E. Lovero, A. J. Morley, C. L. Hurwitz, A. P. Buccino, S. Garcia, and A. H. Barnett. Spikeforest, reproducible web-facing ground-truth validation of automated neural spike sorters. Elife, 9:e55167, 2020.

      (9) M. Pachitariu, S. Sridhar, J. Pennington, and C. Stringer. Spike sorting with Kilosort4. Nature Methods, 21(5):914–921, May 2024. Publisher: Nature Publishing Group.

      (10) C. Windolf, H. Yu, A. C. Paulk, D. Meszéna, W. Muñoz, J. Boussard, R. Hardstone, I. Caprara, M. Jamali, Y. Kfir, D. Xu, J. E. Chung, K. K. Sellers, Z. Ye, J. Shaker, A. Lebedeva, R. T. Raghavan, E. Trautmann, M. Melin, J. Couto, S. Garcia, B. Coughlin, M. Elmaleh, D. Christianson, J. D. W. Greenlee, C. Horváth, R. Fiáth, I. Ulbert, M. A. Long, J. A. Movshon, M. N. Shadlen, M. M. Churchland, A. K. Churchland, N. A. Steinmetz, E. F. Chang, J. S. Schweitzer, Z. M. Williams, S. S. Cash, L. Paninski, and E. Varol. DREDge: robust motion correction for high-density extracellular recordings across species. Nature Methods, 22(4):788–800, Apr. 2025.

      (11) P. Yger, G. L. Spampinato, E. Esposito, B. Lefebvre, S. Deny, C. Gardella, M. Stimberg, F. Jetter, G. Zeck, S. Picaud, et al. A spike sorting toolbox for up to thousands of electrodes validated with ground truth recordings in vitro and in vivo. Elife, 7:e34518, 2018.

    1. eLife Assessment

      This important study investigated the effects of the nutrient microenvironment on key elements of cellular morphology, function, and metabolism in both human iPSC-derived and human fetal RPE cultures. The analysis compared commonly used iPSC-RPE media formulations (MEMα, DMEM-HG/F12 basal media), alongside human plasma-like medium (HPLM) in an effort to model a more physiologically relevant culture environment. The authors provide compelling evidence that standard culture media differ significantly in key aspects of retinal pigment epithelium biology, particularly regarding metabolic profiles. Further clarification of the biological replicates and data normalization methods will help maximize the impact of this manuscript.

    2. Reviewer #1 (Public review):

      Summary:

      The authors utilize both human iPSC-derived RPE and human fetal RPE cultures to interrogate the effect of various types of commonly used cell culture media on several key biological- and disease-relevant RPE properties. These include a comparison of RPE morphology, polarity, transepithelial electrical potential, lipid metabolism, autophagy, and targeted metabolomic profiles across 6 different media compositions.

      Strengths:

      This manuscript is very well written, and data are presented in a well-organized manner. The authors address media composition as a fundamental variable that will influence the interpretation of assays performed in RPE cell cultures, particularly metabolic studies. Figure 6 provides a useful summary of the study's findings across commonly used media types, and the manuscript's discussion offers insight into which media may be best suited to address specific experimental questions. Overall, this manuscript will not only serve as an important resource for vision scientists utilizing RPE culture models, but it also serves to remind the broader cell biology community of the importance of considering the potential (confounding) experimental effect(s) of various culture media and to consider tailoring the selection of culture media types to the specific experimental question.

      Weaknesses:

      (1) While the authors report that iPSCs were obtained from several healthy patients and that at least two clones were generated from each individual, it is not clear to this reviewer whether the experiments with each culture media type were performed on the same set of iPSC-RPE in each case. The authors mentioned iPSC differentiation variability as a limitation, but it would be helpful to understand (and quantify) the experimental variability that may exist with the same culture media using iPSCs from different patients and/or separate iPSC clones from the same patient.

      (2) Since a major purpose of the manuscript is to highlight how cell culture conditions influence RPE biology and metabolism, it would be helpful to also report whether Mycoplasma testing was performed and confirmed to be negative across all cell lines.

      (3) The effect of culture media on mean RPE area and hexagonality was compared in this study. Interestingly, Figure 1E demonstrates higher mean RPE cell area but also substantial variability in cell area for media 2 (MEM-alpha and B27) and media 4 (HPLM and B27). It would be helpful to include a discussion of the potential biological implications of variable cell area across these 2 media types.

      (4) The authors speculate that FBS-containing media may encourage a more mesenchymal or de-differentiated state. This could be experimentally determined by interrogating mesenchymal markers (alpha-SMA, fibronectin, etc) by immunoblot and/or immunofluorescence microscopy, similar to how the RPE markers were evaluated in Figure 1.

      (5) It would be helpful for the discussion section to include a comparison of key differences (where they exist) between iPSC-RPE and fetal RPE across culture media types.

    3. Reviewer #2 (Public review):

      Summary:

      In this study, Lim et al. provide a comprehensive analysis of the metabolic and physiologic effects of different media compositions on iPSC-RPE. This analysis includes commonly used iPSC-RPE media bases (MEMα, DMEM-HG/F12 basal media) as well as human plasma-like medium (HPLM) in attempts to establish a more physiologically relevant culture environment.

      Strengths:

      The analyses in this study provide a very thorough survey of metabolic function as well as an RPE-relevant physiologic characterization. This will be a great resource for optimizing assay conditions for disease-based studies using iPSC-RPE.

      Weaknesses:

      In the Seahorse studies provided in Figure 3. basal readings for OCR are abnormally low compared to Oligomycin treatment and background, suggesting difficulties with the assay. Findings should be taken with caution.

    4. Reviewer #3 (Public review):

      Summary:

      The authors systematically compare six culture-media formulations using induced pluripotent stem cell-derived retinal pigment epithelium and fetal retinal pigment epithelium. They examine cell morphology, marker expression, barrier function, polarized secretion, lipid accumulation, ultrastructure, mitochondrial respiration, glycolytic function, and intracellular and extracellular metabolites. The results demonstrate that culture-medium composition and the choice of serum or B27 supplementation substantially influence retinal pigment epithelium phenotype and metabolism. Rather than identifying a single optimal medium, the study provides a comparative framework to guide medium selection according to the biological question being investigated.

      Strengths:

      The head-to-head comparison of six media under otherwise similar culture conditions addresses an important source of variability in retinal pigment epithelium research. The study uses a broad range of complementary approaches, including imaging, transepithelial resistance, electron microscopy, extracellular flux analysis, and targeted metabolomics. The inclusion of both induced pluripotent stem cell-derived and fetal retinal pigment epithelium increases the potential relevance of the findings across different cell sources. The matched comparisons of serum and B27 supplementation within MEMα and human plasma-like medium are particularly informative because they help distinguish supplement-associated effects from those caused by the basal medium. Overall, the dataset has the potential to serve as a valuable resource for selecting culture conditions and interpreting findings across retinal pigment epithelium studies.

      Weaknesses:

      The most important limitation is that the experimental unit and degree of biological replication are not clearly defined. It is unclear whether individual observations represent independent donors, clones, differentiated lines, culture preparations, wells, images, or sections. This makes it difficult to determine the independence, robustness, and generalizability of several comparisons.

      The metabolic analyses also require additional methodological clarification. For intracellular metabolomics, the culture format, cellular biomass, extraction volume, pooling strategy, and normalization method are not reported sufficiently. Normalization of extracellular measurements to unspent medium accounts for differences in starting metabolite abundance but not for differences in cell number or biomass. Similarly, normalization of intracellular signals to medium 1 does not correct for differences in the amount of cellular material extracted.

      For the Seahorse experiments, the main figures present unnormalized values even though the media produce differences in cell number, size, and protein content. These raw measurements represent total metabolic activity per well and may not reflect activity per cell. It is also unclear how normalization was performed because the Methods describe cell-count and protein measurements from two wells, whereas the stress tests included five to six wells per condition. In addition, measurements obtained after transfer into a common assay medium reflect metabolic adaptations retained from the preceding culture conditions rather than real-time metabolism within the original media.

      Other limitations include insufficient information about the biological replication underlying the sub-RPE deposit analysis and the inability to fully interpret the effects of X-VIVO 10 because its composition is proprietary. Finally, public availability of the underlying metabolomics data would be important for a study intended to serve as a community resource.

    5. Author response:

      We thank the editors and reviewers for their thoughtful feedback on our manuscript. We are encouraged that they recognized the importance accounting for media consideration when modeling in vitro disease phenotypes, as well as the value of this dataset as a resource for the field. We appreciate the points raised regarding biological replicates and data normalization methods, and we plan to address these fully in our formal response and in revisions to the manuscript. Below, we provide preliminary responses to several comments and indicate how we anticipate addressing them in the revised manuscript.

      (1) Reviewer 1 & 3: Unclear definition of iPSC lines/clones used in data generation.

      We thank the reviewers for pointing out this ambiguity in the Methods. The project was completed with multiple donor iPSC lines, with at least 2 clones generated from each line, and each experiment was performed using at least three independent iPSC RPE lines. To minimize confounding variability, we took several precautions where possible: each experiment was performed within the same culture plate (coated with Matrigel from the same lot number), using RPE seeded at the same time to ensure comparable maturity, and RPE of the same passage number were used across multiple experiments to reduce de-differentiation or senescence effects. We agree that RPE derived from different iPSC differentiation batches can vary. For this reason, all iPSC RPE used in this study were generated from a single differentiation attempt. Because the goal of this project was to isolate the impact of nutrient composition on RPE phenotype, rather than to characterize variability arising from clonal or donor differences, we did not stratify our analysis by clone or donor. For imaging-based assays, multiple fields were selected at random from each well to ensure representative sampling. TEM analysis of sub-RPE deposits was performed with n=3 independent filters per medium condition, with three panoramic sections imaged per filter. We will add these details, including the number of clones used per iPSC line, to the revised Methods to clarify experimental unit and level of replication for each assay and will include sample number in legends.

      (2) Reviewer 2: In the Seahorse studies provided in Figure 3. basal readings for OCR are abnormally low compared to Oligomycin treatment and background, suggesting difficulties with the assay. Findings should be taken with caution.

      We thank the reviewer for this careful reading. The apparent discrepancy likely arises from comparing the raw OCR trace (left) versus the background-subtracted “Basal respiration” bar graph (right) in Figure 3A. In the raw trace, basal OCR (~45-65 pmol/min) is appropriately higher than both the oligomycin-treated (~35-45 pmol/min) and background (~30-45 pmol/min) rates, as expected. The bar-graph value is smaller (~7-25 pmol/min) only because the non-mitochondrial rate has been subtracted out, whereas the oligomycin plateau in the raw trace has not. These raw basal values fall within Agilent's recommended starting range for the XFe96 platform (~20-160 pmol/min), consistent with the modest basal energy demand of quiescent, differentiated RPE rather than an assay problem. A recent survey of 530 published Cell Mito Stress Tests [1] found that 17% report the implausible result of maximal OCR below basal OCR, and higher basal rate can compromise the FCCP-stimulated maximal rate. We titrated our cell numbers before this assay to ensure a clear FCCP response. The substantially increased maximal rates in all six media conditions indicate a technically sound assay. We will clarify this calculation in the revised Methods.

      (3) Reviewer 3: The metabolic analyses also require additional methodological clarification. For intracellular metabolomics, the culture format, cellular biomass, extraction volume, pooling strategy, and normalization method are not reported sufficiently. Normalization of extracellular measurements to unspent medium accounts for differences in starting metabolite abundance but not for differences in cell number or biomass. Similarly, normalization of intracellular signals to medium 1 does not correct for differences in the amount of cellular material extracted.

      We agree that these methodological details require clarification. Intracellular metabolomics was performed on RPE lysates collected from 12-well plates (n=3 independent wells/RPE lines per medium), each extracted separately without pooling. RPE were scraped directly into a fixed volume of 300 µL chilled 80% methanol per well, regardless of the medium condition; 10 µL of the resulting lysate was dried together with an internal standard (nicotinamide-D4), reconstituted in 100 µL of mobile phase, and 5 µL was injected for LC-MS/MS analysis. Media were processed in parallel using an identical workflow: 50 µL of conditioned media was collected at 24 and 48 hours, of which 10 µL was mixed with 40 µL cold methanol, and 10 µL of the resulting supernatant was dried with internal standard, reconstituted in 100 µL mobile phase, and 5 µL injected for analysis.

      Biomass data, including average nuclei count and total protein content reported in Supp. Fig. 2D, were obtained from RPE cultured in parallel under identical conditions. Because nuclei count and protein content did not consistently agree with one another across the six media, metabolite intensities were not normalized to either protein content or cell number. Instead, the intensity of each metabolite was instead normalized to Medium 1 to allow relative comparison across conditions, avoiding an additional, potentially skewed layer of correction from an imperfect biomass metric. We acknowledge this as a limitation of the method: since RPE size and biomass differ across media, our fold-changes reflect metabolite pool per well rather than per cell, which could over- or under-represent true per-cell differences in media that yield especially large or small RPE. We will clarify these details in the Methods and Discussion.

      (4) Reviewer 1: Since a major purpose of the manuscript is to highlight how cell culture conditions influence RPE biology and metabolism, it would be helpful to also report whether Mycoplasma testing was performed and confirmed to be negative across all cell lines.

      We thank the reviewer for the suggestion to include this information. To confirm, Mycoplasma testing was performed on all cell lines used in this study with negative results. We will add this information in the revised Methods.

      (5) Reviewer 3: public availability of the underlying metabolomics data would be important for a study intended to serve as a community resource.

      The metabolomics data has been deposited to UCSD Center for Computational Mass Spectrometry (CCMS) repository (Dataset: MSV000095024) and will be made publicly accessible upon publication.

      References:

      (1) Ransy C, Boissan M, Hammad N, Bouaboud A, Issad T, De Dieuleveult M, Miotto B, Ye M, Pasmant E, Bouillaud F. Extracellular flux analyses indicate low ATP yield and require refinement for accurate determination of maximal oxygen consumption rate. Sci Rep. 2026 Jun 10;16(1):18344.

    1. eLife Assessment

      This study reports high-resolution cryo-EM structures of Paracoccus trimethylamine N-oxide demethylase and advances the intriguing hypothesis that the enzyme may be bifunctional, coupling TMAO demethylation to formaldehyde capture at a distal tetrahydrofolate-binding site through an enclosed intramolecular tunnel. Supported by biochemical assays and molecular dynamics simulations, the structural findings are important and potentially of broad interest, particularly the unusual oligomeric architecture and the proposed conduit for a reactive intermediate. However, evidence that the conduit mediates formaldehyde channeling remains incomplete, because the current HCHO-THF measurements do not demonstrate enzyme-dependent transfer or distinguish downstream HCHO consumption from indirect effects of THF on enzyme activity.

    2. Reviewer #1 (Public review):

      Summary:

      Thach et al. report on the structure and function of trimethylamine N-oxide demethylase (TDM). They identify a novel complex assembly composed of multiple TDM monomers and obtain high-resolution structural information for the catalytic site, including an analysis of its metal composition, which leads them to propose a mechanism for the catalytic reaction.

      In addition, the authors describe a novel substrate channel within the TDM complex that connects the N-terminal Zn²⁺-dependent TMAO demethylation domain with the C-terminal tetrahydrofolate (THF)-binding domain. This continuous intramolecular tunnel appears highly optimized for shuttling formaldehyde (HCHO), based on its negative electrostatic properties and restricted width. The authors propose the hypothesis that this channel facilitates the safe transfer of HCHO, enabling its efficient conversion to methylenetetrahydrofolate (MTHF) at the C-terminal domain as a microbial detoxification strategy. Experimental data that shows an involvement of TDM in the reaction of HCHO with THF is less convincing.

      Strengths:

      The authors provide convincing high-resolution cryo-EM structural evidence (up to 2 Å) revealing an intriguing complex composed of two full monomers and two half-domains. They further present evidence for the metal ion bound at the active site and articulate a hypothesis for the catalytic cycle. Substantial effort is devoted to optimizing and characterizing enzyme activity, including detailed kinetic analyses across a range of pH values, temperatures, and substrate concentrations. Furthermore, the authors validate their structural insights through functional analysis of active-site point mutants.

      In addition, the authors identify a continuous channel for formaldehyde (HCHO) passage within the structure and support this interpretation through molecular dynamics simulations. These analyses suggest an exciting mechanism of specific, dynamic, and gated channeling of HCHO. This finding is particularly appealing, as it implies the existence of a unique, completely enclosed conduit that may be of broad interest, including potential applications in bioengineering.

      Weaknesses:

      Although the idea of an enclosed channel for HCHO is compelling, the experimental evidence supporting enzymatic assistance in the reaction of HCHO with THF is less convincing. The linear regression analysis shown in Figure 1C demonstrates a THF concentration-dependent decrease in HCHO; however, it is well established that HCHO and THF can react spontaneously in a non-enzymatic manner, raising the possibility that the observed effect does not require enzymatic involvement. I appreciate the authors' clarification that the data in Figure 1 were not intended to demonstrate enzymatic channeling or catalytic involvement in the HCHO-THF reaction, and that the assay does not distinguish between changes in HCHO production and downstream consumption. The authors' revised statement "Overall, these findings suggest that TDM-mediated TMAO demethylation generates HCHO, which can subsequently react with THF, potentially linking TMAO breakdown to one-carbon metabolism." is appropriately cautious and leaves open the mechanism underlying this process, which is consistent with the evidence presented.

      Overall, the authors were successful in advancing our structural and functional understanding of the TDM complex. They suggest an interesting oligomeric complex composition which in my opinion should be investigated with additional biophysical techniques.

      Additionally, they provide an intriguing hypothesis for a new type of substrate channeling. However, additional kinetic experiments focusing on HCHO and THF turnover by enzymatic proximity effects are required to strengthen this potentially fundamental finding. If this channeling mechanism can be supported by stronger experimental evidence, it would substantially advance our understanding and knowledge of biologic conduits and enable future efforts in the design of artificial cascade catalysis systems with high conversion rate and efficiency, as well as detoxification pathways.

      Comments on revised version.

      It is unfortunate that no additional experimental evidence supporting the proposed channeling mechanism could be provided. In the absence of such evidence, I remain somewhat hesitant to regard the evidence as "solid" rather than "incomplete," particularly given that the authors themselves acknowledge that the current experiments do not establish enzymatic channeling. However, the rest of the manuscript is stronger despite this limitation.

    3. Author response:

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      Thach et al. report on the structure and function of trimethylamine N-oxide demethylase (TDM). They identify a novel complex assembly composed of multiple TDM monomers and obtain high-resolution structural information for the catalytic site, including an analysis of its metal composition, which leads them to propose a mechanism for the catalytic reaction.

      In addition, the authors describe a novel substrate channel within the TDM complex that connects the N-terminal Zn<sup>2+</sup>-dependent TMAO demethylation domain with the C-terminal tetrahydrofolate (THF)-binding domain. This continuous intramolecular tunnel appears highly optimized for shuttling formaldehyde (HCHO), based on its negative electrostatic properties and restricted width. The authors propose that this channel facilitates the safe transfer of HCHO, enabling its efficient conversion to methylenetetrahydrofolate (MTHF) at the C-terminal domain as a microbial detoxification strategy. Experimental data that shows an involvement of TDM in the reaction of HCHO with THF is less convincing.

      Strengths:

      The authors provide convincing high-resolution cryo-EM structural evidence (up to 2 Å) revealing an intriguing complex composed of two full monomers and two half-domains. They further present evidence for the metal ion bound at the active site and articulate a hypothesis for the catalytic cycle. Substantial effort is devoted to optimizing and characterizing enzyme activity, including detailed kinetic analyses across a range of pH values, temperatures, and substrate concentrations. Furthermore, the authors validate their structural insights through functional analysis of active-site point mutants.

      In addition, the authors identify a continuous channel for formaldehyde (HCHO) passage within the structure and support this interpretation through molecular dynamics simulations. These analyses suggest an exciting mechanism of specific, dynamic, and gated channelling of HCHO. This finding is particularly appealing, as it implies the existence of a unique, completely enclosed conduit that may be of broad interest, including potential applications in bioengineering.

      Weaknesses:

      Although the idea of an enclosed channel for HCHO is compelling, the experimental evidence supporting enzymatic assistance in the reaction of HCHO with THF is less convincing. The linear regression analysis shown in Figure 1C demonstrates a THF concentration-dependent decrease in HCHO; however, it is well established that HCHO and THF can react spontaneously in a non-enzymatic manner, raising the possibility that the observed effect does not require enzymatic involvement. I appreciate the authors' clarification that the data in Figure 1 were not intended to demonstrate enzymatic channelling or catalytic involvement in the HCHO-THF reaction, and that the assay does not distinguish between changes in HCHO production and downstream consumption. However, the statement "these findings show that TDM carries out two linked reactions: TMAO demethylation at one active site, and the HCHO produced can condense with THF at the C-terminal domain, connecting TMAO breakdown to one-carbon metabolism" (page 2) still implies a mechanistic and functional coupling that is not supported by the presented data and appears inconsistent with the authors' clarification. In light of this, I recommend revising this statement to avoid implying mechanistic or functional coupling between the two reactions unless additional experimental evidence is provided.

      We thank the reviewer for this clarification. We have revised as per recommendation (page 2).

      “Overall, these findings suggest that TDM-mediated TMAO demethylation generates HCHO, which can subsequently react with THF, potentially linking TMAO breakdown to one-carbon metabolism.”

      Overall, the authors were successful in advancing our structural and functional understanding of the TDM complex. They suggest an interesting oligomeric complex composition which should be investigated with additional biophysical techniques.

      Additionally, they provide an intriguing hypothesis for a new type of substrate channelling. Additional kinetic experiments focusing on HCHO and THF turnover by enzymatic proximity effects would strengthen this potentially fundamental finding. If this channelling mechanism can be supported by stronger experimental evidence, it would substantially advance our understanding and knowledge of biologic conduits and enable future efforts in the design of artificial cascade catalysis systems with high conversion rate and efficiency, as well as detoxification pathways.

      Reviewer #2 (Public review):

      Summary:

      The manuscript reports a cryo-EM structure of TMAO demethylase from Paracoccus sp. This is an important enzyme in the metabolism of trimethylamine oxide (TMAO) and trimethylamine (TMA) in human gut microbiota, so new information about this enzyme would certainly be of interest.

      Strengths:

      The cryo-EM structure for this enzyme is new and provides new insights into the function of the different protein domains, and a channel for formaldehyde between the two domains.

      Weaknesses:

      (1) The proposed catalytic mechanism in this manuscript does not make sense. Previous mechanistic studies on the Methylocella silvestris TMAO demethylase (FEBS Journal 2016, 283, 3979-3993, reference 7) reported that, as well as a Zn2+ cofactor, there was a dependence upon non-heme Fe2+, and proposed a catalytic mechanism involving deoxygenation to form TMA and an iron(IV)-oxo species, followed by oxidative demethylation to form DMA and formaldehyde.

      In this work, the authors do not mention the previously proposed mechanism, but instead just say that elemental analysis "excluded iron". This is alarming, since the previous work has a key role for non-heme iron in the mechanism. The elemental analysis here gives a Zn content of about 0.5 mol/mol protein (and no Fe), whereas the Methylocella TMAO demethylase was reported to contain 0.97 mol Zn/mol protein, and 0.35-0.38 mol Fe/mol protein. It does, therefore, appear that their enzyme is depleted in Zn, and the absence of Fe impacts on the mechanism, as explained below.

      The proposed catalytic mechanism in this manuscript, I am sorry to say, does not make sense, for several reasons:

      (i) Demethylation to form formaldehyde is not a hydrolytic process; it is an oxidative process (normally accomplished by either cytochrome P450 or non-heme iron-dependent oxygenase). The authors propose that a zinc (II) hydroxide attacks the methyl group, which (a) is unprecedented, (b) even if it were possible, would generate methanol, not formaldehyde.

      (ii) The amine oxide is proposed to deoxygenate, with hydroxide appearing on the Zn - unfortunately, amine oxide deoxygenation is a reductive process, for which a reducing agent is needed, and Zn2+ is not a redox-active metal ion;

      (iii) The authors say "forming a tetrahedral intermediate, as described for metalloprotease," but zinc metalloproteases attack an amide carbonyl to form an oxyanion intermediate, whereas in this mechanism, there is no carbonyl to attack, so this statement is just wrong.

      So on several counts the proposed mechanism cannot be correct. Some redox cofactor is needed in order to carry out amine oxide deoxygenation, and Zn2+ cannot fulfil that role. Fe2+ could do, which is why the previously proposed mechanism involving an iron(IV)-oxo intermediate is feasible. But the authors claim that their enzyme has no Fe. If so then there must be some other redox cofactor present. Therefore, the authors need to re-analyse their enzyme carefully and look either for Fe or for some other redox-active metal ion, and then provide convincing experimental evidence for a feasible catalytic mechanism. As it stands the proposed catalytic mechanism is unacceptable.

      Revised version. The authors have essentially not changed the proposed mechanism. They have removed the reference to zinc metalloproteases, but still propose a mechanism mediated only by Zn2+. As explained above, attack by zinc (II) hydroxide is unprecedented and would generate methanol, not formaldehyde, and amine deoxygenation is a reductive process that cannot be fulfilled by Zn2+. So the proposed mechanism is still not feasible at all. The authors now say that "oxidative chemistry....remains unresolved", I'm sorry, but that is not acceptable.

      I have urged the authors to re-examine the metal content of their enzyme, In the Supporting Information (Figure S5) they give ICPMS data that indicates a Zn stoichiometry of 0.5 mol Zn/mol protein, and Fe is not detected. Have the authors analysed for other redox active metals? The authors say that there is no evidence for any other metal binding site, but there is only 50% occupancy of Zn in their protein, so could there be a different metal ion present in place of Zn in the other 50% of the protein, that accounts for the observed activity?

      Since there is clearly a major discrepancy here, the onus is on the authors to explain the discrepancy, rather than just returning with the same data. For example, they could treat the enzyme with EDTA to remove all metals (and check the treated enzyme by ICPMS), and then add different metal ions to test activity with different metals (could even titrate with different molar equivalents of metal ions). They could then test a range of different redox-active metal ions.

      We have re-examined our data and repeated experiments on the reviewer's opinion. We have repeated the IC-PMS several times with different preps, including full scans (data presented). Our enzyme is active, but no iron signal is detected. Moreover, the experimentally determined structure does not support the presence of a non-heme iron-binding site (Bugg TDH, Ramaswamy S., doi:10.1016/j.cbpa.2007.12.007, and other papers). More detailed response in the recommendations to authors.

      (2) Given the metal content reported here, it is important to be able to compare the specific activity of the enzyme reported here with earlier preparations. The authors have now done this in the revised version.

      (3) The consumption of formaldehyde to form methylene-THF is potentially interesting, but the authors say "HCHO levels decreased in the presence of THF", which could potentially be due to enzyme inhibition by THF. Is there evidence that this is a time-dependent and protein-dependent reaction? Not yet addressed.

      We thank the reviewer for this important point. At present, we have not performed detailed time-dependent or protein-dependent analyses to determine whether the observed decrease in HCHO levels in the presence of THF reflects enhanced downstream consumption or indirect effects, such as inhibition of TDM activity by THF. We acknowledge that further kinetic and protein-dependence studies will be important directions for future work.

      Also in Figure 1C, HCHO reduction (%) is not very helpful, because we don't know what concentration of formaldehyde is formed under these conditions; it would be better to quote in units of concentration, rather than %. This point has been addressed by the authors in the revised version.

      (4) Has this particular TMAO demethylase been reported before? It's not clear which Paracoccus strain the enzyme is from; the Experimental Section just says "Paracoccus sp.", which is not very precise. There has been published work on the Paracoccus PS1 enzyme, is that the strain used? Details about the strain are needed, and the accession for the protein sequence. Addressed in the revised version.

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      As noted above, there is still a major problem with the proposed mechanism not being feasible, and remaining questions about the presence or absence of a redox-active metal ion in their enzyme. They should:

      (1) Re-examine for other metal ions (apart from Zn and Fe) using ICPMS. The redox metal ion could, in theory, be some other transition metal ion.

      (2) Seek evidence for the role of metal ions in the activity of this enzyme, for example, by treating with EDTA to remove metal ions, and then adding different metal ions, to correlate activity with a particular metal ion.

      We thank the reviewer for these valuable suggestions regarding the metal identity and its functional role in TDM activity. We also acknowledge the reviewer’s concerns regarding the catalytic mechanism. In response, we have substantially revised Scheme 1 and the associated Discussion text to focus on the observed interactions of the substrate (TMAO) and products (DMA and HCHO) within the Zn<sup>2+</sup>-containing active site, rather than proposing a detailed catalytic mechanism that is not fully supported by the current data.

      (1) We repeated the ICP–MS analysis using a wide-range full-scan survey to examine the presence of additional metal-associated isotopes beyond Zn and Fe. The corresponding experimental details have been added to the revised ICP–MS Methods section (Page 9). Full-scan ICP–MS profiling of purified TDM detected Zn as the predominant associated metal species (Author response image 1A). In contrast, signals corresponding to Fe and other transition metals were either undetectable or present only at trace levels comparable to, or lower than, those observed in the digested HNO<sub>3</sub> solution control. To further validate this observation, we performed targeted ICP–MS quantification for both Zn and Fe on the same purified samples. These measurements confirmed that Fe was below the detection threshold, whereas Zn was consistently detected at an approximate ratio of 0.5 Zn<sup>2+</sup> per protein monomer (Author response image 1B, C, Figure S5).

      The observed 0.5 Zn<sup>2+</sup>-to-protein stoichiometry is consistent with the previously discussed 2 full-length + 2 half-domain (2+2½) assembly. In this complex, only the intact core domains retain the complete metal-binding motif, whereas the truncated half-domains lack the Zn<sup>2+</sup>-binding region. Consequently, only two metal-binding sites are expected per assembled complex, in agreement with the ICP–MS measurements. We additionally note that Zn<sup>2+</sup> was not intentionally supplemented during purification. Based on the current cryo-EM and biochemical data, both metal-binding sites in the full-length subunits appear similarly occupied, with no evidence for asymmetric metal loading.

      Importantly, the previously published FEBS Journal model proposed an Fe<sup>2+</sup>-binding site based on metal analysis and homology modeling rather than direct experimental determination. The residues implicated in Fe<sup>2+</sup> binding are well resolved in our experimental maps and do not define a metal-coordination environment compatible with a second mononuclear metal-binding site. Consistent with the ICP–MS results, the experimental structure provides no evidence of a second metal-binding site. Nevertheless, the enzyme remains catalytically active under these conditions.

      (2) We attempted metal depletion experiments using EDTA treatment to evaluate the functional role of the bound metal ion. However, removal of metal ions resulted in rapid protein aggregation, preventing subsequent activity measurements. These observations suggest that the bound Zn<sup>2+</sup> ion plays an important role in maintaining the structural integrity and stability of the TDM complex. While metal reconstitution experiments would be informative, the aggregation observed following metal depletion precluded a meaningful assessment of alternative metal ions in the current study.

      Author response image 1.

    1. eLife Assessment

      In this valuable study, Robben et al. describe a 3D beta-cell spheroid culture platform, that allows high-throughput monitoring of cytoplasmic calcium concentrations and insulin secretion. The authors demonstrate measurements of calcium signals comparable to those recorded in primary pancreatic islets. The authors validate the method by culturing MIN6 cells in a 3D-culture system, and show solid evidence of its utility by recording calcium signals in a high-throughput format and characterizing these calcium signals using pharmacological tools. This methodology demonstrates the utility of the 3D beta-cell spheroids for screening pharmacological modulators of pancreatic beta-cell function.

    2. Reviewer #1 (Public review):

      [Editors' note: this version has been assessed by the Reviewing Editor without further input from the original reviewers and the authors have satisfactorily answered the previous reviewer's comments.]

      Summary:

      They use cultures of insulinoma MIN6 cells that form spheroids in a micro-patterned PEG-hydrogel to measure Ca2+ oscillations in multiple cells simultaneously.

      Strengths:

      They demonstrate that insulinoma spheroids are formed in multi-well plates and that Ca2+ imaging can be performed on them.

    3. Reviewer #2 (Public review):

      Summary:

      The study by Robben et al., show 3D beta-cell spheroid platform, a valuable tool allowing high-throughput monitoring of cytoplasmic Ca concentrations and insulin secretion, with Ca signals comparable to those recorded in primary islets. The authors demonstrate a solid method to culturing MIN6 cells in a 3D culture system, recording Ca signals in a high-throughput format and characterizing these Ca signals using pharmacological tools, including TRPM3 channel and K-ATP channel modulators. This highlights the utility of the 3D beta-cell spheroid for screening new ion channel modulators in beta-cells of the pancreas.

      Strengths:

      - The study shows that the MIN-6-based 3D beta-cell model is better to study Ca-signaling and insulin secretion compared to 2D culture of single MIN-6 cells.<br /> - The method allows imaging of Ca signaling in many spheroids in parallel followed by collecting medium to measure insulin release and correlate both effects.<br /> - The authors demonstrate that this system is suitable for screening new pharmacological modulators and used as an agonist of the ATP-sensitive potassium channel (diazoxide) and the agonist and antagonist of the TRPM3 channel.

    4. Reviewer #3 (Public review):

      Summary:

      The primary objective of this study is to develop high-throughput screening assays utilizing homogeneous 3D cell cultures that more accurately replicate the intricate architecture and cellular communication found in tissues. The authors have chosen pancreatic islet β-cells as a model system to evaluate agents that modulate insulin release, which is particularly relevant given the increasing prevalence of diabetes mellitus-a significant global health concern. Moreover, the incorporation of human-based 3D spheroids, organoids, or organ-on-chip technologies into drug discovery protocols is essential for enhancing clinical translation, as candidate compounds identified using animal models have often demonstrated limited success in clinical settings.

      Strengths:

      This study was thoughtfully planned and skillfully carried out. The use of micropatterned hydrogels to observe 19 spheroids at once is an ingenious aspect, which has been effectively validated with Ca microfluorography. Overall, I found this investigation to be exceptionally well-executed and free from notable flaws, as the results clearly back up the conclusions. Additionally, the developed method achieved the proposed aims, providing a high-throughput format with 3D cultures. I believe this study deserves publication.

    5. Author response:

      The following is the authors’ response to the original reviews.

      Reviewer #1 (Public review):

      Summary:

      They use cultures of insulinoma MIN6 cells that form spheroids in a micro-patterned PEG-hydrogel to measure Ca <sup>2+</sup> oscillations in multiple cells simultaneously.

      Strengths:

      They demonstrate that insulinoma spheroids are formed in multi-well plates and that Ca <sup>2+</sup> imaging can be performed on them.

      Weaknesses:

      The type of equipment and multi-wells used for the experiments are very specialized to be used as a common tool. Insulinoma cells are tumoral cell lines that divide, unlike primary beta cells. Pancreatic islets are very different from this preparation, as they are highly heterogeneous, whereas these cells all respond equally. It would be good to see the same technique applied to primary cells.

      MIN6 cells do not respond to glucose and other secretagogues in the same way as primary cells, and they cycle, depending on the phase of the cycle to which they are exposed.

      The authors should report the number of cells per spheroid and the number of cells that are alive and dead.

      I would like to examine the effects of calcium channel blockers on calcium transients, and the use of pregnenolone is already described in the literature, but remains less well known.

      MIN6 cells secrete much insulin, because detecting the hormone in ELISAs requires too many primary cells. The authors should discuss the model in greater detail and compare it with primary beta cells. Also, they take 3 mM glucose as the basal concentration, which is low.

      We thank the reviewer for their valuable comments, which have helped us to significantly enhance the quality of our manuscript. We have carefully considered these comments and have revised the manuscript accordingly. A point-by-point rebuttal is provided below.

      Reviewer #1 (Recommendations for the authors):

      (1) The manuscript contains numerous typos, including the combination of numbers and units without a space.

      All spacing inconsistencies between numerical values and unit symbols (e.g. mM, μM, µL, and Hz) have been corrected throughout the text and images. In addition, the following typographical and grammatical errors have been addressed:

      - Abstract: "the frequency of Ca <sup>2+</sup> oscillations correlate" → correlates

      - Figure 2A caption: "200uL" → 200 µL

      - Figure 5A caption: "glimepirde" → glimepiride

      - Figure 5 title: "KATP-antagonists induces" → induce

      - Figure 7D caption: "concentrations" → concentration

      - Figure 7 caption: "Tukey’s post-hoc testm" → Tukey’s post-hoc test

      - Results section header: "increases insulin secretions" → insulin secretion

      - Discussion: "we obtained an EC50 values 7.4 ± 0.4 mM" → "we obtained EC50 values of 7.4 ± 0.4 mM"

      - Discussion: "a EC50 value" → an EC50 value

      - Discussion: "concentration dependence profiles that matches" → match

      - Discussion: "PS-induced activation TRPM3" → activation of TRPM3

      - Conclusion: "found the Islets of Langerhans" → found in the Islets of Langerhans

      - Acknowledgements: "grant agreement No. 955643" → agreement No. 955643

      (2) I suggest showing the experiments in primary beta cells because of the many differences from insulinoma cells, as the most important result is a better culture technique for calcium imaging.

      We thank the reviewer for this valuable suggestion. We did indeed attempt to perform comparable experiments using the Cellartis<sup>®</sup> hiPS Beta Cell Media Kit (Takara, cat. no. Y10108) as a more physiologically relevant cell model. To minimize cellular stress during the transition, we adjusted our protocol by seeding the differentiated cells into the micropatterned plates 24 hours prior to imaging, thereby maintaining the recommended culture conditions for as long as feasible. However, during calcium imaging, the cells failed to respond to either the elevated glucose stimulus or the positive control, suggesting that the cells did not survive the transfer to our plate format with sufficient viability to mount a functional response.

      We acknowledge that the use of primary beta cells or hiPS-derived beta cells would strengthen the physiological relevance of the platform. Nevertheless, we would like to emphasize that the primary objective of this study is to demonstrate the feasibility of high-throughput calcium imaging in 3D cell culture using our optimized protocol: a proof-of-concept that is, by design, independent of the specific cell model employed. Adapting the protocol to accommodate more sensitive or terminally differentiated cell types is a meaningful avenue for future work, but falls outside the scope of the current manuscript. We have added a brief note to the Discussion section to explicitly acknowledge this limitation and to identify hiPS-derived beta cell compatibility as a priority for subsequent optimization.

      (3) These kinds of cultures in three dimensions are interesting, but it has been shown that it is even better to have the liquid flow, simulating blood flow; this can at least be discussed.

      We thank the reviewer for this insightful comment. We agree that the introduction of perfusion-based flow represents a meaningful improvement over static 3D culture systems. We have incorporated this point into the Discussion section, where we now explicitly acknowledge the potential benefits of dynamic culture conditions, including improved viability, insulin secretory function, and morphological integrity of 3D β-cell tissues, and identify the integration of perfusion-induced flow as a potentially meaningful path to explore for future development of the platform.

      Reviewer #2 (Public review):

      Summary:

      The study by Robben et al., show 3D beta-cell spheroid platform, a valuable tool allowing high-throughput monitoring of cytoplasmic Ca concentrations and insulin secretion, with Ca signals comparable to those recorded in primary islets. The authors demonstrate a solid method to culturing MIN6 cells in a 3D culture system, recording Ca signals in a high-throughput format and characterizing these Ca signals using pharmacological tools, including TRPM3 channel and K-ATP channel modulators. This highlights the utility of the 3D beta-cell spheroid for screening new ion channel modulators in beta-cells of the pancreas.

      Strengths:

      - The study shows that the MIN-6-based 3D beta-cell model is better to study Ca-signaling and insulin secretion compared to 2D culture of single MIN-6 cells.

      - The method allows imaging of Ca signaling in many spheroids in parallel followed by collecting medium to measure insulin release and correlate both effects.

      - The authors demonstrate that this system is suitable for screening new pharmacological modulators and used as an agonist of the ATP-sensitive potassium channel (diazoxide) and the agonist and antagonist of the TRPM3 channel.

      Weaknesses:

      - The study is based on only one cell line, the MIN6 insulinoma cells, which may not fully mimic the pancreatic beta-cells within the islet.

      - The authors show only spheroids cultured overnight. A long-term culture is missing to assess beta-cell viability long term function.

      - The authors tested their platform using only two compounds. Testing a larger compound library is necessary to make a clear conclusion about the suitability of the platform for high-throughput screening.

      We thank the reviewer for their valuable comments, which have helped us to significantly enhance the quality of our manuscript. We have carefully considered these comments and have revised the manuscript accordingly. A point-by-point rebuttal is provided below.

      Reviewer #2 (Recommendations for the authors):

      Major Points

      (1) In this study, only 2 pharmacological compounds (for TRPM3, and K channel) were tested. Testing a larger compound library would be necessary to fully demonstrate its suitability for high-throughput screening applications. If this is not possible at the moment, including data on additional pharmacological compounds, e.g., modulators of voltage-gated Ca channels, which are key regulators of Ca signaling in beta cells of the pancreas would strengthen the study. (e.g., use voltage gated Ca channel blocker such as verapamil or nimodipine).

      We thank the reviewer for this constructive suggestion. We would like to clarify that our pharmacological characterization was not limited to two compounds. In total, six compounds spanning two distinct ion channel targets were evaluated: the K-ATP channel modulators diazoxide, glimepiride, tolbutamide and nateglinide, and the TRPM3 modulators pregnenolone sulphate and isosakuranetin. This panel includes both agonists and antagonists across two mechanistically distinct targets, and we believe this is sufficient to demonstrate the platform's suitability for high-throughput compound screening in the context of a proof-of-concept study.

      We nonetheless agree with the reviewer that extending the compound panel to include modulators of additional ion channel classes, such as voltage-gated Ca <sup>2+</sup> channel blockers like verapamil or nimodipine, would further demonstrate the versatility of the platform. We have added a statement to the Discussion explicitly identifying this as a valuable direction for future work.

      (2) Testing another beta-cell line (e.g., INS-1 cells) would strengthen the manuscript.

      We thank the reviewer for this suggestion. We agree that validating the platform using an additional β-cell line, such as INS-1 cells, would further broaden the applicability of the approach. We did indeed attempt experiments with INS-1 cells; however, the results were inconclusive due to cell quality issues at the time of testing, and we were unable to generate reliable data suitable for inclusion in the manuscript.

      We would also like to emphasize that the primary aim of this study was to demonstrate the methodology of the high-throughput screening platform, rather than to provide a comprehensive cross-cell-line validation. In this context, the use of the well-established MIN6 β-cell line is sufficient to serve as a proof-of-principle demonstration of the platform's capabilities.

      Nonetheless, we consider a systematic evaluation of INS-1 cells on this platform as an important and natural next step and have included this explicitly as a future perspective in the Discussion.

      (3) I find the presentation of the results and analysis of Ca oscillation frequency and area under the curve excellent. Could you please provide more details on the analysis method used to quantify the frequency of glucose-induced Ca oscillation. If a custom script was used, sharing this information with the scientific community would be great.

      We thank the reviewer for this positive feedback. We confirm that the Ca <sup>2+</sup> oscillation analysis was performed using a custom Python script (v3.10.11), the key steps of which are described in the Data Analysis section of the experimental procedures.

      In line with our commitment to open and reproducible science, the script will be made publicly available upon acceptance of the manuscript, allowing the broader scientific community to apply, adapt, and build upon the analysis pipeline.

      (4) Please move the Supplementary Figure to the main Figure 3. This will allow a direct comparison between Ca signals in spheroids and in single cell (2D cultures) under identical conditions.

      We thank the reviewer for this suggestion. We have partially incorporated the supplementary figure into the main manuscript. The mean Ca <sup>2+</sup> response of 2D-cultured MIN6 cells at 20 mM glucose has been added as Figure 3D, enabling direct visual comparison with the spheroid data under identical stimulation conditions. The remaining panels of the supplementary figure — showing representative Ca <sup>2+</sup> traces across multiple glucose concentrations and the corresponding dose-response curves for peak frequency and area under the peaks in 2D monolayers — have been retained in the supplementary information, as their inclusion in the main figure would substantially increase its complexity. The figure legend has been updated accordingly.

      (5) A direct comparison of insulin secretion between 3D cultured spheroids and 2D cultures should also be shown.

      We thank the reviewer for this suggestion. We attempted to include a direct comparison of insulin secretion between 3D spheroids and 2D MIN6 monolayers; however, the 2D measurements proved unreliable for quantitative comparison. Insulin values in the 2D condition consistently exceeded the upper detection limit of the ELISA, and inter-well variability was too high to draw meaningful conclusions. We therefore chose not to include this comparison and instead present the Ca <sup>2+</sup> imaging data in Figure 3D as a functional readout enabling direct comparison between the two culture formats under identical stimulation conditions.

      (6) The authors should further discuss the remaining effects of pregnenolone sulphate on insulin secretion.

      We thank the reviewer for this comment. We would like to clarify that in our experimental setup, pregnenolone sulphate and isosakuranetin were applied simultaneously rather than sequentially. As a result, the incomplete inhibition of PS-induced Ca <sup>2+</sup> oscillations and insulin secretion observed in the presence of isosakuranetin may in part reflect a kinetic offset between the two compounds, whereby PS-induced TRPM3 activation and downstream signalling may have been initiated prior to the establishment of effective TRPM3 blockade by isosakuranetin. In addition, as noted in the Discussion, TRPM3 may not be the only molecular target of PS in these spheroids, and alternative signalling pathways may contribute to the residual insulin secretion observed in the presence of the antagonist. We have added a brief clarification to the Discussion to explicitly acknowledge the potential influence of this kinetic limitation on the interpretation of these results.

      (7) Is it possible to collect 3D cultured spheroids after each experiment to measure for example intracellular insulin content or protein levels by Western blot.

      Physical recovery of spheroids from the PEG hydrogel plates for downstream biochemical analysis, such as intracellular insulin content measurements or Western blot, would indeed be a meaningful addition to the platform's capabilities. We would like to note that the firm attachment of spheroids to the glass substrate, while essential for maintaining spheroid positioning during the extensive washing and liquid handling steps, does present a practical challenge for post-experimental recovery. Although we have successfully extracted spheroids of other cell types from comparable plate formats, reliable recovery of the MIN6 β-cell spheroids without compromising their structural integrity has not yet been achieved. We therefore identify the optimization of spheroid recovery as a valuable direction for future development of the platform.

      (8) Ca signals in response to the application of glucose appears more robust in spheroids compared to single MIN6 cells. What are the possible mechanisms underlying this difference. Whole RNA-seq experiments would be one approach to identify differentially expressed genes in 2D versus 3D culture (this is maybe a whole project by itself). An alternative is to look by RT-qPCR analysis for key β-cell markers and genes encoding ion channel and ion channel subunits.

      We thank the reviewer for this thoughtful comment. The more robust Ca <sup>2+</sup> signals observed in 3D spheroids compared to 2D monolayer cultures likely reflect several interconnected factors. First, and importantly, it should be noted that Ca <sup>2+</sup> measurements in 2D monolayer cultures typically represent an averaged signal across a large population of cells, which tends to obscure individual oscillatory events and reduce the apparent amplitude and regularity of Ca <sup>2+</sup> responses. In contrast, our 3D spheroid platform enables Ca <sup>2+</sup> measurements at the level of individual spheroids, allowing discrete oscillatory peaks to be resolved with much greater fidelity. Beyond this methodological distinction, the 3D architecture also promotes enhanced cell-to-cell communication, better recapitulation of in vivo β-cell coupling, and a more physiologically relevant microenvironment, all of which are likely to contribute to the improved oscillatory Ca <sup>2+</sup> dynamics observed.

      We agree with the reviewer that elucidating the transcriptional underpinnings of these differences, through whole RNA-seq or targeted RT-qPCR analysis of key β-cell markers and genes encoding ion channel subunits, would be highly informative and represents an elegant approach to understanding the molecular basis of the observed functional improvements. As the reviewer rightly acknowledges, however, such experiments constitute a substantial research effort in their own right. We have added a statement to the Discussion identifying this as a valuable direction for future investigation, alongside the other platform development priorities already outlined.

      (9) A more detailed discussion of the limitations of the platform and potential strategies to further improve this system would strengthen the manuscript.

      We thank the reviewer for this constructive suggestion. In response, we have expanded the Discussion to provide a more comprehensive overview of the current limitations of the platform and the strategies we envision for future development. Specifically, the revised Discussion now addresses the following points:

      First, the platform is currently optimized for the MIN6 insulinoma cell line, which differs from primary pancreatic beta cells in several important respects, including glucose sensitivity and secretory capacity. Initial attempts to adapt the protocol to hiPS-derived beta cells were unsuccessful, likely due to insufficient cellular viability following transfer to the micropatterned plate format. Optimizing the platform for use with primary beta cells or hiPS-derived beta cells is therefore identified as a priority for future development.

      Second, while the current study demonstrates proof-of-concept pharmacological characterization using six compounds across two mechanistically distinct ion channel targets, extending the compound panel to include modulators of additional ion channel classes, such as voltage-gated Ca <sup>2+</sup> channel blockers, as well as validation using alternative insulinoma cell lines such as INS-1, would further demonstrate the versatility and generalizability of the platform.

      Third, the introduction of perfusion-induced flow, which has been shown to improve viability, insulin secretory function, and morphological integrity of 3D beta-cell tissues under dynamic culture conditions, is identified as an additional avenue for optimization.

      Fourth, while the current platform operates in a 96-well plate format, which already provides a substantial throughput of up to 1824 individual spheroid measurements per plate, adaptation to higher density plate formats such as 384-well or 1536-well plates would be a necessary step towards true high-throughput screening compatible with industrial drug discovery pipelines. Miniaturization of the hydrogel design and adaptation of the molding procedure to accommodate these formats therefore represents an important direction for future development.

      Minor points:

      (10) Please clarify the glucose concentration at which the MIN6 cells the cultured 3D beta cell spheroids were maintained overnight prior to the experiments.

      Prior to the experiments, both 2D MIN6 cells and 3D MIN6 β-cell spheroids were maintained overnight in standard high-glucose DMEM (25 mM glucose), consistent with widely used MIN6 culture protocols. This information has been added to the Methods section.

      (11) In Figure 3A, is the presented Ca trace derived from one single spheroid? Showing representative Ca traces (e.g., 5 i traces per condition) would show the reproducibility of the recordings.

      We thank the reviewer for this suggestion. The trace shown in Figure 3A is indeed derived from a single representative spheroid. To address the concern regarding reproducibility, we refer the reviewer to the updated Figure 3C, which now displays the individual Ca <sup>2+</sup> response traces of all spheroids within a single well in response to 20 mM glucose stimulation. Grey lines represent individual spheroids, while the black line denotes the mean response. This panel illustrates not only the reproducibility of the oscillatory response across the imaged population but also highlights an important consequence of inter-spheroid variability: because individual spheroids oscillate asynchronously, their peaks cancel out when averaged, resulting in a mean trace that appears relatively flat and lacks the oscillatory features visible in individual recordings. We note that this same cancellation effect likely underlies the comparably flat mean response observed for 2D-cultured MIN6 cells in Figure 3D. Additionally, the difference in the initial response profile between 2D and 3D cultures may in part reflect the geometry of the hydrogel microenvironment. In 2D cultures, the glucose stimulus equilibrates rapidly and near-uniformly across the culture plane, potentially driving a more synchronized initial response and the early peak visible in Figure 3D. In contrast, the PEG hydrogel surrounding the 3D spheroids may act as a diffusion barrier, causing the stimulus to reach individual spheroids with variable delay and thereby further desynchronizing response onsets across the population. The figure legend has been updated accordingly.

      (12) In Figure 3A, the potassium application bar looks a bit shifted. Please check and correct if necessary.

      We thank the reviewer for carefully examining the figure. The apparent shift in the potassium application bar is not an error. In these experiments, glucose was added after an initial 10-minute baseline period, and the observed delay in the Ca <sup>2+</sup> response reflects two contributing factors. First, the glucose solution was pipetted at the top of the well, and diffusion to the level of the spheroids introduces a short lag before the stimulus reaches the cells. Second, the spheroid shown in Figure 3A was located at the outer edge of the well, where mixing is slower and the stimulus arrives with additional delay compared to centrally positioned spheroids. Together, these factors account for the offset between the start of the glucose application bar and the onset of the visible Ca <sup>2+</sup> response.

      (13) Is the system also compatible with the ratiometric Ca imaging dye Fura-2?

      The platform is indeed compatible with ratiometric Ca <sup>2+</sup> imaging using Fura-2. The glass-bottom plate format is a prerequisite for Fura-2 imaging due to the requirement for UV excitation at 340/380 nm, and the transparency of the PEG-based hydrogel ensures that the optical properties of the platform are fully compatible with this approach. Furthermore, the µCELL FDSS fluorescence plate imager used in this study supports dual-excitation ratiometric imaging, making it instrumentally compatible with Fura-2 without any additional hardware modifications.

      (14) In Figure 4 legend: 100 µm diazoxide should be corrected to 100 µM diazoxide.

      This has been addressed in the revised manuscript

      (15) Please comment on the cost of spheroid generation compared with conventional 2D MIN6 cultures.

      We thank the reviewer for this relevant question. The cost of spheroid generation using our platform is largely comparable to conventional 2D MIN6 cultures, with the primary additional expense being the specialized micropatterned PEG-based hydrogel plates. All other aspects of the workflow, including cell culture reagents, imaging consumables, and instrumentation, remain identical. It should be noted that providing a precise cost comparison is difficult, as the price of the hydrogel plates represents the dominant variable cost and is subject to change depending on production scale and supplier agreements. Nevertheless, we consider the platform to be cost-effective relative to alternative 3D culture systems, which often require more complex fabrication procedures or proprietary consumables.

      Reviewer #3 (Public review):

      Summary:

      The primary objective of this study is to develop high-throughput screening assays utilizing homogeneous 3D cell cultures that more accurately replicate the intricate architecture and cellular communication found in tissues. The authors have chosen pancreatic islet β-cells as a model system to evaluate agents that modulate insulin release, which is particularly relevant given the increasing prevalence of diabetes mellitus-a significant global health concern. Moreover, the incorporation of human-based 3D spheroids, organoids, or organ-on-chip technologies into drug discovery protocols is essential for enhancing clinical translation, as candidate compounds identified using animal models have often demonstrated limited success in clinical settings.

      Strengths:

      This study was thoughtfully planned and skillfully carried out. The use of micropatterned hydrogels to observe 19 spheroids at once is an ingenious aspect, which has been effectively validated with Ca microfluorography. Overall, I found this investigation to be exceptionally well-executed and free from notable flaws, as the results clearly back up the conclusions. Additionally, the developed method achieved the proposed aims, providing a high-throughput format with 3D cultures. I believe this study deserves publication.

      Weaknesses:

      For an HTS assay, authors should incorporate the Z-factor.

      We thank the reviewer for their valuable comment, which we have directly addressed in the revised version, as outlined below.

      Reviewer #3 (Recommendations for the authors):

      (1) The study is very well performed, but to support the claim of suitability for HTS, the Z-factor of the method should be reported.

      We thank the reviewer for this helpful suggestion. In response, we have now included the Z′-factor in the manuscript to explicitly quantify assay performance and robustness. Specifically, the Z′-factor (0.6) has been added and discussed in the Results section, incorporated into the Discussion to contextualize assay suitability for high-throughput screening, and included in the Methods under “Statistical Analysis,” where its calculation is described.

    1. eLife Assessment

      This potentially valuable work finds that exposure to the odor diacetyl in the worm C. elegans leads to changes in metabolic nutrient-responsive programs, inducing expression of genes that act in the DHAP-glycerol shunt followed by induction of the flavin-containing monooxygenase fmo-2. The genetic and multi-omic evidence is solid and is based on an extensive set of experiments; however, mechanistic support for underlying metabolic and redox changes remains incomplete without more comprehensive metabolic profiling. The mechanism through which diacetyl or its breakdown products is sensed did not involve known receptors of diacetyl, so it remains unknown whether sensory neurons are involved in these responses.

    2. Reviewer #1 (Public review):

      Short overview:

      This work demonstrates that exposure to the odor diacetyl in C. elegans first induces the expression of genes that act in the DHAP (dihydroxyacetone phosphate) - glycerol shunt, followed by induction of the flavin-containing monooxygenase fmo-2. This diacetyl exposure also increases the survival of starved animals. However, the mechanism through which diacetyl is sensed by the animal does not involve any of the known receptors of diacetyl, and it is unknown whether any sensory neurons are involved in these responses.

      Summary:

      This manuscript shows that diacetyl exposure induces metabolic remodeling in C. elegans, where genes in the DHAP (dihydroxyacetone phosphate) - glycerol shunt pathway are activated. This leads to a secondary activation of the flavin-containing monooxygenase fmo-2, which is a longevity-promoting gene upon dietary restriction. The activation of fmo-2 is consistent with the authors' observations that diacetyl exposure increases the survival of starved animals, but not of fed animals. The authors also show that diacetyl exposure improves the animals' resistance to hyperosmotic stress, which is consistent with an increase in glycerol and glycerolipids in these animals. Together, the authors show that a volatile odor or odors can induce gene expression changes that promote metabolic changes and survival under certain environmental conditions.

      Strengths:

      Through transcriptomics, metabolomics and lipidomics of wild type, with or without diacetyl, and phenotypic analyses of mutant or RNAi-treated animals, the authors delineate a mechanism through which an odor or odors remodel metabolism and increase survival. They show that genes of the DHAP-glycerol shunt are activated during diacetyl exposure, and that these in turn activate fmo-2, which is known to promote longevity under different stressors, like food deprivation. The authors also show that an upstream transcription factor, the co-activator MDT-15, is required for the diacetyl-mediated initial activation of the DHAP-glycerol shunt, and thereby of fmo-2. The involvement of MDT-15 and, to a lesser extent, one of its partner nuclear hormone receptors (NHRs), NHR-49, might also explain the lipidomic changes that accompany diacetyl exposure.

      Weaknesses:

      The authors tested the involvement of the ODR-10 receptor, which is the known receptor for attractive concentrations of diacetyl, the same concentrations used in this study. Surprisingly, ODR-10 is not required for any of the DHAP-glycerol shunt or fmo-2 induction. In addition, they found that loss of sensory cilia (in the background mutation of the NHR daf-12) enhances diacetyl-mediated fmo-2 induction, which suggests that wild-type sensory cilia inhibit fmo-2 and likely synergizes with the DAF-12 receptor. Considering that chemosensory receptors, like ODR-10, are normally localized to the ciliary endings, their observations suggest a different mechanism through which the animals process this odor. There are at least two possibilities.

      One, diacetyl might affect the animals by permeating their cuticles. However, the authors did not test if any neurons are involved in their phenotypes. The AWA neuron, where ODR-10 is expressed, is required to sense attractive concentrations of diacetyl. Thus, what happens when the AWA neurons are ablated or lost?

      Two, diacetyl is both light- and temperature-sensitive and can break down into two other odors, acetoin and 2, 3-butanediol. Acetoin is sensed by the chemoattractive AWA neurons (Siddiqui et al, eLife 2024, 101936.1), but not by ODR-10 (Zhang et al., PNAS 1997, vol 94, pp 12162-12167). Thus, this odor would have a different receptor. Although C. elegans chemosensory receptors have been found within the cilia, there is a formal possibility that some of the sensory receptors will also be found on other parts of some sensory neurons, like the dendrites, as in Drosophila (Joseph and Carlson, Trends Genet 2016, vol 31, pp 683-695). Regarding the other diacetyl derivative, 2, 3-butanediol, it has not been shown to elicit any chemosensory responses in C. elegans (Siddiqui et al. eLife 2024, vol 13, RP101936), but 2, 3-butanediol has been linked to the microbiome of fmo knockout mice (Said et al., Metabolomics 2025, vol 21, 170). Thus, it is possible that it is the breakdown products of diacetyl that elicit the responses the authors see.

      Finally, the authors' data contradict the observations made by Park et al (Aging Cell 2021, vol 20, e13300). Park et al previously showed that diacetyl exposure reduces the survival of food-deprived animals, although this phenotype is also independent of ODR-10 or of the SRI-14 receptor for aversive concentrations of diacetyl.

    3. Reviewer #2 (Public review):

      Short overview:

      This study presents potentially important findings showing that DHAP-glycerol shunt involved in energy balance is regulated by food availability in a widely used C. elegans model. The genetic evidence supporting this conclusion is solid and is based on an extensive set of experiments; however, key metabolic measurements and comprehensive metabolic profiling are not provided, limiting the strength of the conclusions about the underlying metabolic and redox changes.

      Comments:

      Giorda and colleagues report interesting findings demonstrating that the DHAP-Gro3P shuttle is modulated by food availability in C. elegans. Although the authors provide multiple interesting observations in worms, supported by an extensive number of experiments, the metabolic aspect of the study requires additional development. It appears that targeted lipidomics and metabolomics analyses were performed, but the corresponding datasets are largely absent from the manuscript. Only a very limited subset of lipid species is presented in Fig. 2D. What about triglycerides? It would be highly informative to include comprehensive lipidomic profiles covering major lipid classes. A similar concern applies to the metabolomics data. Where are the measurements of Gro3P, DHAP, and glycerol? The authors state that their LC-MS method was unsuccessful and that glycerol levels were ultimately measured using a commercial kit. Given that glycerol production and excretion appear to be major output across many of the experiments presented, this approach is not entirely satisfactory. Reliable GC-MS based methods are available for the quantification of all major components of this pathway, including Gro3P, DHAP, and glycerol (derivatization helps to preserve these species, especially glycerol).

      Furthermore, comprehensive LC-MS/GC-MS-based metabolic profiling should be included. Metabolites reported and organized by pathway (e.g., glycolysis, TCA cycle, pentose phosphate pathway) would provide a broader understanding of the metabolic consequences of DHAP-glycerol shunt activation.

      Finally, because the DHAP-glycerol shunt is closely linked to cellular redox homeostasis, it would be important to determine how its activation affects intracellular pyridine nucleotide pools, and measurements of NAD+, NADH, NADPH, NADP+ would substantially strengthen the mechanistic conclusions and provide direct evidence for alterations in cellular redox state.

    4. Reviewer #3 (Public review):

      Short overview:

      This study asks whether the perception of a volatile and attractive cue, diacetyl, leads to changes in metabolic nutrient-responsive programs in C. elegans. Using multi-omic and genetic evidence, they connect the transcriptional response to diacetyl to early activation of the DHAP-glycerol shunt, suggesting worms may activate this metabolic pathway in preparation for food intake. This work also identifies transcription factors involved, and while it does not test whether this response is diacetyl-specific, could identify a conserved pathway of food intake preparation.

      Summary:

      This study asks whether C. elegans can use a volatile food cue alone, in the absence of ingestion, to anticipate nutrient availability. Using the attractive odorant diacetyl, the authors show that fasting worms rapidly induce the DHAP-Glycerol shunt, a metabolic pathway normally associated with glucotoxicity and hyperosmotic stress, and that this response depends on the transcription factor MDT-15. This drives measurable metabolic rewiring (glycerol and phosphatidylglycerol accumulation) and confers protection against subsequent hyperosmotic stress. Prolonged diacetyl exposure further triggers a second, HLH-30-dependent wave of fmo-2 expression linked to depleted NTP levels, resulting in enhanced heat tolerance and food-seeking behavior upon refeeding. Together, the authors build a testable model for a pathway linking sensory cue detection to gene expression, metabolism, and physiological changes.

      Strengths:

      The central finding that smell alone without ingestion activates a metabolic-stress adaptation pathway is highly interesting and supported by a convergence of methods including RNA-seq, transcriptional reporters, metabolomics, and functional behavior assays. The transcriptomic time course distinguishes two temporally separable gene expression waves (the shunt at 30 min and fmo-2 at 90 min), and epistasis experiments comparing food status and osmotic stress (Fig. 1i-j) convincingly argue diacetyl acts as a food-predictive cue rather than mimicking osmotic stress. A particular strength is the test of the relationship between the two pathways identified in the study. RNAi knockdown of DHAP-Glycerol shunt enzymes block fmo-2 induction, and pretreatment with salt then switching to food deprivation shows that it is the shunt's activation state that drives fmo-2 expression. This is further reinforced by depletion of energetic mechanism (NTP/ATP). Figure 5 extends this upstream to MDT-15 validated through gene expression, metabolomics, and functional readouts. Throughout the study, transcriptional and metabolic findings are paired with functional outcomes such as hyperosmotic protection, heat stress resistance, or survival, strengthening the paper's model. Together, the authors largely achieve their aim of establishing that a food-related olfactory cue is sufficient to trigger an anticipatory metabolic and transcriptional program. The core claim that diacetyl sensing activates the shunt and subsequent fmo-2 induction and physiological protection is well supported by convergent genetic and biochemical experiments.

      Weaknesses:

      The authors show that diacetyl-induced fmo-2 induction does not require canonical olfactory sensing since neither diacetyl receptor mutants (odr-10, sri-14) nor a cilia-deficient strain (daf-19; daf-12) blocked the response. However, the pathway characterized here is defined almost entirely through diacetyl, and it remains unclear whether the anticipatory response reflects general food-predictive olfaction or a diacetyl-specific effect. This distinction is important because the authors' central hypothesis is whether volatile food cues in general can be used to anticipate nutrient availability. Testing a limited number of additional attractive odorants, ideally sensed through distinct chemosensory receptors, for a limited number of phenotypes, would establish whether this pathway is generalized to food-predictive smells or only diacetyl. This would increase the impact of the work. Identifying the mechanism(s) for diacetyl perception would as well, but is much more challenging and less likely to be feasible in this work.

      The RNA-seq and reporter data disagree on when fmo-2 induction begins, and thus do not fully support that there are two temporally distinct waves. RNA-seq (sampled at 5, 15, 30, and 90 min) shows fmo-2 is not a significantly differentially expressed gene at 30 min and only reaches significance at 90 min, while the fmo-2 transcriptional reporter (Fig.S3a) shows detectable induction as early as 30 minutes. Other readouts in the paper use the reporter for fmo-2 but qPCR for shunt genes, further muddling the timing since mRNA should precede reporter visualization. Since reporter signal generally lags transcript level detection, this discrepancy could be further clarified through a time-course qPCR (like what was tested for the shunt genes) between 30 and 90 minutes. This would help establish whether the two waves are separated by time or whether this reflects a difference in detection method.

      Minor weaknesses:

      In Figure 1i, gpdh-1 induction is compared between control and 200 mM NaCl (3 hr exposure), with food present or absent. While this directly tests food status, a 3-hour exposure approaches the ~5-6-hour window previously shown to be sufficient for salt-food associative learning to form (worms move toward the high-salt side of a gradient plate if pretreated with high salt and food). This raises the possibility that, in the food-present condition, part of the measured gpdh-1 response could reflect an emerging learned association between salt and food forming during the assay itself, rather than purely reflecting an interaction between nutrient status and osmotic stress signaling. Testing a shorter exposure window (e.g., under 1 hour, as used for the diacetyl exposure in panel j) would help clarify this.

      Thrashing is used throughout the paper as the primary readout for hyperosmotic stress resistance, including the central result that diacetyl protects against subsequent hyperosmotic stress (Fig. 2g). However, the authors don't explain why this was chosen as the main stress-resistance readout.

      nhr-49 knockdown reduces gpdh-1 induction in diacetyl-mediated hyperosmotic protection but has no effect on development or survival on sustained hyperosmotic stress, raising the question of how the role of nhr-49 may be distinct from mdt-15 in this context.

    5. Author response:

      We thank the Reviewing Editor and the reviewers for their highly constructive feedback and their positive assessment of our study. We plan to submit a revised manuscript that addresses these critiques through textual revisions, contextualization of our data, and explicit discussion of the study's limitations.

      To address the comments from Reviewers 1 and 3 regarding the sensing mechanism and cue specificity, we agree that the precise sensor for diacetyl remains an open question. While we cannot specifically pinpoint the mechanism with our current data, we hypothesize that this response relies on a non-canonical sensing mechanism, given our multiple negative results for canonical diacetyl receptors (odr-10, sri-14), signalling pathways, and cilia-defective mutants (daf-19; daf-12). We will revise the text to suggest the mechanism could be cilia-independent or cell-autonomous. Additionally, we will explicitly frame the investigation of AWA-ablated worms, diacetyl derivatives such as acetoin, and additional volatile food cues as important future directions to establish the generalizability of this response.

      Regarding the contrasting survival phenotypes observed by Park et al.24 raised by Reviewer 1, our manuscript currently discusses how chronic odour exposure represses the longevity benefits of dietary restriction48, hypothesizing that this may stem from age-dependent olfactory decline49 and the confounding variable of olfactory learning8. To make this connection clearer, we will explicitly cite Park et al. in this section to directly link their findings with our hypothesis that repeated odour exposure without a nutritional reward extinguishes its efficacy as an anticipatory cue.

      In response to Reviewer 2’s feedback on our metabolic profiling, we will ensure it is clear in the text that we highlighted the lipid species most prominently affected by diacetyl, explicitly noting that triglycerides were not significantly altered. We will also better direct readers to Supplementary Table 1, which contains the comprehensive lists of all metabolites and lipids detected in our metabolomics and lipidomics, and quantifies how they are affected by the treatments in our study. As noted in our manuscript, both DHAP and Gro3P were successfully detected in our metabolomics platform but were not significantly altered by diacetyl exposure. Glycerol, however, was measured via a commercial enzymatic assay because it was not detected by our specific LC-MS platform. We will acknowledge that independently measuring the effects on cellular redox states and utilizing GC-MS for broader metabolic profiling are valuable future directions.

      Finally, to address Reviewer 3’s queries regarding experimental readouts and nhr-49, we will clarify our rationale for using thrashing as our primary readout for hyperosmotic stress. Because diacetyl exposure triggers an acute induction of the DHAP-glycerol shunt, we specifically chose an acute behavioural readout to match this rapid timeline. Regarding nhr-49, we will add a new point to the discussion proposing that the distinct phenotypes may come down to expression thresholds. We hypothesize that while nhr-49 RNAi partially reduces gpdh-1 expression, this residual level of expression might still be sufficient to allow development and survival during sustained hyperosmotic stress."

      References:

      (8) Choi, J. I., Yoon, K., Kalichamy, S. S., Yoon, S.-S. & Lee, J. I. A natural odor attraction between lactic acid bacteria and the nematode Caenorhabditis elegans. ISME J. 10, 558–567 (2016).

      (24) Park, S. et al. Diacetyl odor shortens longevity conferred by food deprivation in C. elegans via downregulation of DAF‐16/FOXO. Aging Cell 20, e13300 (2021).

      (48) Zhang, B., Jun, H., Wu, J., Liu, J. & Xu, X. Z. S. Olfactory perception of food abundance regulates dietary restriction-mediated longevity via a brain-to-gut signal. Nat. Aging 1, 255–268 (2021).

      (49) Suryawinata, N. et al. Dietary E. coli promotes age-dependent chemotaxis decline in C. elegans. Sci. Rep. 14, 5529 (2024).

    1. eLife Assessment

      This valuable study examines how abstraction and metacognition relate to transdiagnostic dimensions of psychopathology. It combines a reward-learning task, computational modelling and a large questionnaire. The evidence linking reduced metacognitive sensitivity to a compulsive hypersensitivity dimension is solid, but the evidence for the abstraction findings is currently incomplete. This work will be of interest to researchers studying metacognition and computational approaches to psychiatry.

    2. Reviewer #1 (Public review):

      Summary:

      This work investigated the associations between abstraction and metacognition in the context of reward-guided learning and transdiagnostic symptom dimensions in a sample of N = 249 participants. Participants completed a reward-guided learning task and confidence judgements. Transdiagnostic dimensions used to examine associations with abstraction and metacognition were based on a large existing dataset. Findings showed that in the examined sample, a Compulsive Hypersensitivity dimension was negatively associated with abstraction and metacognitive sensitivity, while a Social Withdrawal dimension was positively associated with metacognitive sensitivity. These data add to the existing literature on associations between metacognition and transdiagnostic symptom dimensions and extend previous work on the association between abstraction and these dimensions.

      Strengths:

      (1) The study addresses an interesting research question and uses a transdiagnostic approach. While part of the research question is a replication (metacognition), the additional inclusion of an abstraction parameter is highly valuable.

      (2) Methodologically, the study is strong. Specifically, the implementation of an experimental task to estimate parameters, the hierarchical Bayesian modelling, the parameter recovery analyses, the bootstrap regression (including corrections for multiple comparisons) and the control of relevant covariates and response tendencies are quite impressive.

      (3) The Open Science approach is laudable in general. The study was preregistered and provides open data and open code. Deviations from the preregistration are transparently reported (e.g., bootstrap regression, exclusion criteria). In this vein, the high number of robustness analyses provided in the supplements is very much appreciated.

      (4) More generally, the extensiveness of the supplements is particularly valuable.

      (5) Finally, it is very useful that the discussion takes into account potential alternative explanations of the findings.

      Weaknesses:

      (1) High number of exclusions:

      As the authors mention (also in their limitations section), more than half of the participants had to be excluded (only 249 out of 512 participants remained), which is substantial. Specifically, 203 participants were excluded because they failed attention checks, 58 failed comprehension questions on the confidence scale, 25 had a reading time of the instruction page below 5 seconds, and 68 showed a performance that was too poor (the sum is probably higher than 263 because these overlap). This high number of excluded cases resulted in a substantially smaller analysed sample than planned (corresponding to approximately 77% power instead of 90%) and potentially limited both statistical power and generalisability. It might even be the case that highly impulsive individuals were excluded systematically. That is, the exclusion criteria may themselves be associated with psychiatric symptoms, so the analysed sample may no longer (fully) represent the target population. It would be interesting to see whether the results also hold when including the excluded participants (because excluded and included participants differ in attentional and impulsivity-related symptoms). A sensitivity analysis would be beneficial.

      (2) Deviations from preregistration:

      While the preregistration of the study is positive in general, some aspects raise doubts here. First, there have been quite a few deviations from the preregistration. For instance, the primary outcome variable for abstraction has been changed (initially: proportion of blocks in which the Abstract RL model had a better fit than the Feature RL model; instead: mean posterior responsibility of the Abstract RL model across blocks), and this appears to have influenced the findings (e.g., Supplementary Figures: 8 vs. 9). Generally, these deviations introduce additional researcher degrees of freedom and therefore warrant a more detailed justification (or replication in future studies). Second, it appears that the preregistration was uploaded only to an OSF folder and not formally registered with a timestamp. However, while an update to the document has been made according to the metadata, the content still seems to be the same as the original one (created on Jun 11, 2025).

      (3) Combined measurement of choices and metacognition:

      Choice behaviour and metacognition were not measured independently (i.e., they were measured using a single slider). This challenges the interpretation of results concerning metacognitive sensitivity, as the authors note themselves in the limitations section (even if the effect of starting position was small). It should be clarified whether responses exactly at the centre of the slider were possible or not (and what range the slider had).

      (4) Generalisability of findings:

      It is questionable whether transdiagnostic factors derived from a Japanese population can be transferred to the sample in this work, especially because it seems to be composed of a heterogeneous international sample comprising participants from multiple countries (South Africa, United States, United Kingdom, Poland, others). While it is appreciated that the authors validate the transdiagnostic factors through a new exploratory factor analysis and by correlating item loadings across samples, the correlations are only moderate and point at least to some degree of variation. In addition, for factor 2 (social withdrawal?), the correlation seems to be mainly existent due to two clusters of items, challenging the validation approach in itself. Also, apart from the correlations themselves, the absolute values of the item loadings appear to vary largely. Considering this, the conclusion that "transdiagnostic symptom dimensions appear broadly consistent across samples (even across countries)" (p. 9) definitely goes too far.

      (5) Parameter recovery analysis: The parameter recovery analysis is appreciated; however, the recovery of the learning rate in particular seems to be rather low (r=.40). This raises concerns regarding the validity of the parameter estimates.

      (6) Effect sizes: The reported regression coefficients appear relatively small in magnitude (i.e., betas of the associations between metacognition/abstraction and transdiagnostic dimensions appear to be in the range around .02-.04). If these are standardised estimates, the corresponding effect sizes are relatively small. If not, I recommend reporting standardised effect sizes in addition.

      Overall, the study provides relevant replications and new insights into the association between reward-guided learning behaviour (abstraction, metacognition) and transdiagnostic symptom dimensions. It should be noted that effect sizes appear to be rather low, however. The analytic approach is generally strong. At the same time, the high number of exclusions, the limitations regarding the validation of the transferred transdiagnostic factor structure, and the limited recovery of the learning rate parameter clearly raise important questions regarding the robustness and generalizability of findings. Given this, any interpretations regarding potential therapeutic implications (e.g., p. 10) should be made with caution.

    3. Reviewer #2 (Public review):

      Summary:

      In this study, Oka and colleagues recruited an online sample to complete a previously validated abstraction task (Cortese et al., 2021) alongside confidence ratings and a large psychiatric questionnaire battery, which included a variety of methods to screen out inattentive or otherwise biased responders. Questionnaire item scores were combined with factor weights from a large dataset to estimate transdiagnostic factor scores. A computational model was then fit in a hierarchical manner to the abstraction task data, with individuals' fit to an "Abstract RL" model used as a metric of individual-level abstraction ability, and metacognitive bias and sensitivity were estimated from the confidence ratings. Associations between these task-derived measures and both dimensional and symptom-level measures of psychopathology were then estimated using multiple regression. The key findings were that, while metacognitive sensitivity and abstraction ability were associated with symptom-level scores, the associations with transdiagnostic dimensions - higher compulsivity associated with lower abstraction ability and metacognitive sensitivity; higher social withdrawal was associated with higher metacognitive sensitivity - were interpreted by the authors as more coherent.

      Strengths:

      (1) Robust screening for inattentive responders through catch questions (Zorowitz et al., 2023), as well as incorporating recent recommendations regarding response bias (Sarna et al., 2026).

      (2) Assessed the cross-cultural generalisability of the imported factor weights by comparing item loadings from a large external sample against a de novo exploratory factor analysis in their sample.

      (3) Directly compares a theory-driven model-defined abstraction metric to metacognition in relation to dimensional and symptom-level measures of psychopathology.

      (4) Pre-registered analyses, with deviations from pre-registration clearly stated.

      Weaknesses:

      (1) The abstraction metric (mean posterior responsibility of the Abstract RL model) differed from the pre-registered metric and has not been validated here for reliability (e.g., split-half across the blocks or similar).

      (2) Model recovery is not shown, so it's not clear whether the Abstract RL and Feature RL models are fully dissociable in this task design.

      (3) Metacognitive measures are behaviourally defined (AUROC2 for sensitivity and mean confidence for bias), but models do not correct for task accuracy, which may be related to both.

      (4) Dimensional and symptom regressions differ: the dimensions are entered in one model, but the symptom measures are entered into separate regressions and the marginal effects corrected for multiple comparisons. If I've understood this correctly, this means that the dimensional coefficients are partial associations adjusting for the other two factors, whereas the symptom-level coefficients are marginal and FDR-corrected, making it difficult to directly compare them.

      Additional questions and context:

      (1) Could split-half reliability (e.g. odd vs even blocks) be reported for the abstraction measure? Relatedly, a model recovery/confusion analysis for the two abstraction models, and/or posterior predictive checks showing that the two models generate behaviour resembling that of participants would help establish that the responsibility metric is able to dissociate the different abstraction strategies.

      (2) Supplementary Table 2 shows the results for the pre-registered discrete proportion metric - here, there is limited evidence (p=0.220) of an association between abstraction and compulsivity, so saying they are "almost consistent" is perhaps a little overstated. Though the argument for using the alternative continuous metric is justified in the text, it's not quite clear whether the difference is due to the inference method or the abstraction metric itself - the bootstrapped analysis of the pre-registered metric is not reported, nor is the analysis without bootstrapping of the continuous metric (I think this may have been what Supplementary Table 1 was meant to report, but currently it's identical to Supplementary Table 5). In addition, it might be helpful if the correlation between the two metrics were presented graphically.

      (3) In the Methods and Supplement, the authors mention that they had pre-registered running a sensitivity analysis including excluded participants. This might be interesting given the high exclusion rate, and given that most exclusions were not based on task behaviour (chance-level choosing). If there is concern about shifting group-level parameter distributions, then this could be explicitly included in the model by including an offset on group-level parameters (i.e., interaction term) on excluded participants, which would allow them to systematically differ in model parameters. Alternatively, one could at least estimate the abstraction and metacognition metrics in the excluded sample (perhaps restricted to those excluded on questionnaire-based criteria rather than task performance) to see whether they do indeed differ.

      (4) How do factor scores relate to task accuracy - do those with higher compulsivity perform worse, and is this plausibly related to less abstraction?

      (5) How do the factors extracted here compare to those in other studies, such as those from Gillan et al. (2016, eLife)? In particular, it'd be interesting to know what questionnaires/symptoms in the "Compulsive hypersensitivity" factor in the present study overlap with the Compulsive behaviour/intrusive thoughts factor from that earlier work, as the latter has been strongly associated with metacognitive measures - higher metacognitive efficiency for anxious/depression, lower metacognitive efficiency for compulsive behaviour - in previous work (Rouault et al., 2018). That three-factor structure also included a factor they labelled "Social withdrawal" - is it similar to the one presented here, or is the one here (including distress) more like their anxious depressive factor?

      (6) In the Discussion, the authors state "Our findings are also consistent with previous converging evidence linking compulsive tendencies to less efficient computation and a preference for familiar over goal-directed action". That the Feature RL model might fairly be called less efficient is reasonable, but I'm not sure how the reduction of features in the Abstract RL model relates to goal-directed action (they're both model-free RL algorithms).

      (7) The Discussion also mentions "models that integrate abstract and metacognitive representations" - was there a reason these could not be applied in the present study?

    4. Reviewer #3 (Public review):

      Summary:

      This is an interesting study investigating the relationship between transdiagnostic symptom profiles and computational parameters of abstraction and metacognition. The authors found that a transdiagnostic dimension they term compulsive hypersensitivity was negatively related to abstraction ability and metacognitive sensitivity, and a dimension termed social withdrawal was positively related to metacognitive sensitivity. Overall, the question of whether and how higher-level cognitive processes relate to transdiagnostic psychiatric factors is interesting and addresses a relevant gap in the literature.

      Strengths:

      This is an overall well-designed study, and the authors have clearly given considerable thought to data quality and validation during the study design. This is evident, for instance, in the use of multiple approaches to assess inattentive responding and acquiescence tendencies. In addition, the use of advanced modelling approaches for the task data represents a strength and allows the authors to capture individual differences in task behaviour in a more nuanced and mechanistic manner in comparison to what would have been possible with model-agnostic analyses.

      Weaknesses:

      Nevertheless, I would like to raise several concerns, detailed below.

      (1) My most pressing concern regards the exclusion rate of roughly half the sample. Of the 512 participants who were tested, only 249 (48.6%) were included in the final analyses, meaning that more than half of the recruited participants were excluded. Although the authors state that these exclusions were based on preregistered criteria, this high exclusion rate still warrants further investigation. Pre-registering exclusion criteria does not eliminate the potential for selection bias or establish that the resulting sample is representative of the recruited sample. In fact, the authors even report that the excluded participants differed significantly on several psychiatric scores. I believe that a detailed account of the exclusion, including how included and excluded participants differed on relevant demographic or study variables, would be beneficial.

      Additionally, the authors state that a sensitivity analysis including participants excluded from the primary analysis was preregistered but was not conducted because the excluded and included participants differed significantly. This does not seem a sufficient justification for omitting a preregistered sensitivity analysis; in fact, systematic differences between included and excluded participants warrant this kind of sensitivity analysis. I agree with the authors that this may lead to changes in task parameter estimates. However, I believe this change would be an informative result of the sensitivity analyses rather than a methodological problem to avoid.

      Finally, I would like the authors to clarify some inconsistencies in the preregistered exclusion criteria. In the pre-registration, they first state that all participants who fail any infrequency question will be excluded. Later in the pre-registration, they state that participants with two or more failed questions will be excluded and that a sensitivity analysis will be done, in which participants with only one failed question are retained. Neither of these criteria matches what is reported in the manuscript ("Attention check + infrequency item mistake more than 2").

      (2) The current sample size of n = 249 participants is not sufficient for a factor analysis with 176 questionnaire items, and I believe that the solution of using factor weights from a previous larger sample is generally sensible. I also commend the authors for wanting to assess the validity of this approach. However, both the methods and results reported for the comparison between the factor analysis based on the current, smaller sample and the large, previous sample are insufficient and warrant substantially more detail. Firstly, it is unclear to me whether the three-factor solution in the factor analysis with the current, smaller sample was empirically grounded or whether three factors were extracted to match the three-factor solution from the larger sample. For both factor analyses, I would recommend reporting the factor extraction methods, the criterion used to determine the number of factors, and whether alternative factor solutions were considered. Secondly, it is unclear what exactly was correlated across the two factor analyses (e.g., factor scores, item loadings, factor weights?). Finally, I do not believe that the result of significant correlation is sufficient to conclude that the dimensions replicate between samples, particularly given they indicate only moderate correspondence (r = 0.46, for instance, corresponds to only 21% shared variance), thereby providing very limited evidence for replication of the factor structure.

      These concerns are particularly pressing given that the authors state the consistency of transdiagnostic symptom dimensions across samples and countries as a primary result in the discussion.

      (3) It is also unclear to me whether any exclusion was applied based on the first part of the Deary-Liewald reaction time task. The supplement implies that it was ("This task result was used to [...] exclude participants who exhibit problematic behaviours"), but I could not find a corresponding report in the manuscript.

      (4) The manuscript refers to an attentional check questionnaire used to identify and exclude inattentive participants, but does not provide a reference for this measure or describe the questions included. Given the substantial overall exclusion rate, it is particularly important that all exclusion procedures and criteria are described in sufficient detail to allow readers to assess their appropriateness and reproducibility. The authors should therefore provide the relevant reference and/or report the specific questions and criteria used to determine inattention.

      Relatedly, when describing the control-neutral items in the Supplementary Materials, the authors state that further details are provided in the Supplementary Materials. However, I cannot find another, separate Supplementary Material that may contain this information

    5. Author response:

      We thank the editors for the eLife Assessment and the reviewers for their thorough and constructive evaluation of our manuscript.

      We are glad that the evidence for reduced metacognitive sensitivity in relation to the compulsive hypersensitivity dimension was considered solid. In hindsight, we agree that the evidence for the abstraction findings is currently incomplete, and we plan to address this with additional analyses as outlined below (to clarify the robustness of the abstraction metric and its association with symptom dimensions).

      Below, we outline how we plan to address the main points raised by the reviewers, grouped thematically, given the overlap across the three reviews.

      A full, detailed point-by-point response accompanied by the corresponding analyses will follow in our formal revision response.

      (1) Exclusion rate and lack of relevant sensitivity analysis

      In the revision, we will report a detailed comparison of included versus excluded participants on demographic and symptom variables, and we will conduct the originally preregistered sensitivity analysis to estimate abstraction and metacognition metrics, including the excluded sample, and compare these metrics with psychopathological variables, rather than omitting this analysis.

      (2) Abstraction metric and its deviation from the preregistered definition

      We recognise that our primary measure of abstraction differs from the preregistered metric (the proportion of blocks better fit by the Abstract RL model), and that this change appears to affect the pattern of results. In the revision, we will report split-half reliability for the abstraction metric, present the bootstrapped analysis of the preregistered metric, and provide a clearer justification for the switch, making the source of the discrepancy (metric versus inference method) transparent.

      (3) Generalisability of the transdiagnostic factor structure

      We acknowledge that our claim of consistency between samples of the transdiagnostic dimensions requires more support and more careful framing. In the revision, we will provide full methodological detail on both factor analyses (extraction method, criteria for the number of factors, etc.), and we will moderate our interpretation, especially in the discussion, to reflect the actual strength of correspondence rather than describing the structure as broadly consistent.

      (4) Parameter and model recovery

      We will extend our recovery analyses to include a model recovery/confusion analysis between the Abstract RL and Feature RL models, and we will investigate the source of the relatively low learning-rate recovery (e.g., by testing recovery stratified by block length and without injected noise), reporting these results in the revision.

      (5) Documentation of the attention-check procedure

      We will provide all details on the attention-check items and criteria used to identify inattentive participants, including a reference where applicable, and we will standardise terminology (e.g., infrequency vs. inattention items) throughout the manuscript and supplementary information.

      (6) Other methodological and presentational clarifications

      We will review the manuscript, methods, and supplementary materials to resolve the remaining methodological ambiguities and presentational issues raised by the reviewers. This includes, among other points, clarifying whether the reported regression coefficients are standardised and correcting labelling inconsistencies, missing references/DOIs, and other minor textual issues throughout.

    1. eLife Assessment

      This valuable study compares the anti-HIV activity of human natural killer cell subsets and identifies CD56<sup>dim</sup>CD16<sup>dim</sup> cells as a potentially key effector population against HIV-infected autologous T cells. The evidence supporting the central conclusions seems incomplete because the interpretation may be substantially confounded by activation-induced CD16 downregulation, making it difficult to distinguish intrinsic functional differences from phenotypic changes that occur during target-cell engagement. Additional experiments using purified NK-cell subsets and appropriate controls would substantially strengthen the evidence supporting the major conclusions.

    2. Reviewer #1 (Public review):

      Howell et al investigate the functional capacities of CD16dim CD56dim NK cells, including their activity against HIV-1-infected T cells. The authors provide an extensive characterization of the functional activity of different NK cell subsets derived from peripheral blood. CD16 is an important receptor expressed on NK cells, and previous studies have demonstrated that CD16 expression changes depending on the activation status of NK cells - one important regulator of CD16 expression is proteolytic shedding/cleavage of CD16 on activated NK cells by the metalloprotease ADAM17. In vitro activation of NK cells, for example in response to K562 cells or other target cells, results in a rapid downregulation of the expression of CD16 on the surface of NK cells, unless an ADAM17 inhibitor is added. This is an important point to consider in the interpretation of the presented results. Overall, the manuscript includes many data in nine figures plus supplemental figures, and would benefit from some focusing of the results.

      (1) Figure 1<br /> The observation that CD16dim NK cells responded more strongly by degranulation to K562 cells and HIV-1-infected cells could be due to the shedding of CD16 following activation. In other words, more strongly activated NK cells express higher levels of CD107 but also shed CD16, resulting in higher CD107 expression in CD16low NK cells. The authors should investigate this, for example by performing the degranulation assays shown in Figure 1 in the presence and absence of an ADAM17 inhibitor.

      (2) Figure 2<br /> The authors sorted CD16dim and bright NK cells for these experiments and observed higher lysis of HIV-1-infected CD4+ T cells. Important controls should be included in these experiments - how strong was the lysis of HIV-1-uninfected CD4+ T cells by these different NK cell subsets? It also appears that the results shown were derived using NK cells from one donor, and "representative of two independent sort experiments performed with separate donors, each yielding similar results". Why are the authors now showing the respective data? One or two experiments appear too few to come to these conclusions. To support the broad conclusions drawn by the reviewers, the experiments should be performed in a larger number of individuals.

      (3) Figure 3<br /> It appears that experiments were performed again using bulk NK cell populations, and superior degranulation and killing frequencies by CD16dim NK cells might reflect different levels of activation again, as described above for Figure 1. The same applies to Figure 4 - lower degranulation events in CD16bright NK cells are consistent with lower activation of these cells, resulting in less CD16 downregulation. Also, it is not clear to the reviewer why CD107a expression and killing frequencies decrease with higher effector-to-target ratios (Figure 3).

      (4) Pages 19-25<br /> It would be helpful if the authors could provide some conclusions regarding their findings - it is very difficult for the reader to follow the many reported frequencies and p-values. What does this actually mean? Overall, the results appear to follow prior observations that licensed (KIR3DL+) NK cells respond more strongly than unlicensed (KIR3DL1neg) NK cells. The consistent observation within these different subanalyses that CD16dim NK cells degranulate more than CD16bright NK cells is probably the result of activation-induced CD16 downregulation in these assays, as mentioned above. Providing two-way ANOVA analysis results for these very many observations would furthermore require, in the opinion of the reviewer, adjustments for multiple comparisons.

      (5) Figures 5 and 6<br /> These figures demonstrate that NK cell-mediated activation by HIV-1-infected cells depends on NKG2D ligands and can be inhibited by blocking this interaction - this is consistent with data presented by the Barker group and others previously, and does not provide new information.

      (6) Figure 7<br /> The authors extended their functional analyses of NK cells to ADCC function. It is very well established that CD16 is downregulated in the context of ADCC following activation of NK cells. Consistent with this, higher degranulation is observed by CD16dim NK cells.

      (7) The data using ADAM17 inhibition in the final figures of the manuscript<br /> These data are of interest, but should be presented in a more structured way. First of all, does the addition of ADAM17 inhibitors change the overall proportion of CD16bright and dim NK cells following activation, independent of whether these cells degranulate or not? Overall, the proportion of CD16dim NK cells that degranulate appears to be reduced in the presence of the ADAM inhibitor, which is consistent with reduced CD16 shedding and maintenance of CD16 expression on activated NK cells - and this is supported by the increase in CD107a-positive NK cells that express CD16 (Figure 8a). Overall, the differences between CD16bright and dim NK cells in their level of activation appear to disappear in the presence of an ADAM17 inhibitor, based on the data shown in Figure 8b, suggesting that CD16 downregulation is occurring in response to activation of NK cells as a consequence of CD16 shedding, and can be inhibited by an ADAM17 inhibitor.

      Taken together, many of the data presented in the manuscript are consistent with the very well-established downregulation of CD16 expression on activated NK cells, suggesting that the observed association between reduced CD16 expression on CD56dim NK cells and enhanced effector functions is a consequence of higher activation of these NK cells.

    3. Reviewer #2 (Public review):

      Summary:

      This study investigates the cytotoxic activity of human NK-cell subsets against autologous HIV-1-infected CD4 T cells and identifies CD56dimCD16dim NK cells as the dominant effector population. The authors propose that this subset possesses superior cytotoxic activity compared with CD56dimCD16bright NK cells and could therefore represent an attractive target for HIV cure strategies. While the study addresses an important and clinically relevant question, several of its major conclusions rely on assumptions that are not adequately supported by the experimental design. In particular, CD16 is treated as a stable phenotypic marker throughout most of the study despite its well-established and rapid downregulation following NK cell activation.

      Strengths:

      (1) The study addresses an important and clinically relevant question regarding which NK cell subset is responsible for the elimination of autologous HIV-1-infected cells. To the best of my knowledge, this is the first study directly comparing the anti-HIV functional activities of CD56dimCD16dim vs CD56dimCD16bright NK cells.

      (2) The experiments performed with purified NK cell subset (Figure 2) provide some evidence that CD56dimCD16dim NK cells possess enhanced cytotoxic activity relative to CD56dimCD16bright NK cells. This experimental approach is considerably more convincing than the analyses performed on mixed NK cell populations and should be expanded throughout the study.

      Weaknesses:

      (1) The central conclusion is weakened by the use of CD16 as a stable phenotypic marker. CD16 is well established to be rapidly downregulated following NK-cell activation and target cell (K562 or infected cells) engagement through ADAM17-mediated shedding. NK cell shedding regulates NK cell effector functions by promoting target cell detachment, boosting serial killing capacity, and preventing overstimulation. Therefore, NK cells displaying a CD56dimCD16dim phenotype after co-culture cannot be assumed to represent a pre-existing subset with intrinsically superior cytotoxic activity, but may instead correspond to activated CD56dimCD16bright NK cells that have downregulated CD16 during the assay. Because the vast majority of the functional experiments classified NK cell subsets based on post-assay CD16 expression, it is difficult to distinguish intrinsic functional differences between NK cell subsets from activation-induced phenotypic conversion. This limitation affects the interpretation of most of the study's principal findings.

      (2) The "killing frequency" analysis presented in Figure 3 is based on a mathematical estimate rather than a direct experimental measurement. Since total target cell killing is measured in mixed NK cell populations, it cannot be attributed to individual NK cell subsets. This experiment must be repeated using purified NK cell subsets.

      (3) The serial degranulation assay presented in Figure 4 does not directly measure serial target cell killing and therefore does not support the conclusion that CD56dimCD16dim NK cells possess superior serial killing capacity. Furthermore, the increased serial degranulation observed in the CD16dim population could simply reflect activation-induced CD16 downregulation rather than an intrinsic property of this subset. This experiment should therefore be repeated using purified NK cell subsets.

      (4) The finding that CD56dimCD16dim NK cells exhibit greater ADCC activity is somewhat counterintuitive given the central role of CD16 in mediating ADCC. Moreover, these experiments are likely confounded by activation-induced CD16 downregulation, which is expected to be even more pronounced during ADCC. Thus, the apparent superiority of the CD56dimCD16dim subset may simply reflect the conversion of activated CD56dimCD16bright NK cells into the CD16dim gate rather than intrinsically greater ADCC activity. To directly compare the intrinsic ADCC capacity of each subset, these experiments should be repeated using purified NK cell populations prior to target-cell stimulation.

    4. Author response:

      We thank the editors and all three reviewers for their careful and constructive evaluation of our manuscript. We recognize that a single concern, the possibility that cells classified as CD56<sup>dim</sup>CD16<sup>dim</sup> after co-culture represent activated CD56<sup>dim</sup>CD16<sup>bright</sup> cells that have shed CD16 rather than a pre-existing subset, underlies the majority of the comments. We therefore address this concern first, in a central response, and then respond to each reviewer and editor comment in turn. Where a comment relates to this shared concern, we point to the central response rather than repeating the argument.

      The analytical and presentational revisions described below are complete: the statistical analyses have been re-run with corrections for multiple comparisons, the Discussion has been rewritten, the figures and supplemental tables have been renumbered and corrected, and existing data on pre-stimulation receptor expression and NKp30 have been incorporated. These will appear in the revised manuscript. The new experiments described will be completed within approximately six to eight weeks and provided with the revised manuscript.

      Central are CD56<sup>dim</sup>CD16<sup>dim</sup> cells a pre-existing subset, or activated CD56<sup>dim</sup>CD16<sup>bright</sup> cells that have shed CD16?

      We agree with the premise that CD16 is rapidly shed by ADAM17 upon activation, and that classifying subsets by post-assay CD16 expression alone cannot, on its own, distinguish a pre-existing subset from activation-induced conversion. For this reason, our conclusion does not rest on post-assay classification. The evidence below, from experiments already in the manuscript, argues against activation-induced conversion, and we will strengthen it with the expanded sorted-subset experiments described at the end.

      Cells sorted before target exposure establish the advantage independently of any during-assay shedding (Fig. 2, unchanged in the revised manuscript).

      The most direct evidence comes from subsets purified before the assay. In Fig. 2, NK cells were sorted into CD56<sup>dim</sup>CD16<sup>dim</sup> and CD56<sup>dim</sup>CD16<sup>bright</sup> populations before any exposure to target cells, and their cytolytic function was measured as specific lysis of autologous HIV-infected T cells. Purified CD16<sup>dim</sup> cells lysed infected targets approximately twice as efficiently as purified CD16<sup>bright</sup> cells across the effector-to-target range, reaching 77.58% versus 39.18% at 1:1. The difference was significant at 1:4 (p = 0.0008), 1:2 and 1:1 (both p < 0.0001); at the lowest ratio tested, 1:8, specific lysis was low in both subsets and the difference did not reach significance (p = 0.2121; Supplemental Table 4). The additional donors described below will allow this comparison to be made across a larger data set, including at the lowest ratios where specific lysis is low in both subsets.

      Because the subsets are defined by sorting before target contact, and because the readout is direct target lysis rather than post-assay CD16 gating, this advantage cannot arise from activation-induced CD16 shedding during the assay. Public reviewer 2 and the peer reviewer both identified this experiment as the strongest evidence in the manuscript. Its interpretation is secure; what it requires is additional donors for statistical robustness, which we provide in the planned expansion below.

      The two subsets respond to ADAM17 inhibition in opposite directions (Figs. 7 and 9, now Figures 6 and 8).

      If CD56<sup>dim</sup>CD16<sup>dim</sup> cells were simply CD56<sup>dim</sup>CD16<sup>bright</sup> cells that had shed CD16, the two would be one population sampled at different points along a shedding continuum, and inhibiting ADAM17 would move them in the same direction. Instead, ADAM17 inhibition moves them in opposite directions. In the antibody-dependent degranulation assay (Fig. 7B, now Figure 6B), ADAM17 inhibition increased CD56<sup>dim</sup>CD16<sup>bright</sup> degranulation but decreased CD56<sup>dim</sup>CD16<sup>dim</sup> degranulation across VRC01 concentrations. The same opposition is seen when serial degranulation is resolved by the number of degranulation events per cell (Fig. 9, now Figure 8): ADAM17 inhibition increased multiple degranulation events in CD56<sup>dim</sup>CD16<sup>bright</sup> cells, with cells undergoing three events rising from 0.79% to 5.12%, while in CD56<sup>dim</sup>CD16<sup>dim</sup> cells it reduced them, with three events falling from 15.16% to 2.24% and the non-degranulating fraction rising from 61.56% to 92.13%.

      A single population would not be expected to respond to the same perturbation in opposite directions, and these observations are difficult to reconcile with the CD16<sup>dim</sup> cells being activated CD16<sup>bright</sup> cells; rather, they point to two functionally distinct subsets with opposite dependence on ADAM17 activity. This is consistent with our model, in which CD16<sup>dim</sup> cells use ADAM17-mediated shedding to detach and serially re-engage, whereas CD16<sup>bright</sup> cells are hindered by the loss of CD16.

      The CD16<sup>dim</sup> degranulation advantage is driven by NKG2D through a mechanism separable from ADAM17 (Fig. 8, now Figure 7).

      The change in subset frequency on exposure to VRC01-treated infected cells (Fig. 8B, now Figure 7B) is abolished by anti-NKG2D even though VRC01 and ADAM17 remain present, indicating that this frequency shift is driven by NKG2D-dependent activation rather than by antibody-CD16 engagement alone.

      The degranulation data in Fig. 8A (now Figure 7A) show that the two perturbations act differently in the two subsets. In CD56<sup>dim</sup>CD16<sup>dim</sup> cells, both reduce the response, and the combination reduces it further than either alone: from 12.60% under vehicle to 5.24% with anti-NKG2D (p < 0.0001), 5.05% with ADAM17 inhibition (p < 0.0001), and 2.03% with both (p < 0.0001 versus vehicle; p < 0.0001 versus anti-NKG2D alone; p = 0.0001 versus ADAM17 inhibition alone). If anti-NKG2D acted only by removing the activation trigger for ADAM17, that is, if NKG2D and ADAM17 lay on a single linear pathway, blocking the pathway at two points would not be expected to add to the effect of either alone. The further reduction therefore indicates that NKG2D and ADAM17 contribute through separable mechanisms.

      In CD56<sup>dim</sup>CD16<sup>bright</sup> cells the two perturbations act in opposite directions. Anti-NKG2D reduced degranulation from 2.58% to 0.92% (p = 0.0263), whereas ADAM17 inhibition increased it to 3.99% (p = 0.0679). NKG2D therefore supports the response of CD56<sup>dim</sup>CD16<sup>bright</sup> cells while ADAM17 activity constrains it, the reverse of the pattern in CD56<sup>dim</sup>CD16<sup>dim</sup> cells, where ADAM17 activity is required. Two populations differing only in the extent to which they have shed CD16 would not be expected to respond to the same two perturbations in opposite ways.

      Supporting evidence: pre-sorted subsets are stable and differ before stimulation.

      Two further observations support a pre-existing subset. First, we have directly tracked the fate of each subset sorted before target exposure. NK cells were sorted into CD16<sup>bright</sup> and CD16<sup>dim</sup> subsets, exposed to HIV-infected cells for one hour, and reanalyzed for CD16 expression. One hour is the point at which we observe the highest frequency of degranulating cells in both subsets (Figure 3—Figure Supplement 1A in the revised manuscript), and therefore the point at which activation-induced shedding would be most likely to be detected. Sorted CD16<sup>bright</sup> cells remained predominantly CD16<sup>bright</sup> (approximately 62%); of those that lost CD16, most became CD16<sup>negative</sup> (approximately 35%) rather than CD16<sup>dim</sup> (approximately 4%). Sorted CD16<sup>dim</sup> cells likewise shifted predominantly to a CD16<sup>negative</sup> phenotype (approximately 75%). Activation-induced CD16 shedding therefore directs cells of both subsets toward the CD16<sup>negative</sup> gate rather than generating the CD16<sup>dim</sup> population from CD16<sup>bright</sup> cells. These data are presented in Author response image 1.

      Author response image 1.

      Phenotype of sorted CD56<sup>dim</sup>CD16<sup>bright</sup> and CD56<sup>dim</sup>CD16<sup>dim</sup> NK cells after exposure to HIV-infected T-cells. NK cells were sorted into CD56<sup>dim</sup>CD16<sup>bright</sup> and CD56<sup>dim</sup>CD16<sup>dim</sup> subsets, exposed to purified autologous productively HIV-1<sup>SHM-1</sup>-infected T cells for 1 hour at a 1:1 effector-to-target cell ratio, and reanalyzed for CD16 expression. Bars show the percentage of each sorted NK cell population (CD56<sup>dim</sup>CD16<sup>bright</sup> and CD56<sup>dim</sup>CD16<sup>dim</sup>) falling into the CD16<sup>bright</sup>, CD16<sup>dim</sup>, and CD16<sup>negative</sup> gates after exposure, as the mean ± standard deviation (SD) of three replicates. One hour is when the highest frequency of degranulating cells is observed in both subsets.

      Second, the subsets differ before any stimulation. CD56<sup>dim</sup>CD16<sup>dim</sup> cells express higher NKG2D than CD56<sup>dim</sup>CD16<sup>bright</sup> cells before any target-cell contact. In the no-target condition, NKG2D was 1182 gMFI higher on CD56<sup>dim</sup>CD16<sup>dim</sup> cells (p < 0.0001), a difference of approximately 1.6-fold; across all conditions tested the subset means were 2521.5 versus 1772.2 gMFI, or 1.42-fold (two-way ANOVA: subset F(1, 20) = 1897, p < 0.0001, 70.97% of the total variation; Fig. 6, now Figure 5; Supplemental Table 18). This difference is specific to NKG2D: NKp46, measured on the same cells in the same wells, did not differ between the subsets in the no-target condition (mean difference 45.67 gMFI, p = 0.0668), and was higher on CD56<sup>dim</sup>CD16<sup>bright</sup> cells when targets were present. The subsets therefore differ in NKG2D density before activation, and not in activating receptor density generally.

      Planned strengthening.

      To place this beyond doubt, we will expand the sorted-subset experiments, performing the specific-lysis assay (Fig. 2) and the antibody-dependent degranulation assay on subsets purified before target exposure across additional donors, together with uninfected-target controls. If cell yields from the sort permit, we will also perform the serial degranulation assay on sorted subsets; because the CD56<sup>dim</sup>CD16<sup>dim</sup> subset constitutes fewer than 5% of CD56<sup>dim</sup> NK cells and the serial degranulation assay requires four sequential labelling and washing steps, we cannot commit to this in advance of the sort. These experiments require sorting and primary-cell work and will be completed within approximately six to eight weeks and provided with the revised manuscript.

      We note that performing every functional assay in this study on sorted subsets is not feasible within the scope of this revision. Sorting the CD56<sup>dim</sup>CD16<sup>dim</sup> subset, which constitutes fewer than 5% of CD56<sup>dim</sup> NK cells, from a sufficient number of donors to repeat the full panel of assays would require resources beyond those currently available to us. We have therefore prioritized the specific-lysis and antibody-dependent degranulation assays, which bear most directly on the concern raised by the reviewers and the editor, and will extend the approach to the remaining assays as resources allow.

      Public Reviews:

      Reviewer #1 (Public review):

      Overall organization

      Overall, the manuscript includes many data in nine figures plus supplemental figures, and would benefit from some focusing of the results.

      We agree, and we have reduced the main figures from nine to eight. The NKG2D ligand histograms, previously Figure 5A, have been removed for the reason given in our response to the peer reviewer, Recommendation 4. The degranulation against wild-type and ΔVpr-infected targets, previously Figure 5B, and the degranulation and killing frequency measurements in mixed populations, previously Figure 3, have been moved to the supplementary material. The inhibitory receptor analyses have been reduced from approximately 1,800 words and more than 80 reported p-values to approximately 700 words and 24, with the detail retained in the supplemental tables. Each Results section now opens with a statement of the principal finding before the supporting data, and the Discussion synthesizes what the findings mean rather than restating them.

      Figure 1 and Figure 1—figure supplement 2

      The observation that CD16dim NK cells responded more strongly by degranulation to K562 cells and HIV-1-infected cells could be due to the shedding of CD16 following activation. In other words, more strongly activated NK cells express higher levels of CD107 but also shed CD16, resulting in higher CD107 expression in CD16low NK cells. The authors should investigate this, for example by performing the degranulation assays shown in Figure 1 in the presence and absence of an ADAM17 inhibitor.

      We thank the reviewer for raising this important point, which we recognize is shared by all three reviewers and the editors, and which we address in full in the central response above. We note for clarity that the K562 data are not in Figure 1; they are presented in Figure 1—figure supplement 2, both in the reviewed preprint and in the revised manuscript. They were included to reproduce a previously established finding (Amand et al., Front Immunol 2017;8:699), and were not intended as a central experimental claim; the mechanistic focus of this study is the response to HIV-infected cells. Notably, that same study addressed the question by sorting the CD56<sup>dim</sup> subsets before stimulation rather than gating after, and our study applies the same approach and extends it to the HIV-infected setting. Four lines of evidence argue against the interpretation that CD16<sup>dim</sup> degranulation reflects activation-induced CD16 shedding of CD16<sup>bright</sup> cells:

      (1) In cells sorted before target exposure, purified CD16<sup>dim</sup> cells lyse HIV-infected targets approximately twice as efficiently as purified CD16<sup>bright</sup> cells across the effector-to-target range, significantly so at 1:4 and above (Fig. 2; Supplemental Table 4); because the subsets are defined before any activation and the readout is direct target lysis rather than post-assay CD16 gating, this advantage cannot arise from shedding during the assay.

      (2) The two subsets respond to ADAM17 inhibition in opposite directions, both in the magnitude of degranulation (Fig. 7B, now Figure 6B) and in the number of serial degranulation events per cell (Fig. 9, now Figure 8): ADAM17 inhibition increased degranulation in CD16<sup>bright</sup> cells but decreased it in CD16<sup>dim</sup> cells.

      (3) Blocking NKG2D together with ADAM17 reduced CD16<sup>dim</sup> degranulation below either treatment alone (Fig. 8A, now Figure 7A), indicating that the CD16<sup>dim</sup> advantage is driven by NKG2D through a mechanism separable from ADAM17-mediated shedding.

      (4) The subsets also differ before any stimulation, with NKG2D 1182 gMFI higher on CD16<sup>dim</sup> cells in the no-target condition (p < 0.0001) and no corresponding difference in NKp46 under the same condition (Fig. 6, now Figure 5).

      We will further strengthen these findings by expanding the sorted-subset experiments across additional donors, as described in the central response.

      Figure 2

      The authors sorted CD16dim and bright NK cells for these experiments and observed higher lysis of HIV-1-infected CD4+ T cells. Important controls should be included in these experiments - how strong was the lysis of HIV-1-uninfected CD4+ T cells by these different NK cell subsets? It also appears that the results shown were derived using NK cells from one donor, and "representative of two independent sort experiments performed with separate donors, each yielding similar results". Why are the authors now showing the respective data? One or two experiments appear too few to come to these conclusions. To support the broad conclusions drawn by the reviewers, the experiments should be performed in a larger number of individuals.

      We thank the reviewer for these constructive points.

      (1) Uninfected-target control. We agree this is an important control and will include lysis of uninfected autologous CD4 T cells by the sorted CD16<sup>dim</sup> and CD16<sup>bright</sup> subsets in Figure 2, confirming that the observed lysis is specific to HIV-infected targets. We note that the corresponding CD107a degranulation controls against uninfected targets are presented in Figure 1—Figure Supplement 4C of the revised manuscript.

      (2) Number of donors and presentation of data. We agree that the conclusions require more than the representative donor shown. As described in the central response, we will expand these sorted-subset experiments to a larger number of individuals and will present the data from all donors rather than a single representative experiment. Because they require cell sorting and primary-cell work, these experiments will be completed within approximately six to eight weeks and provided with the revised manuscript.

      Figures 3 and 4

      It appears that experiments were performed again using bulk NK cell populations, and superior degranulation and killing frequencies by CD16dim NK cells might reflect different levels of activation again, as described above for Figure 1. The same applies to Figure 4 - lower degranulation events in CD16bright NK cells are consistent with lower activation of these cells, resulting in less CD16 downregulation. Also, it is not clear to the reviewer why CD107a expression and killing frequencies decrease with higher effector-to-target ratios (Figure 3).

      (1) Activation-induced shedding in bulk experiments (Figs. 3 and 4, now Figure 2—figure supplement 1 and Figure 3). We agree that these figures use bulk NK cell populations gated by CD16, and we address the underlying shedding concern in full in the central response. The concern that the lower serial degranulation of CD16<sup>bright</sup> cells in the direct-killing assay simply reflects lower activation and therefore less shedding is addressed directly by our ADAM17-inhibition data. At 0 µg/mL VRC01, that is, in the absence of antibody, ADAM17 inhibition already affects the two subsets differently rather than in the same direction (Figs. 7B and 7C, now Figures 6B and 6C), as would be expected if they were one population differing only in activation level. This differential response is also seen across the antibody-dependent conditions in Figs. 7B and 9 (now Figures 6B and 8). As described in the central response, we will additionally repeat the specific-lysis and antibody-dependent degranulation measurements on subsets purified before target exposure across additional donors, and the serial degranulation assay as well if cell yields from the sort permit.

      (2) Decrease in CD107a and killing frequency at higher effector-to-target ratios (Fig. 3, now Figure 2—figure supplement 1). This reflects the nature of the readout. CD107a mobilization is measured per effector cell, as the percentage of NK cells that degranulate, and is therefore maximized when targets are in excess. At low effector-to-target ratios, nearly every NK cell can encounter and engage a target, yielding a high percentage of CD107a-positive cells; at high ratios, targets become limiting, so a large fraction of NK cells never contact a target and remain unstimulated, and the rapid destruction of the limited target pool further reduces the stimulus available to the remaining cells. This lowers the measured per-effector degranulation frequency even as the absolute number of targets killed is maintained, and it is distinct from a lysis assay, which measures the fate of the target population and accordingly rises with increasing effector-to-target ratio. The same per-effector readout behavior applies to the degranulation data shown in Figs. 5C and 6C (now Figures 4A and 4B, and Figure 5C). This explanation has been added to the revised Discussion.

      Pages 19-25

      It would be helpful if the authors could provide some conclusions regarding their findings - it is very difficult for the reader to follow the many reported frequencies and p-values. What does this actually mean? Overall, the results appear to follow prior observations that licensed (KIR3DL+) NK cells respond more strongly than unlicensed (KIR3DL1neg) NK cells. The consistent observation within these different subanalyses that CD16dim NK cells degranulate more than CD16bright NK cells is probably the result of activation-induced CD16 downregulation in these assays, as mentioned above. Providing two-way ANOVA analysis results for these very many observations would furthermore require, in the opinion of the reviewer, adjustments for multiple comparisons.

      We thank the reviewer, and we have addressed this in three ways.

      (1) Readability. We agree that these sections were difficult to follow as presented. We have rewritten them, opening each with a statement of the principal finding before the supporting statistics, and reducing the inhibitory receptor section from approximately 1,800 words and more than 80 reported p-values to approximately 700 words and 24, with the detail retained in the supplemental tables. The Discussion now synthesizes what the findings mean.

      (2) CD16<sup>dim</sup> degranulation in these subanalyses. The consistent observation that CD16<sup>dim</sup> cells degranulate more than CD16<sup>bright</sup> cells across these subanalyses is addressed in full in the central response, where several lines of evidence, including subsets sorted before target exposure (Fig. 2) and the opposite responses of the two subsets to ADAM17 inhibition (Figs. 7B and 9, now Figures 6B and 8), argue against activation-induced CD16 downregulation as the explanation.

      (3) Multiple comparisons. We agree, and we have re-analyzed these comparisons, applying the post-hoc test matched to each comparison structure: Dunnett's where every subset is compared against a single designated subset, Tukey's where all pairwise comparisons are of interest, and Šidák’s where a prespecified subset of comparisons is of interest. Adjusted p-values are reported throughout, and the design and post-hoc test used for each figure and panel are given in a new supplemental table. The streamlining described above has also reduced the number of comparisons reported in the main text.

      Figures 5 and 6

      These figures demonstrate that NK cell-mediated activation by HIV-1-infected cells depends on NKG2D ligands and can be inhibited by blocking this interaction - this is consistent with data presented by the Barker group and others previously, and does not provide new information.

      We agree that the dependence of NK-cell recognition of HIV-infected cells on NKG2D and its ligands is established, including in our own earlier work (Ward et al., PLoS Pathog 2009;5(10):e1000613) and by others, and we do not present that dependence as a novel finding.

      On review, the histograms in Figure 5A were reproduced from that earlier study and should not have been included without attribution. We have removed that panel and cited the original finding in its place. Figures 5B and 5C are both new results from this study, and both are retained. Figure 5B, which shows that NK cells degranulate in response to wild-type HIV-infected targets but not to ΔVpr-infected or uninfected targets, establishing that the degranulation response in this system depends on Vpr, becomes Figure 4—figure supplement 1 in the revised manuscript. Figure 5C, the NKG2D blockade experiment, becomes Figure 4A and 4B.

      We would also distinguish Figure 6 (now Figure 5), which we consider a substantive finding rather than a restatement of the established NKG2D-ligand dependence. That figure shows that NKG2D expression differs at the level of the individual subsets, and that the difference is present before stimulation and is specific to NKG2D. In the no-target condition, NKG2D was 1182 gMFI higher on CD56<sup>dim</sup>CD16<sup>dim</sup> cells (p < 0.0001), approximately 1.6-fold, while NKp46 measured on the same cells in the same wells did not differ (mean difference 45.67 gMFI, p = 0.0668). This provides a candidate mechanism for the superior effector function of the CD56<sup>dim</sup>CD16<sup>dim</sup> subset, in addition to their serial-degranulation capacity, and it bears directly on the central question of whether the two subsets differ intrinsically rather than as a consequence of activation. The revised text presents the established NKG2D-ligand dependence as context while making the subset-level NKG2D difference, and its mechanistic significance, more prominent.

      Figure 7

      The authors extended their functional analyses of NK cells to ADCC function. It is very well established that CD16 is downregulated in the context of ADCC following activation of NK cells. Consistent with this, higher degranulation is observed by CD16dim NK cells.

      We agree that CD16 is downregulated during antibody-dependent responses, and this is precisely why we included the ADAM17-inhibition experiments within Figure 7 (now Figure 6), to determine whether the higher degranulation of CD56<sup>dim</sup>CD16<sup>dim</sup> cells is a consequence of that shedding or a property of a distinct subset. As detailed in the central response, these experiments argue against the shedding interpretation. In Fig. 7B (now Figure 6B), inhibiting ADAM17 affects the two subsets in opposite directions: it increases the degranulation of CD56<sup>dim</sup>CD16<sup>bright</sup> cells while decreasing that of CD56<sup>dim</sup>CD16<sup>dim</sup> cells. If the CD16<sup>dim</sup> cells were simply CD16<sup>bright</sup> cells that had shed CD16, blocking shedding would be expected to move the two in the same direction; the opposite responses instead indicate two distinct populations with opposite functional dependence on ADAM17 activity. The same opposition is seen when serial degranulation is resolved by the number of events per cell (Fig. 9, now Figure 8). Thus, while CD16 downregulation during antibody-dependent responses is well established, these data indicate that the superior response of the CD56<sup>dim</sup>CD16<sup>dim</sup> subset is not explained by it. This interpretation is now explicit in the revised Discussion.

      ADAM17-inhibition data (final figures)

      These data are of interest, but should be presented in a more structured way. First of all, does the addition of ADAM17 inhibitors change the overall proportion of CD16bright and dim NK cells following activation, independent of whether these cells degranulate or not? Overall, the proportion of CD16dim NK cells that degranulate appears to be reduced in the presence of the ADAM inhibitor, which is consistent with reduced CD16 shedding and maintenance of CD16 expression on activated NK cells - and this is supported by the increase in CD107a-positive NK cells that express CD16 (Figure 8a). Overall, the differences between CD16bright and dim NK cells in their level of activation appear to disappear in the presence of an ADAM17 inhibitor, based on the data shown in Figure 8b, suggesting that CD16 downregulation is occurring in response to activation of NK cells as a consequence of CD16 shedding, and can be inhibited by an ADAM17 inhibitor.

      We thank the reviewer for these suggestions, which we have used to present the ADAM17-inhibition data more clearly.

      (1) Effect on subset proportions, independent of degranulation. This is shown in Fig. 8B (now Figure 7B). Because the two subsets differ greatly in baseline frequency, with CD56<sup>dim</sup>CD16<sup>bright</sup> cells constituting the large majority of CD56<sup>dim</sup> NK cells before stimulation, a change in raw bulk proportion is small and difficult to interpret, for example a shift from roughly 95% to 92.5% of the bright population. To place the two subsets on comparable footing, the panel reports, for each subset, the frequency following target exposure minus its frequency in the matched unstimulated condition. Presented this way, ADAM17 inhibition clearly reduces the activation-associated change in subset proportions, consistent with reduced CD16 shedding. This normalization is stated in the legend and is now described in the Results text so that the analysis is not overlooked.

      (2) Interpretation of the ADAM17-inhibition data. We agree that CD16 downregulation occurs as a consequence of activation-induced shedding and is prevented by ADAM17 inhibition; this is not in dispute. We would, however, offer an additional observation that bears on whether the between-subset functional difference is itself a product of that shedding. In Fig. 8A (now Figure 7A), combining ADAM17 inhibition with NKG2D blockade reduces CD56<sup>dim</sup>CD16<sup>dim</sup> degranulation to 2.03%, below both anti-NKG2D alone at 5.24% and ADAM17 inhibition alone at 5.05% (p < 0.0001 and p = 0.0001 respectively). If the CD16<sup>dim</sup> advantage were solely a consequence of CD16 shedding, and if NKG2D blockade acted only by reducing that shedding, the combination could not reduce degranulation further than ADAM17 inhibition alone.

      The same figure also shows that the two perturbations act in opposite directions within the CD56<sup>dim</sup>CD16<sup>bright</sup> subset: anti-NKG2D reduced their degranulation from 2.58% to 0.92% (p = 0.0263), whereas ADAM17 inhibition increased it to 3.99% (p = 0.0679). NKG2D therefore supports the response of CD56<sup>dim</sup>CD16<sup>bright</sup> cells while ADAM17 activity constrains it, the reverse of the pattern in CD56<sup>dim</sup>CD16<sup>dim</sup> cells. Together with the opposite responses of the two subsets to ADAM17 inhibition described in the central response, this indicates that NKG2D and ADAM17 contribute through separable mechanisms and that the two subsets are not one population at different stages of shedding. These data are now presented in a more structured form and the interpretation is explicit in the revised text.

      Summary statement

      Taken together, many of the data presented in the manuscript are consistent with the very well-established downregulation of CD16 expression on activated NK cells, suggesting that the observed association between reduced CD16 expression on CD56dim NK cells and enhanced effector functions is a consequence of higher activation of these NK cells.

      We appreciate the reviewer articulating the central concern so clearly. We agree that CD16 downregulation on activated NK cells is well established and occurs in our assays; where we reach a different conclusion is on whether the enhanced function of the CD56<sup>dim</sup>CD16<sup>dim</sup> subset is a consequence of that downregulation. As set out in the central response, three observations argue that it is not: the advantage is present in cells sorted into subsets before any target contact, where post-assay CD16 changes cannot apply (Fig. 2); the two subsets respond to ADAM17 inhibition in opposite directions, both in magnitude (Fig. 7B, now Figure 6B) and in the number of serial degranulation events per cell (Fig. 9, now Figure 8), which is difficult to reconcile with their being one population at different activation levels; and blocking NKG2D together with ADAM17 reduces CD16<sup>dim</sup> degranulation below either alone (Fig. 8, now Figure 7), indicating that the advantage is driven by NKG2D through a mechanism separable from shedding. We therefore interpret the association between low CD16 and enhanced function not as activation-induced downregulation of a single population, but as a property of a distinct, pre-existing subset. This interpretation is stated and defended explicitly in the revised Discussion, and will be strengthened with the expanded pre-sorted experiments.

      Reviewer #2 (Public review):

      (1) The central conclusion is weakened by the use of CD16 as a stable phenotypic marker. CD16 is well established to be rapidly downregulated following NK-cell activation and target cell (K562 or infected cells) engagement through ADAM17-mediated shedding. NK cell shedding regulates NK cell effector functions by promoting target cell detachment, boosting serial killing capacity, and preventing overstimulation. Therefore, NK cells displaying a CD56dimCD16dim phenotype after co-culture cannot be assumed to represent a pre-existing subset with intrinsically superior cytotoxic activity, but may instead correspond to activated CD56dimCD16bright NK cells that have downregulated CD16 during the assay. Because the vast majority of the functional experiments classified NK cell subsets based on post-assay CD16 expression, it is difficult to distinguish intrinsic functional differences between NK cell subsets from activation-induced phenotypic conversion. This limitation affects the interpretation of most of the study's principal findings.

      We thank the reviewer for this careful and well-articulated concern, which we recognize as the central issue of the review, and which we address in full in the central response above. We agree with the reviewer's premises: CD16 is rapidly shed by ADAM17 upon activation, and this shedding is itself functionally important, promoting target detachment, supporting serial engagement, and limiting overstimulation. Indeed, ADAM17-mediated shedding is integral to the serial degranulation mechanism we propose. We also agree that classifying subsets by post-assay CD16 expression alone cannot, on its own, distinguish a pre-existing subset from activation-induced conversion.

      For this reason, our conclusion does not rest on post-assay classification. As detailed in the central response, the CD56<sup>dim</sup>CD16<sup>dim</sup> advantage is demonstrated in cells sorted into subsets before any target contact, where the readout is direct lysis rather than post-assay gating and where activation-induced shedding therefore cannot account for the difference (Fig. 2), an experiment both this reviewer and the peer reviewer identify as the strongest in the manuscript. This is reinforced by evidence that the two subsets are functionally distinct rather than one population caught at different stages of shedding: they respond to ADAM17 inhibition in opposite directions, both in the magnitude of degranulation (Fig. 7B, now Figure 6B) and in the number of serial degranulation events per cell (Fig. 9, now Figure 8); blocking NKG2D together with ADAM17 reduces CD16<sup>dim</sup> degranulation below either alone, indicating a mechanism separable from shedding (Fig. 8, now Figure 7); and the subsets differ before stimulation, with NKG2D 1182 gMFI higher on CD16<sup>dim</sup> cells and no corresponding difference in NKp46 (Fig. 6, now Figure 5). We will strengthen this further by expanding the sorted-subset experiments across additional donors, and the revised text rests the manuscript's conclusions explicitly on the pre-sorted data.

      (2) The "killing frequency" analysis presented in Figure 3 is based on a mathematical estimate rather than a direct experimental measurement. Since total target cell killing is measured in mixed NK cell populations, it cannot be attributed to individual NK cell subsets. This experiment must be repeated using purified NK cell subsets.

      We agree that the killing frequency in Fig. 3 (now Figure 2—figure supplement 1C) is a mathematical estimate rather than a direct measurement. We would add that this is intrinsic to the metric: killing frequency is a derived quantity whether calculated from mixed or purified populations, so repeating it on purified subsets would not convert it into a direct measurement. This analysis has been moved to the supplementary material, and its limitations are stated in the Discussion, namely that killing frequency is an estimate and should be interpreted as such. Direct, subset-resolved killing is instead provided by Figure 2, in which NK cells sorted before target exposure show that purified CD56<sup>dim</sup>CD16<sup>dim</sup> cells lyse HIV-infected targets more efficiently than purified CD56<sup>dim</sup>CD16<sup>bright</sup> cells; this is the measurement on which our conclusion regarding direct killing rests, and it is the experiment we will expand across additional donors.

      (3) The serial degranulation assay presented in Figure 4 does not directly measure serial target cell killing and therefore does not support the conclusion that CD56dimCD16dim NK cells possess superior serial killing capacity. Furthermore, the increased serial degranulation observed in the CD16dim population could simply reflect activation-induced CD16 downregulation rather than an intrinsic property of this subset. This experiment should therefore be repeated using purified NK cell subsets.

      We agree that Fig. 4 (now Figure 3) measures serial degranulation, not serial killing directly. The text has been revised throughout to describe this as serial degranulation rather than serial killing, so that our conclusions match what was measured. Regarding the concern that increased serial degranulation in CD16<sup>dim</sup> cells reflects activation-induced CD16 downregulation, we address this in the central response; the opposite responses of the two subsets to ADAM17 inhibition (Figs. 7B and 9, now Figures 6B and 8) argue against that interpretation. As the reviewer suggests, we will repeat the serial degranulation assay on subsets purified before target exposure if cell yields from the sort permit. We note that this assay requires four sequential labelling and washing steps and that the CD56<sup>dim</sup>CD16<sup>dim</sup> subset constitutes fewer than 5% of CD56<sup>dim</sup> NK cells, so the number of sorted cells recovered may be limiting; we will report the outcome either way.

      (4) The finding that CD56dimCD16dim NK cells exhibit greater ADCC activity is somewhat counterintuitive given the central role of CD16 in mediating ADCC. Moreover, these experiments are likely confounded by activation-induced CD16 downregulation, which is expected to be even more pronounced during ADCC. Thus, the apparent superiority of the CD56dimCD16dim subset may simply reflect the conversion of activated CD56dimCD16bright NK cells into the CD16dim gate rather than intrinsically greater ADCC activity. To directly compare the intrinsic ADCC capacity of each subset, these experiments should be repeated using purified NK cell populations prior to target-cell stimulation.

      We agree that the greater antibody-dependent response of CD56<sup>dim</sup>CD16<sup>dim</sup> cells is counterintuitive given the central role of CD16, and we regard it as an informative finding rather than an artifact. As set out in our revised Discussion, the surface density of gp120 on HIV-infected primary T-cells is approximately 6.4 × 10<sup>2</sup> molecules per cell (Vasiliver-Shamis et al., 2008), two orders of magnitude below high-density antigens such as CD20 on Raji cells at approximately 5 × 10<sup>4</sup> molecules per cell (Lallemand et al., 2017). Antibody-dependent responses against HIV-infected cells therefore proceed under conditions of limiting antigen, and both CD56<sup>dim</sup> subsets face the same constraint. What differs between them is not the constraint but the NKG2D available to meet it: CD56<sup>dim</sup>CD16<sup>dim</sup> cells carry higher NKG2D before target contact and respond more strongly, despite their lower CD16. Consistent with a requirement for a second signal under these conditions, ADAM17 inhibition reduced CD56<sup>dim</sup>CD16<sup>dim</sup> degranulation even at 0 µg/mL VRC01 (Figs. 7B and 7C, now Figures 6B and 6C), where no antibody is present to engage CD16.

      Regarding the concern that this superiority reflects conversion of CD16<sup>bright</sup> cells into the CD16<sup>dim</sup> gate, we address this in full in the central response; the opposite responses of the two subsets to ADAM17 inhibition, in both degranulation magnitude (Fig. 7B, now Figure 6B) and serial degranulation (Fig. 9, now Figure 8), argue against it. As the reviewer recommends, we will directly compare the intrinsic antibody-dependent capacity of each subset using cells purified before target-cell stimulation, extending the pre-sorted approach of Figure 2 to the antibody-dependent setting across additional donors.

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      (1) As discussed in the public review, the majority of the functional assays should be repeated using purified NK cell subsets. This approach would eliminate the confounding effect of activation-induced CD16 downregulation and allow the intrinsic functional properties of each subset to be directly compared.

      We agree, and this is the central experimental commitment of our revision. As described in the central response, we will repeat the specific-lysis and antibody-dependent degranulation assays using NK cell subsets purified before target-cell exposure, so that the intrinsic functional properties of each subset are compared directly and are not subject to activation-induced changes in CD16 expression. We will also perform the serial degranulation assay on sorted subsets if cell yields permit; that assay requires four sequential labelling and washing steps, and the CD56<sup>dim</sup>CD16<sup>dim</sup> subset constitutes fewer than 5% of CD56 <sup>dim</sup> NK cells, so we cannot commit to it in advance of the sort. This extends the pre-sorted approach already used in Figure 2, which the reviewer identifies as the strongest evidence in the manuscript, across additional donors and across the functional readouts. These experiments require cell sorting and primary-cell work and will be completed within approximately six to eight weeks and provided with the revised manuscript.

      (2) The experiments performed with purified NK cell subsets in Figure 2 provide the strongest evidence supporting the authors' conclusion that CD56dimCD16dim NK cells exhibit greater direct cytotoxicity against HIV-infected target cells. These data are the most convincing in the manuscript because they are not confounded by post-assay changes in CD16 expression. However, unlike the other functional assays, no representative gating strategy or raw flow cytometry plots are provided, and the results appear to be based on a single representative experiment. Given the importance of these data to the manuscript's central conclusion, this experiment should be expanded to include biological replicates from additional donors, representative flow cytometry plots, and validation using additional HIV-1 infectious molecular clones.

      We appreciate the reviewer identifying the sorted-subset experiments in Figure 2 as the strongest evidence for our conclusion, and we agree these data warrant expansion. In the revised manuscript we will:

      (1) Expand the experiment to include biological replicates from additional donors, with all donors shown rather than a single representative experiment.

      (2) Provide the representative gating strategy and flow cytometry plots for the sorted subsets. The reviewer is correct that these should be included, and we will add them, including for the expanded experiments.

      (3) Validate the finding using additional HIV-1 strains. We note, for clarity, that the virus used throughout this study is a primary patient isolate (HIV-1<sup>SHM-1</sup>), as stated in the Materials and Methods, rather than an infectious molecular clone. We have now compared NK cell degranulation against autologous CD4<sup>positive</sup> T-cells productively infected with HIV-1<sup>SHM-1</sup>, with the X4-tropic infectious molecular clone HIV-1<sup>NL4-3</sup>, and with the R5-tropic laboratory-adapted strain HIV-1<sup>BaL</sup>, at three effector cell to target cell ratios. CD56 <sup>dim</sup> CD16 <sup>dim</sup> cells degranulated more than CD56 <sup>dim</sup>CD16<sup>bright</sup> cells against every virus at every ratio, in all nine comparisons at p < 0.0001. Both subsets responded less to HIV-1<sup>NL4-3</sup> and HIV-1<sup>BaL</sup> than to HIV-1<sup>SHM-1</sup>, and did so in proportion: the ratio of CD56 <sup>dim</sup>CD16 <sup>dim</sup> to CD56 <sup>dim</sup> CD16<sup>bright</sup> degranulation ranged from 2.7 to 3.8 across all nine conditions. The magnitude of the response therefore varies with the virus, whereas the relationship between the two subsets does not. These data are included in the revised manuscript as Figure 1—figure supplement 3, with the statistical analysis in a supplemental table.

      The experiments described in points 1 and 2 require cell sorting and primary-cell work and will be completed within approximately six to eight weeks and provided with the revised manuscript.

      (3) The mechanism underlying the enhanced effector function of CD56dimCD16dim NK cells remains unclear. Although the phenotypic characterization presented in Figure 5 (and related supplement figures) is informative, NK cell receptor expression was assessed after target cell stimulation, when it may already have been altered by activation and CD16 downregulation. Receptor expression should therefore be evaluated prior to stimulation. In addition to NKG2D, the authors should also consider assessing additional activating receptors, notably NKp30, which has recently been implicated in the elimination of autologous HIV-1-infected cells (PMID: 41079618).

      We agree that receptor expression should be assessed before stimulation, and in fact it is. In Fig. 6 (now Figure 5), the receptor gMFI data include the no-target condition, showing that CD56<sup>dim</sup>CD16<sup>dim</sup> cells express 1182 gMFI more NKG2D than CD56 <sup>dim</sup>CD16<sup>bright</sup> cells before any target-cell contact (p < 0.0001), approximately 1.6-fold, and therefore before any activation-induced change in receptor expression. NKp46, measured on the same cells in the same wells, did not differ between the subsets under the same condition (mean difference 45.67 gMFI, p = 0.0668), indicating that the difference is specific to NKG2D rather than a general difference in activating receptor density. The baseline condition is now labelled explicitly as 1:0 in the revised figure and described as such in the Results text.

      Regarding NKp30, we have assessed this receptor. Degranulation did not differ between NKp30 positive and NKp30 negative cells within either CD56<sup>dim</sup> subset, whereas both CD56<sup>dim</sup>CD16<sup>dim</sup> groups exceeded both CD56<sup>dim</sup>CD16<sup>bright</sup> groups regardless of NKp30 status, indicating that the enhanced degranulation of the subset is not attributable to NKp30, paralleling our finding for NKp46. These NKp30 data are included in the revised manuscript as Figure 5—figure supplement 1, with the corresponding statistical analysis in a supplemental table. We note that this analysis addresses whether NKp30 accounts for the difference between the subsets; it does not exclude a role for NKp30 in NK-cell recognition of HIV-infected cells more generally, consistent with the study the reviewer cites, which we now discuss.

      (4) In Figure 5, the histograms corresponding to the uninfected and ΔVpr conditions appear to be identical. If this is indeed the case, this represents a serious concern, as these are two distinct experimental conditions and should not be represented by the same flow cytometry plot. This raises the possibility of an inadvertent panel duplication. The authors should carefully verify the figure and replace the duplicated panel if necessary.

      We thank the reviewer for this careful observation. On review, the histograms in Figure 5A were reproduced from our earlier study (Ward et al., PLoS Pathog 2009;5(10):e1000613) and should not have been included without attribution. We have removed that panel and cite the original finding in its place.

      Figures 5B and 5C are both new results from this study, and both are retained. Figure 5B, which shows that NK cells degranulate in response to wild-type HIV-infected targets but not to ΔVpr-infected or uninfected targets, becomes Figure 4—figure supplement 1 in the revised manuscript. Figure 5C, the NKG2D blockade experiment, becomes Figure 4A and 4B.

      (5) The ADCC experiments and calculation require additional methodological clarification, particularly the analyses presented in Figure 7C. Although NK cell degranulation is commonly used as a surrogate marker of ADCC, the data presented in Figure 7 do not appear to isolate the antibody-dependent component of the response. To specifically quantify ADCC-mediated degranulation, the degranulation induced by HIV-infected target cells alone (i.e., in the absence of VRC01) should be subtracted from that measured in the presence of VRC01. Notably, in Figure 7C (DMSO), the CD56dimCD16dim population appears to exhibit similar levels of degranulation in the absence and presence of VRC01, suggesting that antibody-dependent degranulation may be limited in this subset, which does not support the author's conclusions.

      We thank the reviewer for raising this, and we agree that the antibody-dependent and antibody-independent components of the response should be distinguished. We would, however, respectfully argue against the subtraction as a means of doing so, and we believe the experiments already in the manuscript address the underlying question more directly.

      The condition without VRC01 is not a background to be removed. It is the NKG2D-driven response of the same cells to the same infected targets, measured through the same degranulation machinery, and it is one of the principal findings of the study. Subtracting it treats the two components as though they were independent and additive, when both converge on a single immunological synapse and a single degranulation event per cell. The difference between the two conditions is therefore not the antibody-dependent response; it is the increment in total degranulation produced by adding antibody, which is a different quantity and one that carries no clean interpretation at the level of the individual cell.

      The question the reviewer raises, whether the antibody-dependent component differs between the subsets, is answered directly by the two-way ANOVA of these data. Across the VRC01 titration, the effect of NK cell subset accounts for 85.91% of the total variation (F(1, 16) = 424.0, p < 0.0001) and the effect of VRC01 concentration for 10.18% (F(3, 16) = 16.74, p < 0.0001), while the subset × VRC01 interaction is not significant (F(3, 16) = 1.112, p = 0.3732) and accounts for 0.68% (Supplemental Table 20 in the revised manuscript). The absence of an interaction means that adding antibody raises the response of both subsets by a comparable amount, and that the difference between the subsets is the same at every VRC01 concentration tested. This is a statistical statement about the antibody-dependent component, obtained without subtracting one condition from another.

      We agree with the implication the reviewer draws from this, and we state it plainly in the revised Discussion: the antibody-dependent increment is modest in both subsets. We attribute this to the very low surface density of gp120 on HIV-infected primary T-cells, approximately 6.4 × 10<sup>2</sup> molecules per cell (Vasiliver-Shamis et al., 2008), which is two orders of magnitude below high-density antigens such as CD20 on Raji cells, approximately 5 × 10<sup>4</sup> molecules per cell (Lallemand et al., 2017). Under these conditions the antibody-dependent signal available to any NK cell is limited, and this applies equally to both subsets.

      Where we differ from the reviewer is on the conclusion this supports. That the antibody-dependent increment is modest in both subsets does not weaken our central claim, which is comparative: at every VRC01 concentration tested, including in the presence of antibody, CD56<sup>dim</sup>CD16<sup>dim</sup> cells degranulate more than CD56<sup>dim</sup>CD16<sup>bright</sup> cells against antibody-coated HIV-infected targets. That comparison is what the manuscript reports, and it is unaffected by how the response is partitioned between its antibody-dependent and antibody-independent components.

      We also note that the experiments in Figs. 7B and 7C (now Figures 6B and 6C) do isolate a component of the response experimentally rather than arithmetically. Inhibiting ADAM17 removes the contribution that depends on CD16 turnover, and it does so in opposite directions in the two subsets, reducing CD56<sup>dim</sup>CD16<sup>dim</sup> degranulation and increasing that of CD56<sup>dim</sup>CD16<sup>bright</sup> cells at every VRC01 concentration. These are direct experimental manipulations of the antibody-dependent pathway, and they are more informative than the arithmetic difference between two conditions.

      Finally, we take the reviewer's point that the analyses in Fig. 7C require clearer explanation. In the revised manuscript we state explicitly what is plotted, namely the percentage of each CD16 subset among CD107a positive CD56<sup>dim</sup> NK cells, we describe the background subtraction that applies to all CD107a data in this study, and we report the statistical analysis of each panel in full.

      (6) It is also unclear how the authors interpret the effects of ADAM17 inhibition. While ADAM17 inhibition increases the ADCC activity of the CD56dimCD16bright population, it simultaneously decreases that of the CD56dimCD16dim population. An alternative explanation is that inhibition of CD16 shedding prevents activated CD56dimCD16bright NK cells from transitioning into CD56dimCD16dim during the assay. This possibility should be discussed and experimentally addressed, as it provides a plausible alternative interpretation of the observed phenotype.

      We thank the reviewer for articulating this alternative, which we address in full in the central response. We agree that the opposite effects of ADAM17 inhibition on the two subsets are central to interpreting these experiments, and we interpret them as evidence that the two are distinct populations rather than one transitioning into the other.

      The reviewer's alternative, that ADAM17 inhibition prevents CD56<sup>dim</sup>CD16<sup>bright</sup> cells from transitioning into the CD56<sup>dim</sup>CD16<sup>dim</sup> gate, predicts that blocking shedding should reduce the CD56<sup>dim</sup>CD16<sup>dim</sup> population by cutting off its supply from CD16<sup>bright</sup> cells. Three observations argue against this being the explanation for the functional difference. First, the effect is not merely a change in population size but a change in per-cell function in opposite directions: ADAM17 inhibition increases the number of serial degranulation events in CD56<sup>dim</sup>CD16<sup>bright</sup> cells while decreasing them in CD56<sup>dim</sup>CD16<sup>dim</sup> cells (Fig. 9, now Figure 8), which is difficult to explain if the dim cells were simply bright cells prevented from converting. Second, blocking NKG2D together with ADAM17 reduces CD56<sup>dim</sup>CD16<sup>dim</sup> degranulation below either treatment alone (Fig. 8A, now Figure 7A); if NKG2D blockade acted only by reducing the shedding that drives the putative transition, the combination could not exceed the effect of ADAM17 inhibition alone. Third, we have tracked the fate of each subset sorted before target exposure: at one hour, when degranulation is maximal, only approximately 4% of sorted CD16<sup>bright</sup> cells were found in the CD16<sup>dim</sup> gate, while approximately 35% had moved to the CD16<sup>negative</sup> gate (Author response image 1 accompanying this response). Shedding therefore directs CD16<sup>bright</sup> cells past the CD16<sup>dim</sup> gate rather than into it.

      We discuss this alternative explicitly in the revised Discussion and will address it further experimentally by repeating these assays on subsets purified before target exposure, where no transition can occur during the assay.

      Reviewing Editor Comments:

      The conclusion that CD56dimCD16dim NK cells are intrinsically superior effectors against HIV-infected target cells requires additional evidence because CD16 is rapidly downregulated following NK-cell activation. Throughout most of the study, NK-cell subsets are classified after target-cell encounter, making it difficult to distinguish pre-existing CD56dimCD16dim cells from activated CD56dimCD16bright cells that have undergone ADAM17-mediated CD16 shedding. The authors should repeat functional experiments using NK-cell subsets purified before target-cell exposure and determine the extent to which ADAM17 inhibition alters subset frequencies and functional readouts. These experiments are essential to establish whether the observed functional differences reflect intrinsic biology rather than activation-induced phenotypic conversion.

      We thank the editor for this clear synthesis of the central concern, which we address in full in the central response above. In brief, our conclusion does not rest on post-encounter classification: the CD56<sup>dim</sup>CD16<sup>dim</sup> advantage is established in cells sorted into subsets before any target contact, using direct lysis as the readout (Fig. 2), and is reinforced by the opposite responses of the two subsets to ADAM17 inhibition in both degranulation magnitude (Fig. 7B, now Figure 6B) and serial degranulation (Fig. 9, now Figure 8), by the separable contributions of NKG2D and ADAM17 (Fig. 8, now Figure 7), and by pre-stimulation differences between the subsets, with NKG2D 1182 gMFI higher on CD16<sup>dim</sup> cells and no corresponding difference in NKp46 (Fig. 6, now Figure 5). We agree these questions are central and will repeat the functional experiments on subsets purified before target exposure, and the effect of ADAM17 inhibition on subset frequencies is presented explicitly in Figs. 7C and 8B (now Figures 6C and 7B).

      Major conclusions should be supported by more rigorous experimental validation. In particular, the sorted NK-cell experiments should be expanded using multiple independent donors, include killing of uninfected target cells as controls and provide representative gating strategies and flow cytometry plots. Likewise, the current analyses of killing frequency, serial killing, and ADCC should be strengthened by direct measurements using purified NK-cell subsets rather than mathematical estimates or analyses performed in mixed NK-cell populations.

      We agree and will strengthen the validation as follows. The sorted-subset experiments will be expanded across multiple independent donors, with all donors shown. Uninfected-target controls will be included for the sorted-cell lysis experiments (Fig. 2); the corresponding CD107a controls against uninfected targets are already presented in Figure 1—Figure Supplement 4C of the revised manuscript. Representative gating strategies and flow cytometry plots will be provided for the sorted-cell experiments, as for our other assays. Regarding direct measurement: the killing-frequency metric (Fig. 3, now Figure 2—figure supplement 1C) is a mathematical estimate whether derived from mixed or purified populations, and it has been moved to the supplementary material with this limitation noted in the Discussion, while direct, subset-resolved killing is provided by the pre-sorted lysis experiment (Fig. 2), which we will expand; the serial degranulation assay (Fig. 4, now Figure 3) is now described as serial degranulation rather than serial killing, and will be repeated on purified subsets if cell yields from the sort permit; and the antibody-dependent comparison will be performed on subsets purified before stimulation.

      Some aspects of the data analysis and presentation require clarification. The authors should evaluate receptor expression before target-cell stimulation, clarify the ADCC analyses and interpretation of ADAM17 inhibition, verify the apparent duplicated flow-cytometry panel, apply appropriate statistical corrections for multiple comparisons where necessary, and streamline the presentation by emphasizing the principal conclusions rather than extensive descriptive analyses.

      We have addressed each of these. Receptor expression before stimulation is shown in the gMFI data of Fig. 6 (now Figure 5) at the no-target condition, where CD56<sup>dim</sup>CD16<sup>dim</sup> cells carry 1182 gMFI more NKG2D than CD56<sup>dim</sup>CD16<sup>bright</sup> cells (p < 0.0001) with no corresponding difference in NKp46 (p = 0.0668); this condition is now labelled explicitly as 1:0 in the figure and described as such in the Results text. The antibody-dependent analyses and the interpretation of ADAM17 inhibition are clarified in the revised text, as detailed in our responses to the three reviewers and the central response.

      On the duplicated panel: the histograms in Figure 5A were reproduced from Ward et al. (2009) and should not have been included without attribution. That panel has been removed and the original finding is cited in its place. Figures 5B and 5C are both new results and are retained, becoming Figure 4—figure supplement 1 and Figures 4A and 4B respectively.

      We have applied appropriate multiple-comparison corrections, using the post-hoc test matched to each comparison structure and reporting adjusted p-values throughout; the design and post-hoc test used for each figure and panel are given in a new supplemental table. Finally, we have streamlined the presentation by opening each Results section with a statement of the principal finding before the supporting data, and by reducing the inhibitory receptor section from approximately 1,800 words and more than 80 reported p-values to approximately 700 words and 24, with the detailed data retained in the supplemental tables and their significance synthesized in the Discussion.

      References

      Amand M, Iserentant G, Poli A, Sleiman M, Fievez V, Sanchez IP, Sauvageot N, Michel T, Aouali N, Janji B, Trujillo-Vargas CM, Seguin-Devaux C, Zimmer J. 2017. Human CD56<sup>dim</sup>CD16<sup>dim</sup> cells as an individualized natural killer cell subset. Frontiers in Immunology 8:699. doi:10.3389/fimmu.2017.00699.

      Lallemand C, Liang F, Staub F, Simansour M, Vallette B, Huang L, Ferrando-Miguel R, Tovey MG. 2017. A novel system for the quantification of the ADCC activity of therapeutic antibodies. Journal of Immunology Research 2017:3908289. doi:10.1155/2017/3908289.

      Vasiliver-Shamis G, Tuen M, Wu TW, Starr T, Cameron TO, Thomson R, Kaur G, Liu J, Visciano ML, Li H, Kumar R, Ansari R, Han DP, Cho MW, Dustin ML, Hioe CE. 2008. Human immunodeficiency virus type 1 envelope gp120 induces a stop signal and virological synapse formation in noninfected CD4+ T cells. Journal of Virology 82:9445-9457. doi:10.1128/JVI.00835-08.

      Ward J, Davis Z, DeHart J, Zimmerman E, Bosque A, Brunetta E, Mavilio D, Planelles V, Barker E. 2009. HIV-1 Vpr triggers natural killer cell-mediated lysis of infected cells through activation of the ATR-mediated DNA damage response. PLoS Pathogens 5(10):e1000613. doi:10.1371/journal.ppat.1000613.

    1. eLife Assessment

      The cell wall and plasma membrane are tightly associated in plant cells, and even after plasmolysis, sites of strong contact remain between the plasma membrane and cell wall. Here, the authors provide solid data to implicate an Arabidopsis lectin-like receptor kinase in plasma-membrane-to-cell-wall contact sites, and propose that these attachment sites are of importance for plant resistance to hyperosmotic stress, yet the evidence remains incomplete without key controls. These important results broaden our view on dynamic modulation of plasma membrane topology and its physiological consequences.

    2. Reviewer #1 (Public review):

      This work identifies two lectin receptor-like proteins as cell wall -plasma membrane anchors that contribute to persistent attachment sites maintained during plasmolysis. The authors propose that these attachment sites are important for plant resistance to hyperosmotic stress. Although the existence of persistent attachment sites between the cell wall and the plasma membrane, particularly evident in plasmolysed cells, was recognised long ago, the molecular components tethering the two cellular components remain largely unknown. Therefore, the findings presented by Arico et al. address an important question in plant cell biology.

      Through a screening of potential anchors, they found that the overexpression of fluorescently tagged LecRK-I.9, LecTM, AGP18 and AT14A in Nicotiana benthamiana increased the density of Hechtian strands in plasmolysed cotyledon epidermis cells. They show that for LecRK-I.9* (* indicates kinase-dead version) and LecTM, this effect depends on the Lectin domain. Focusing on LecRK-I.9*, the overexpressed Lectin domain localized to cell walls, accumulating in certain foci. The authors interpret this as a possible preference for certain cell wall composition. I find these results convincing and the methodology robust.

      The second part of the manuscript, however, relies on interpretations that, in my opinion, are not fully supported by the presented evidence. The authors move to Arabidopsis thaliana and show that overexpression of both LecRK-I.9* and LecRK-I.9*ΔLec fluorescent reporters also localizes to the plasma membrane and Hechtian strands in plasmolysed cells, although the density of Hechtian strands is not quantified. Thus, it is unclear whether LecRK-I.9* promotes Lectin domain-dependent strong cell wall-plasma membrane attachment sites in Arabidopsis.

      The authors focus on the formation of big signal clusters at the plasma membrane in response to hyperosmotic treatment. They studied the dynamics of the clusters upon treatment application and observed increased mobility of LecRK-I.9* compared to LecRK-I.9*ΔLec and a plasma membrane marker. They then analysed the abundance and size (not the mobility) of these clusters on different plasma membranes facing cell walls that have or are predicted to have different mechanical and chemical properties, identifying differences between LecRK-I.9* and LecRK-I.9*ΔLec. Although the reduced mobility of LecRK-I.9 relative to LecRK-I.9ΔLec is consistent with an interaction between the lectin domain and the cell wall, it does not by itself demonstrate that the observed clusters correspond to CW attachment sites. Moreover, the relation between the clusters and Hechtian strands (bona fide cell wall-plasma membrane attachments) is not explored. Likewise, the differential clustering observed on different cell faces is intriguing but could have different interpretations.

      Finally, the physiological relevance of the proposed anchoring mechanism is supported by osmotic stress assays, but the scoring method relies on manual classification of resistant seedlings and could benefit from a more objective quantitative readout.

      In summary, although I find all these results valuable, I find that the methodology is not completely adequate and that several aspects of the data interpretation require additional support or clarification before the central conclusions can be fully justified.

    3. Reviewer #2 (Public review):

      Summary:

      The manuscript submitted by Arico and co-workers describes the impact of two lectin-domain-containing proteins (LecRK-I.9* and LecTM) on the formation and persistence of Hechtian strands. Based on a survey of selected candidates, overexpression of these two proteins resulted in an increase in Hechtian strand formation. Removal of the lectin domains and expression of this variant did not alter HS formation compared to the WT. In addition, the existence of the lectin domain reduced protein mobility, probably due to interactions with the wall. Last but not least, overexpression of LecRK-I.9 increased the resistance of plants towards water loss conditions.

      Strengths:

      The study seems well conducted, but may require some small additions. While the results themselves seem not surprising, I think that this is a valuable demonstration of the cell wall binding ability of lectin proteins and its physiological and microscopical consequences.

      Weaknesses:

      At this stage, some of the study would benefit from some additional quantification.

    4. Reviewer #3 (Public review):

      Summary:

      The cell wall and plasma membrane are tightly associated in plant cells, but even after plasmolysis, sites of strong contact remain between the plasma membrane and cell wall. Several hypotheses have been introduced about the molecular makeup of these sites of PM-CW adhesion (e.g., Rui et al 2026 Cell; Qin et al 2026 Current Biol; Pérez-Sancho et al 2025 Cell). Here, the authors implicate two transmembrane lectin proteins in PM-CW adhesion via overexpression in Nicotiana benthamiana and via Arabidopsis knockout phenotypes for one of these candidates, the lectin receptor kinase LecRK-1.9. They further show that the PM-CW adhesion function of LecRK-1.9 requires the extracellular domain, suggesting that this lectin-like domain may interact with the cell wall. Interestingly, LecRK-1.9 has also been implicated in extracellular ATP binding in the context of biotic and abiotic stress responses (e.g., Choi et al 2014 Science).

      Strengths:

      Overall, the work is carefully conducted with high-quality imaging and quantitative image analysis. The results present an interesting candidate for future studies of plasma membrane-to-cell-wall attachment.

      Weaknesses:

      There are two major caveats to this work. First, all work was conducted with the kinase-dead version of LecRK-1.9, which eliminates a significant biological function of this protein, as evidenced by the major differences in expression pattern of wild-type vs kinase-dead LecRK-1.9. Second, controls are essential to document the expression levels of different protein variants and controls, since LecRK-1.9 expression is correlated with Hechtian strand density.

    1. eLife Assessment

      This study presents a valuable metagenomic analysis of the gut microbiome in sickle cell disease (SCD) patients, revealing associations between bacteriophage, host immunity, and SCD pathophysiology. This solid dataset and in-depth analysis will be of broad interest to researchers studying SCD, immunology, phage biology, and the microbiome, adding to the small but growing literature suggesting a microbial component to SCD.

    2. Reviewer #1 (Public review):

      Summary:

      In this manuscript, Flamholz and colleagues use metagenomic sequencing to profile the microbiome of individuals with sickle cell disease (SCD), the most common genetic blood disorder in the world. To build on previous studies that found dysbiosis in SCD, this manuscript aims to examine whether changes in either bacterial species or bacteriophages correlate with inflammatory hallmarks of the disease. The authors claim that sickle cell dysbiosis does not correlate with inflammatory hallmarks of the disease including aged neutrophil numbers, a cell type previously highlighted in preclinical sickle cell microbiome work.

      Strengths:

      The primary strength of this paper is the investigation into disease associated changes in bacteriophages. This is an entirely novel idea in the sickle cell field, and based on the current results, may be an important, under-recognized disease hallmark. It is unclear, however, if phages are "the chicken or the egg" in terms of sickle cell inflammatory profiles; do these increases in phage number simply result from other disease process or are they in anyway contributing to disease pathophysiology?

      Weaknesses:

      The authors addressed many of the initial manuscript weaknesses in their revision. In particular, they have softened language regarding sickle cell dysbiosis and its causative role in disease pathology. This is particularly appropriate given the lack of correlation between dysbiosis and immune factors in this single-center study.

    3. Reviewer #2 (Public review):

      Summary:

      The study analyzes stool metagenomes from 98 SCD patients and 46 controls, with SCD and control groups matched on age, race, sex, and ethnicity. The authors report lower Shannon diversity, lower Firmicutes/Bacteroidetes ratio, loss of health-associated taxa, increased disease-associated indicators, altered butyrate/fatty-acid metabolism pathways, and enrichment of provirus/prophage fractions in SCD. They further correlate aged-like neutrophils and prophage fractions with inflammatory cytokines. The main strength is that this is not just another 16S comparison. The use of whole-community metagenomics, immune profiling, neutrophil assays, and clinical metadata makes the study more biologically interesting than prior small SCD microbiome papers. The main weakness is that the causal and mechanistic interpretation is too strong. The data support an association between SCD status and microbiome/virome features, but they do not yet establish a clear "axis of pathophysiology." The provirus findings are intriguing, but require stronger statistical control, better validation, and more cautious interpretation.

      Strengths:

      The major strengths of the study include the clinically relevant disease setting, the use of whole-community sequencing, the integration of microbial, immune-cell, cytokine, and clinical measurements, and the novel attention to bacterial virus-related features. A particularly interesting aspect of the work is the analysis of virus-like elements integrated into bacterial genomes. The authors report that these elements are enriched in the gut microbial communities of patients with sickle cell disease and are associated with several inflammatory signals in blood. This observation is potentially important because it suggests that the microbial contribution to inflammation in sickle cell disease may involve not only bacteria but also bacterial virus-related genetic elements.

    4. Reviewer #3 (Public review):

      Summary:

      In this manuscript, Flamholz et al. sought to determine whether consistent and significant interactions exist between the gut microbiome and disease pathology in sickle cell disease (SCD). By sequencing and analysing metagenomes from faecal samples collected from 98 SCD patients and 46 control subjects, they identified community-level shifts in both the bacterial and proviral gut microbiome of SCD patients. They further reported correlations between the proviral microbiome and multiple blood cytokines, whereas similar associations were not observed for the bacterial microbiome.

      Strengths:

      This work includes the largest SCD cohort analysed to date, enabling analysis with relatively strong statistical power. In addition to profiling the bacterial microbiome, the study also examines the gut proviral microbiome, thereby providing a more comprehensive investigation of the topic. The newly generated metagenomic dataset will also be valuable for further meta-analysis by the wider community. Overall, the authors have largely achieved their aims.

      Weaknesses:

      This study represents a single-centre cross-sectional investigation, and most findings remain correlative in nature. Additional mechanistic and/or longitudinal evidence would be required to unravel causality and the underlying mechanism.

    5. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      This study presents a valuable metagenomic analysis of the gut microbiome in sickle cell disease (SCD) patients, revealing associations between bacteriophage, host immunity, and SCD pathophysiology. While these data are interesting and helpful for hypothesis generation, they are deemed incomplete; additional experiments would be needed to test causality and to provide mechanistic insight. Despite these limitations, this work will be of broad interest to researchers studying SCD, immunology, phage biology, and the microbiome, adding to the small but growing literature suggesting a microbial component to SCD.

      The authors would like to thank the reviewers for thorough and constructive comments on our manuscript. We have made major updates to the manuscript addressing the following points and suggestions from the three reviewers: (1) assessing HbAS/AA genotype influence on microbiome composition; (2) conducting the requested beta diversity analysis, (3) conducting the requested sensitivity analysis to assess the impact of disease severity and therapy on microbiome and virome features; (4) modifying our language to clearly state that our results do not indicate causality or mechanism of microbiome interactions with sickle cell disease pathophysiology; (5) improved discussion of the phage results and their strengths and limitations; (6) additional changes throughout for clarity and correction of errors. We have changed the title to “Bacterial and viral gut microbiome alterations characterize microbiome-immune-pathophysiology axes in Sickle Cell Disease.” These additions have greatly improved our work and presentation and we are grateful to the reviewers and our editors. We have indicated where specific changes were made in response to the public reviews below.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      In this manuscript, Flamholz and colleagues use metagenomic sequencing to profile the microbiome of individuals with sickle cell disease (SCD), the most common genetic blood disorder in the world. To build on previous studies that found dysbiosis in SCD, this manuscript aims to examine whether changes in either bacterial species or bacteriophages correlate with inflammatory hallmarks of the disease. The authors claim that sickle cell dysbiosis does not correlate with inflammatory hallmarks of the disease, but instead, aged neutrophil numbers and bacteriophages do. Appropriate control subjects and additional analyses are needed to support that conclusion.

      Strengths:

      The primary strength of this paper is the investigation into disease-associated changes in bacteriophages. This is an entirely novel idea in the sickle cell field, and based on the current results, may be an important, under-recognized disease hallmark. It is unclear, however, if phages are "the chicken or the egg" in terms of sickle cell inflammatory profiles; do these increases in phage number simply result from other disease processes, or are they in any way contributing to disease pathophysiology?

      Weaknesses:

      A primary weakness of the manuscript is the fact that the majority of individuals included in the control group maintain sickle cell trait (HbAS genotype). Although typically asymptomatic, it is unclear if this genotype is associated with microbial changes that would not be observed in a true control group (HbAA genotype). This is a significant limitation that may limit the ability to draw conclusions from the current data set.

      Another key weakness is the lack of beta diversity assessment. Although decreased alpha diversity is observed in individuals with SCD, and specific bacterial taxa are differentially abundant following multivariate analyses, there is no overall comparison of bacterial community composition between individuals with SCD and controls. Prior to drawing conclusions about the relationship (or lack thereof) between the SCD microbiome and inflammatory markers, it is important to know if this study did indeed find disease-associated changes in microbiome composition.

      It is unclear which individuals were used for aged neutrophil (AN) and molecular data assessments. For example, were children who were still receiving penicillin prophylaxis included in these specific assessments? Given the authors' previous work demonstrating that antibiotic treatment decreases AN pathology, it seems critical to limit all AN/molecular analyses to older subjects who are not on daily penicillin treatment (if possible).

      A minor weakness is the continued use of "disease" vs. "healthy" indicators as primary microbiome metrics that are used for molecular correlations. The lack of metric specificity - and lack of discussion regarding which diseases were used to generate these indicators (how similar/different are they to sickle cell?) - could be said to make these metrics essentially meaningless.

      We thank the reviewer for their helpful comments and suggestions. We want to first note that patients on prophylactic penicillin within six months of sample collection were excluded from the study due to the known impact of antibiotics on gut microbiomes, this has been clarified in the main text. We have now included an analysis evaluating the influence of control genoype (HbAA/HbAS) on our microbiome and virome results. To evaluate whether control genotype influenced major microbiome and virome features, analyses were restricted to control participants only. Controls were stratified by genotype as HbAA or HbAS. Four significant microbiome and virome features were tested: F:B ratio, Shannon diversity, provirus fraction, and virus count. HbAA and HbAS controls were compared using two-sided Mann-Whitney U tests. Benjamini-Hochberg FDR correction was applied across the four tested features. HbAS and HbAA controls did not differ significantly for F:B ratio, Shannon diversity, provirus fraction, or virus count. The inclusion of HbAA/AS strengthens our results with respect to the observation that sickle cell disease patient microbiomes remain significantly different from sickle trait (HbAS) controls. These results are reported in the new Supplemental Table 6.

      We have now included a beta diversity analysis using MetaPhlAn species profiles. Beta diversity analyses were performed in Python using pandas and NumPy for data processing, scikit-bio for distance calculations and PERMANOVA, scikit-learn for ordination-related computations, statsmodels for multiple-testing correction where applicable, and matplotlib for visualization.

      For the primary disease/control comparison, samples were grouped as control or SCD. For the genotype control sensitivity analysis, samples were restricted to HbAA and HbAS individuals as described above. Species detected in at least 10% of included samples were retained for beta diversity analysis. To account for the compositional structure of metagenomic relative abundance data, species profiles were transformed using a centered log-ratio transformation after addition of a small pseudocount to accommodate zero values. Aitchison distances were calculated from the CLR-transformed species profiles. Statistical significance of group separation was assessed by PERMANOVA using 999 permutations. For the control versus SCD comparison, PERMANOVA was performed between the two disease-status groups. For the HbAA versus HbAS control comparison, PERMANOVA was performed among controls only.

      In the SCD cohort, beta diversity differed significantly between controls and SCD participants by Aitchison distance after CLR transformation (R<sup>2</sup> = 0.030, p = 0.001). In contrast, HbAA and HbAS controls did not differ significantly in beta diversity (R<sup>2</sup> = 0.024, p = 0.282), supporting the conclusion that the observed SCD/control separation was not driven by control genotype composition. These methods and results are now reported in the manuscript.

      The manuscript describing the microbiome health and disease indicators was submitted to eLife jointly with this manuscript as a package; eLife declined to review the indicator manuscript. Briefly, this study conducted a cross-disease meta-analysis of 38 studies comprising 8,204 samples and identified 100 bacterial taxa or “indicators” that are weakly but consistently associated with health or disease across diverse conditions, including, but not limited to, inflammatory bowel disease, colorectal cancer, type 2 diabetes. The indicator taxa were validated in an independent cohort of Graves’ disease patients. We currently cite an older version of this work posted as a preprint. The manuscript is currently under review at another journal and we will update this manuscript with the updated citation when it is available.

      We have addressed other recommendations from this reviewer as follows. We cite and discuss previous SCD rodent model work observing decreased butyrate in disease, and we have updated our results and discussion sections regarding associations between the microbiome and virome and clinical and molecular features.

      Reviewer #2 (Public review):

      Summary:

      The study analyzes stool metagenomes from 98 SCD patients and 46 controls, with SCD and control groups matched on age, race, sex, and ethnicity. The authors report lower Shannon diversity, lower Firmicutes/Bacteroidetes ratio, loss of health-associated taxa, increased disease-associated indicators, altered butyrate/fatty-acid metabolism pathways, and enrichment of provirus/prophage fractions in SCD. They further correlate aged-like neutrophils and prophage fractions with inflammatory cytokines. The strength is that this is not just another 16S comparison. The use of whole-community metagenomics, immune profiling, neutrophil assays, and clinical metadata makes the study more biologically interesting than prior small SCD microbiome papers. The main weakness is that the causal and mechanistic interpretation is too strong. The data support an association between SCD status and microbiome/virome features, but they do not yet establish a clear "axis of pathophysiology." The provirus findings are intriguing, but require stronger statistical control, better validation, and more cautious interpretation.

      Strengths:

      The major strengths of the study include the clinically relevant disease setting, the use of whole-community sequencing, the integration of microbial, immune-cell, cytokine, and clinical measurements, and the novel attention to bacterial virus-related features. A particularly interesting aspect of the work is the analysis of virus-like elements integrated into bacterial genomes. The authors report that these elements are enriched in the gut microbial communities of patients with sickle cell disease and are associated with several inflammatory signals in blood. This observation is potentially important because it suggests that the microbial contribution to inflammation in sickle cell disease may involve not only bacteria but also bacterial virus-related genetic elements.

      Weaknesses:

      The evidence for this proposed immune-related mechanism is incomplete. The study is cross-sectional and largely based on associations, so it cannot determine whether these virus-like elements drive immune activation, reflect immune activation, or are linked indirectly through disease severity, treatment history, or other clinical factors. The main limitations are the single-center design, modest sample size for some immune measurements, limited ability to control for treatment and disease heterogeneity, and the need for clearer multiple-testing correction in the correlation analyses. In particular, stronger adjustment for available clinical factors such as hydroxyurea use, transfusion history, pain admissions, genotype, and other markers of disease burden would help readers judge how specific the microbial and viral findings are to sickle cell disease itself. (REVISION POINT 3)

      Overall, the authors largely achieve their descriptive aim of identifying gut microbial differences associated with sickle cell disease. The evidence is solid for the presence of broad microbial community differences, but incomplete for the stronger conclusion that virus-like elements form a pathophysiological immune axis. The work will likely be useful to researchers studying the microbiome, inflammation, and sickle cell disease, especially as a hypothesis-generating dataset. Its impact would be strengthened by more cautious interpretation, stronger control of clinical confounders, clearer statistical correction, and future longitudinal or experimental studies to test causality.

      We thank the reviewer for their helpful comments and suggestions. We want to first note that patients on prophylactic penicillin within six months of sample collection were excluded from the study due to the known impact of antibiotics on gut microbiomes, this has been clarified in the main text. We have tempered our interpretation of our results, making clear that we are not arguing that either prophages or bacteria are causal or mechanistically associated with SCD biology and pathology. We have strengthened our control of clinical confounders, and added clearer statistical correction, as described below, with corresponding updates to the manuscript. We look forward to conducting future studies to test causality and understand mechanism.

      We have now done sensitivity analysis within SCD patients to determine whether our microbiome and virome results associate with treatment and clinical severity. To evaluate whether microbiome and virome features were explained by clinical or demographic heterogeneity within the SCD cohort, we restricted analyses to SCD participants. We fit a separate multivariable regression model for each feature. Each model included age, sex, hydroxyurea use, transfusions in the past year, and acute care utilization in the past year as predictors.

      feature_z <sup>~</sup> age_z + sex_F + HU + log1p(TxPastYr)_z + log1p(AcuteCarePastYr)_z

      Non-negative abundance, ratio, pathway, viral, and count-like variables were log-transformed to reduce skew, using feature-specific pseudocounts for zero-containing microbiome/virome variables and ln (1 + x) transformation for count covariates. Diversity and indicator scores were not log-transformed. Continuous variables were then standardized to Z-scores before modeling. Models were fit using ordinary least squares with HC3 robust standard errors. FDR correction was applied separately for each model term across the tested microbiome and virome features. No microbiome or virome feature showed an FDR-significant association with hydroxyurea use, transfusions in the past year, or acute care utilization in the past year. These results are reported in the manuscript and in the new Supplemental Table 7.

      We have reported multiple-testing correction results for all associations between microbiome and virome features and clinical and molecular features and updated manuscript figures accordingly.

      To evaluate relationships between microbiome/virome features and clinical or immune markers within the SCD cohort, we performed Spearman correlation analyses. Microbiome and virome features were organized into four prespecified feature groups: community metrics, taxa, functional pathways, and viral features. Clinical and immune markers were grouped into marker sets for visualization and multiple-testing correction, including hematologic clinical markers, hemolysis markers, creatinine, acute care burden, and cytokines/chemokines. Spearman correlation coefficients were calculated for each feature–marker pair. Benjamini-Hochberg FDR correction was applied within each prespecified feature group by marker group block. Nominal associations were defined as p < 0.05, FDR-significant associations as q < 0.05, and trends as q < 0.10. In the heatmap figure, boxes now indicate nominal p < 0.05, asterisks indicate q < 0.05, and daggers indicate q < 0.10.

      We have addressed other recommendations from this reviewer as follows. We have modified our language describing prophage/immune associations. We have revised the Methods to clarify how abundance data were processed before MaAsLin2 modeling. Taxonomic profiles were analyzed using MaAsLin2 with total-sum scaling normalization and log transformation, while pathway profiles were analyzed without additional normalization because the input pathway table had already been normalized prior to MaAsLin2 analysis; MaAsLin2 log transformation was then applied. We agree that relative abundance metagenomic data are compositional, and we have revised the text to clarify that these analyses identify covariate-adjusted associations with transformed relative abundance rather than absolute abundance. We also now note this as a limitation of the study. We also note the limitations of F:B as a metric. We have added text describing the need for further analysis of the prophages to understand their patterns of host range and transmission and their associations with features such as shared geography, health care exposure, and diet. Finally, we have updated Figure 5 to reflect our updated analysis with significance indicated.

      Reviewer #3 (Public review):

      Summary:

      In this manuscript, Flamholz et al. sought to determine whether consistent and significant interactions exist between the gut microbiome and disease pathology in sickle cell disease (SCD). By sequencing and analysing metagenomes from faecal samples collected from 98 SCD patients and 46 control subjects, they identified community-level shifts in both the bacterial and proviral gut microbiome of SCD patients. They further reported correlations between the proviral microbiome and multiple blood cytokines, whereas similar associations were not observed for the bacterial microbiome. Based on these findings, the authors propose the existence of a viral-immune axis in SCD pathophysiology and targetable functional alterations in the gut microbiome.

      Strengths:

      This work includes the largest SCD cohort analysed to date, enabling analysis with relatively strong statistical power. In addition to profiling the bacterial microbiome, the study also examines the gut proviral microbiome, thereby providing a more comprehensive investigation of the topic. The newly generated metagenomic dataset will also be valuable for further meta-analysis by the wider community. Overall, the authors have largely achieved their aims.

      Weaknesses:

      However, this study represents a single-centre cross-sectional investigation, and most findings remain correlative in nature. In particular, the claim that the study identifies targetable functional alterations in the gut microbiome for disease treatment may be somewhat overstated. Although the reported functional module changes in SCD patients are intriguing, additional mechanistic and/or longitudinal evidence would be required before these features can realistically be considered targetable.

      We thank the reviewer for their helpful comments and suggestions. We have now noted in the text that additional mechanistic and longitudinal studies are required before we can target the microbiome and virome in SCD and clarified that this is a single-centre, cross-sectional. We have further made modifications to the manuscript to clarify cohort features (specifically, age and race were matched, other baseline characteristics were balanced), to properly describe the Shannon diversity metric, and to fix several errors that this reviewer caught.

    1. eLife Assessment

      This valuable study leveraged high-definition transcranial direct current stimulation to the left dorsolateral prefrontal cortex to provide strikingly large effects on procrastination behavior over an extended time span. The cross-sectional, longitudinal study and testing of competing models provided solid evidence for how stimulating dlPFC impacts reveal-world procrastination behavior. Whether these results generalize to a larger population will be a significant future direction. This work will be of interest to those interested in cortical function, procrastination, and related states.

    2. Reviewer #4 (Public review):

      Summary:

      The current study tested the effects of repeated sessions of tDCS targeting the DLPFC on procrastination behavior. The main outcome is that anodal versus sham DLPFC tDCS reduces procrastination behavior on both a short-term and a long-term scale up to six months after the stimulation sessions.

      Strengths:

      The current study tests competing models of procrastination with state-of-the-art high-definition transcranial electric stimulation. The study assesses stimulation effects on procrastination on both a short-term and a long-term scale, suggesting that repeated stimulation of the prefrontal cortex reduces procrastination on a time scale of up to six months.

      Comments on revised version.

      The manuscript has already been reviewed and revised before, and it seems that the quality of the manuscript has substantially improved as a result of this revision process. I agree with the other reviewers that one must be cautious with drawing conclusions regarding the cognitive mechanisms underlying this effect, as many different cognitive functions are implemented by the DLPFC. The effect sizes are surprisingly large, but I am satisfied with the reasons provided by the authors for the large effect sizes.

      The authors successfully addressed my previous concerns on the manuscript.

    3. Author response:

      The following is the authors’ response to the previous reviews.

      Public Reviews:

      Reviewer #1 (Public review)

      Summary:

      The authors report the results of a tDCS brain stimulation study (verum vs sham stimulation of left DLPFC; between-subjects) in 46 participants, using an intense stimulation protocol over 2 weeks, combined with an experience-sampling approach, plus follow-up measures after 6 months.

      Strengths:

      The authors are studying a relevant and interesting research question using an intriguing design, following participants quite intensely over time and even at a follow-up time point. The use of an experience-sampling approach is another strength of the work.

      Comments on revised version.

      With the last round of revisions, the authors have now addressed my concerns.

      Thank you to re-review this revision, and we all appreciate you kindly contributing to substantially improve the conceptualization, statistics and statements for this manuscript.

      Reviewer #4 (Public review):

      Summary:

      The current study tested the effects of repeated sessions of tDCS targeting the DLPFC on procrastination behavior. The main outcome is that anodal versus sham DLPFC tDCS reduces procrastination behavior on both a short-term and a long-term scale up to six months after the stimulation sessions.

      Strengths:

      The current study tests competing models of procrastination with state-of-the-art high-definition transcranial electric stimulation. The study assesses stimulation effects on procrastination on both a short-term and a long-term scale, suggesting that repeated stimulation of the prefrontal cortex reduces procrastination on a time scale of up to six months.

      Weaknesses:

      The manuscript has already been reviewed and revised before, and it seems that the quality of the manuscript has substantially improved as a result of this revision process. I agree with the other reviewers that one must be cautious with drawing conclusions regarding the cognitive mechanisms underlying this effect, as many different cognitive functions are implemented by the DLPFC.

      We do appreciate you to take valuable time offering those insightful and helpful comments on this revised manuscript. As you kindly raised, this revision has redrawn conclusions and statements on domain-specific mechanistic roles of DLPFC in interpreting why this neuromodulation treatments are effective.

      One aspect of the current results that puzzles me is the strength of the current stimulation effects. Meta-analyses suggest that tDCS shows only small-to-moderate effect sizes (with Cohen's d around 0.5). While the authors report no effect sizes for their statistical models, the small p values, in combination with the unusually small sample size of 18 participants per group, suggests that the effect size must be rather large. Can the authors provide an estimate of the effect size of their stimulation effects? If they are considerably larger than to be expected, could the authors give an explanation for why their stimulation setup is showing much stronger effects than comparable high-definition tDCS studies on cognition or decision making?

      Thank you for raising this very crucial question in effect size determination. We fully understand that this large effect size makes you puzzled, and that the limited sample size indeed attenuates detectable power in the statistics. We completely agree that reporting the actual effect sizes is essential for interpreting the magnitude of our findings, and we appreciate the opportunity to clarify why the observed effects in our study appear substantially larger than the small-to-moderate effect sizes typically reported in meta-analyses of tDCS studies on cognition and decision-making.

      Following your suggestion, we have calculated the effect sizes for our primary outcomes. Based on the simple effect analyses (pre- vs. post-neuromodulation within the active neuromodulation group), we computed Cohen’s d for the within-group changes: for task-execution willingness, d = 2.37 (95% CI [1.49, 3.25]); for the actual procrastination rate, d = 1.52 (95% CI [0.87, 2.16]).

      We acknowledge that these effect sizes are considerably larger than the typical d ≈ 0.5 reported in the tDCS literature. After careful consideration, we attribute this discrepancy to three methodological and conceptual differences between our study and typical cognitive/decision-making tDCS studies. First, the small-to-moderate effect sizes are observed in studies utilizing a single-session tDCS protocol, yet our study employed an intensive 7-session HD-tDCS protocol over 15 days. Therefore, multi-session tDCS that induces cumulative, activity-dependent long-term potentiation (LTP)-like plasticity may substantially amplify and consolidates behavioral effects compared to single-session stimulation (Ke et al., 2023; Zhong et al., 2021). Second, most tDCS studies on cognition recruit healthy young adults who often perform near ceiling on laboratory tasks, leaving little "room for improvement" and thereby constraining the observable effect size. In our study, we strictly screened for severe chronic procrastinators. Because our participants had severe baseline deficits in task execution, the "ceiling space" for behavioral improvement was much larger, naturally inflating the observable effect size of the intervention. Lastly, given all the procrastinators completed tasks in the last session (0% procrastination rate in the active group, without within-group variance), mathematically, this boundary variable (0% vs 100%) artificially inflates the effect size estimate when calculating Cohen’s d with a near-zero post-test standard deviation.

      Nevertheless, as a sensitivity analysis, those findings are confirmed by Beta regression model addressing the risks of boundary variables, indicating that the potential inflation of effect sizes is statistically acceptable:

      In summary, while the observed Cohen's d values are unusually large, they are contextually justified by the cumulative nature of our multi-session protocol, the targeted clinical-like population, and the mathematical properties of bounded behavioral metrics.

      Results Section (Page 9, Line 441-444)

      “... For procrastination willingness, results showed a statistically significant interaction effect between multi-session neuromodulations and groups (β = -7.84, SE = 1.80, t = -4.36, DF = 45.6, p < .001, Cohen d = 2.37, 95% CI: 1.49-3.25; Fig. 3A and Fig. S2a).”

      Results Section (Page 9, Line 454-458)

      “... Similarly, a statistically significant interaction effect was identified here (β = -7.37, SE = 2.40, t = -3.02, DF = 46.6, p = .004, Cohen d = 1.52, 95% CI: 0.87-2.16), and the simple effect analysis further revealed decreased actual procrastination rates after ms-tDCS in the active neuromodulation group.”

      Regarding the strengths of the stimulation effects, I moreover found remarkable that the post-test procrastination rate was 100% in all (!) participants in the DLPFC group (figure 3F). I admit that it is hard to trust results that have no individual variation at all. This means that all participants are perfect responders to tDCS, which is again at variance what one typically expects for tDCS (where one usually has many non-responders). Do the authors have an explanation for this?

      Thank you for raising this highly important and helpful comment. Indeed, we fully understand that this result (a 100% task complete rate among all participants in the DLPFC group) is somewhat extraordinary. This pattern was equally striking to us when unblinded the data. After carefully scrutinizing the data and statistics, we are thrilled to confirm that this pattern is true. In support of this observation, we were gratified to receive numerous thank-you letters from participants who engaged in active neuromodulation. They expressed gratitude to us, and reported that they have substantially ameliorated procrastination behavior in real-life activities after completing the trial. While this does not constitute formal scientific evidence, we are also glad to see the benefits of this neuromodulation for those procrastinators.

      Two reasons could account for this pattern herein. One interpretation is to attribute this pattern to “floor effect”. In the present study, the procrastination rate was calculated as 1 minus the task-completion rate (e.g., 80%, 60%, 40%) by the deadline. At last stimulation sessions (#6 and #7), all the participants completed their real-life tasks before the deadline, yielding a 0% (1 minus 100% completion rate) procrastination rate, without any between-individual variation. Thus, rather than there being no individual variation in procrastination, this scalar – the procrastination rate - is too insensitive to capture subtle differences per se. For instance, although participants #1 and #2 both showed a 0% procrastination rate - meaning that both completed their tasks before the deadline - Participant #1 might have completed it 3 hours before the deadline, whereas Participant #2 might have completed it only 10 minutes before. In this case, the “scalar inflation” emerges to let us perceive that both participants have equivalent procrastination rates, although participant #2 may have a higher procrastination level than #1. As conceptually defined in the field, procrastination is contextualized as “not completing a task before the deadline”. Thus, if this task is completed before the deadline, regardless of whether it was finished close to or far in advance of the deadline, this case is defined as “no procrastination”. In the present study, the primary outcome is whether a participant procrastinated on a real-life task before the deadline in real-world settings, irrespective of when she/he completed this task. Thus, this scalar - procrastination rate - fits our conceptualization of procrastination.

      Another reason is the potential accumulative effects from sequential multi-session tDCS stimulation, as we explained above. As shown in Mann-Kendall trend tests, the procrastination rates show a significant linear downtrend in the active neuromodulation group across sessions, even after removing sessions #6 and #7. This indicates that the improvements of going against procrastination may be sequentially accumulative along with the increase in sessions, implying a potential “dose-dependent effect”. Despite a speculative interpretation, this “dose-dependent effect” in neuromodulation has been well-documented in previous studies, showing the robustly linear association between the number of sessions and effectiveness (c.f., Cole et al., 2020; Hutton et al., 2023; Sabé et al., 2024; Schulze et al., 2018). Therefore, although this extreme pattern is somewhat extraordinary compared to previous observations, it makes sense.

      We also conducted robustness check by removing sessions #6, #7, and both, to validate whether this results were biased by “scalar inflation”. We do believe that this analysis could support statistical robustness to go against potential biases from extreme cells. By doing so, we found that all the group*treatment_day interaction effects remained significant when removing either session #6 or session #7 (or even both, all p-values < .05), indicating high statistical robustness. Please see Table S3 and Table S4.

      Taken together, in spite of their being extraordinary, we confirm that those findings are statistically robust to extreme outliers. As you kindly suggested, we have added those findings of the robustness check into the revised Supplemental Materials section.

      In any case, I am surprised by the rather small sample size. Due to the small effect sizes for tDCS, it is common to have a minimum of 30 subjects per group in between-subject designs. According to G*Power, a between-subject design with 17 subjects per group could detect only relatively large effect sizes of Cohen's d = 0.99 (alpha = 5%, power = 80%, independent-samples t-test). As explained above, this is far above the effect size that can be expected for tDCS. In addition, small samples bear the risk that results strongly depend on outliers in the data, which might explain the strong effect size observed in the current study. The small sample size should be discussed as a major limitation of the current study and that the results need to be replicated by studies with larger sample sizes. Moreover, to rule out that the results are driven by outlier in the data, the authors should show individual data points in all plots showing empirical data.

      We sincerely thank you for this highly constructive and methodologically sound critique. We completely agree that sample size is a critical consideration in tDCS research, and that visualizing individual data points is essential to rule out the possibility that our findings are driven by outliers.

      We acknowledge that our sample size is smaller than the ~30 per group often recommended for detecting small-to-moderate effects in general cognitive tDCS meta-analyses. We have determined this a priori effect size based on the existing work we published previously (Xu et al., 2023, J Exp Psychol Gen;152(4):1122-1133). In our pilot study (Xu et al., 2023), we identified a significant interaction effect between the single-session tDCS stimulation (active vs sham) and time (pre-test vs post-test) (t = 2.38, p = .02, n = 27; 95% CI [0.14, 1.49]) for changing procrastination willingness in the laboratory settings, indicating a medium effect size. Based on this specific empirical foundation, GPower indicated that a total sample size of 34 (17 per group) was sufficient to achieve 80% power (please see GPower output below). To account for potential attrition, we aimed to recruit 36 participants (18 per group), ultimately retaining 46 participants (23 per group) after exclusions. While we stand by this a priori justification, we fully agree with your overarching point that this remains a constraint.

      We completely agree with your observation regarding the plots. The apparent absence of data points in the previous versions of Figures 3B and 3F was not due to data exclusion, but rather to severe overplotting. Because multiple participants in the active neuromodulation group achieved identical scores (e.g., 0% procrastination rate or 100% task-execution willingness in later sessions), their data points perfectly overlapped, making it appear as though only ~10 points were present. As you helpfully suggested, we now employ jittered scatter plots with adjusted transparency, ensuring that all 23 individual data points per group are clearly visible, even when values are identical. As these revised figures demonstrate, the significant group differences reflect a consistent, cohort-wide shift rather than the influence of isolated outliers. Those

      As you rightly suggested, we have explicitly framed the small sample size as a major limitation and emphasized the necessity for large-scale replication. We have strengthened the wording in the Limitations section to explicitly mention the risk of outlier dependency and the need for larger cohorts.

      Legend Section (Page 28, Line 1161-1163)

      “… To ensure transparency and rule out outlier-driven effects, individual data points for all participants (N=23 per group) are overlaid on the bars using a jittered distribution to prevent overplotting of identical values.”

      Discussion Section (Page 13, Line 691-696)

      “… a major limitation of the current study is the relatively small sample size (total N = 46). While this was determined a priori based on our specific pilot study, small samples inherently bear a higher risk of being influenced by outliers and may overestimate effect sizes compared to large-scale meta-analytic expectations for tDCS. Therefore, these findings warrant caution in generalization and necessitate rigorous replication in larger, adequately powered cohorts.”

      Related to this, in the figure showing individual data points (3B/F), I count only around 10 data points per tDCS group for the 18 participants per group. I ask the authors to modify the plot that the data points from all participants can be seen (for example, by adding some noise on the x-axis for participants with the same value on the y axis).

      Thank you for this kind reminder. As we replied above, those plots have been redrawn by adding the jitters, which favor the readability as you kindly suggested.

      Another surprising aspect of the data is that repeated sessions of tDCS change procrastination behavior up to six months after stimulation. Do the authors think that their tDCS setup leads to such long-lasting neuroplastic changes, and if yes, can they cite prior work where similar dosages of tDCS also showed such long-lasting effects? Or could the results be explained by learning effects, for example because participants in the DLPFC group learned during the repeated tDCS sessions that it feels internally rewarding to finish one's tasks instead of procrastinating them, and they still benefit from this kind of "learned industriousness" 6 months later? In any case, in my view it is important to be more specific about how seven sessions of tDCS can affect behavior half a year later.

      We sincerely thank the reviewer for this highly insightful and thought-provoking comment. The concept of "learned industriousness" is particularly apt and captures a crucial alternative mechanism that we must address. We agree that explaining how seven sessions of tDCS can affect behavior half a year later requires a nuanced discussion of both neurobiological and behavioral learning mechanisms.

      Regarding the first point, we do believe that our multi-session protocol can induce long-lasting neuroplastic changes. While single-session tDCS effects are typically transient, cumulative neurobiological evidence demonstrates that repeated, multi-session protocols (typically ranging from 5 to 10 sessions) can induce activity-dependent, long-term potentiation (LTP)-like plasticity that consolidates over time (Agboada et al., 2020; Au et al., 2017; Jannati et al., 2023). Our 7-session protocol falls squarely within this range of "intensified dosing" designed to promote such consolidation. Meta-analyses and empirical studies on multi-session tDCS have shown that such protocols can produce behavioral and neurophysiological effects lasting weeks to months, particularly when targeting prefrontal regions involved in value-based decision-making and cognitive control (e.g., Brunoni et al., 2013; Sabé et al., 2024; Woodham et al., 2025).

      Furthermore, we completely agree with you for this alternative explanation regarding learning effects. It is highly plausible that participants in the active group, experiencing reduced task aversiveness and increased outcome value during the intervention, learned that completing tasks is internally rewarding. This aligns perfectly with the psychological concept of "learned industriousness" (Eisenberger, 1992), where the reinforcement of effortful behavior makes future engagement more likely. We explicitly acknowledge that repeated exposure to the experience-sampling protocol and the positive feedback of task completion could facilitate this kind of behavioral learning. More importantly, we argue that the learning effects are not bad things in this neuromodulation, and the learning effect and the neuroplasticity may be synergistic. The sham control group underwent the exact same experience-sampling protocol, reported real-life tasks, and had the identical opportunity for "learned industriousness" through feedback. However, as identified in the half-year follow-up, the sham group did not exhibit the same progressive improvement during the intervention, nor did they sustain a significant reduction in procrastination at the 6-month follow-up (their rates returned to near-baseline levels). This divergence suggests that while learning may play a role, the active neuromodulation likely provided the necessary neuroplastic "boost" (e.g., by enhancing prefrontal value-encoding circuits) that facilitated, accelerated, and consolidated this learning, making the behavioral change durable. Without the neuromodulatory enhancement, the mere exposure to the protocol was insufficient to produce long-term change.

      As you kindly suggested, we have explicitly incorporated this nuanced discussion into the revised manuscript, by citing relevant literature on multi-session tDCS plasticity, explicitly acknowledging the "learned industriousness" hypothesis, and reiterating the limitation of having only a single follow-up point.

      Discussion Section (Page 12, Line 627-634)

      “... Despite statistically supporting the TDM, we acknowledge that alternative neurocognitive mechanisms could contribute to the observed reductions in procrastination. For instance, repeated exposure to the experience-sampling protocol may have enhanced participants’ awareness of task progress or facilitated feedback-based learning, thereby increasing the subjective value of goal completion independent of DLPFC neuromodulation. Participants in the active group may have learned during the repeated sessions that completing tasks feels internally rewarding, thereby benefiting from a form of “learned industriousness” (Eisenberger, 1992) that persists months later.”

      Discussion Section (Page 14, Line 719-723)

      “... we explicitly note that a single 6-month follow-up timepoint cannot definitively establish the stability or trajectory of these effects. Future studies incorporating multiple longitudinal assessments (e.g., 1-month, 3-month, 6-month, 12-month) are required to substantiate claims about long-term retention and to disentangle the precise contributions of neuroplasticity versus behavioral learning.”

      Lastly, the link to the data repository works, but I could not inspect the data because I was asked to request access to the data, which I did not do in order to remain anonymous.

      Thank you a lot to take invaluable to review our data and code in this repository. As we reported previously, all the data and code to support those findings have been deposited in the eLife online submission system for your reviews and scrutiny before this manuscript is formally published. As the editorial policy of eLife on VOR (Version of Record) instructed, to prevent from mixture of codes and data across multiple round of revisions, those data and codes in the final version would be released once this paper is formally published. Please do not worry for the anonymity policy. This is a public peer review, and it thus enables those helpful comments that you kindly suggested to be public when this manuscript is formally published. Again, thank you to substantially contribute on this revised manuscript by sharing those helpful suggestions.

      Recommendations for the authors:

      Editors note: We encourage the authors to consider the remaining reviewer concerns and revise the manuscript accordingly.

      Thank you so much for this warm and kind reminder. We have addressed all of those concerns that remained by the new Reviewer #4, point-by-point. All the co-authors do appreciate you for handling our manuscript, and for contributing those fruitful and helpful comments. We do believe that the quality of this manuscript has been substantially improved, benefiting from this editorial process.

      References

      Agboada, D., Mosayebi-Samani, M., Kuo, M. F., & Nitsche, M. A. (2020). Induction of long-term potentiation-like plasticity in the primary motor cortex with repeated anodal transcranial direct current stimulation - Better effects with intensified protocols? Brain Stimulation, 13(4), 987–997. https://doi.org/10.1016/j.brs.2020.04.009

      Au, J., Karsten, C., Buschkuehl, M., & Jaeggi, S. M. (2017). Optimizing transcranial direct current stimulation protocols to promote long-term learning. Journal of Cognitive Enhancement, 1(1), 65–72. https://doi.org/10.1007/s41465-017-0007-6

      Brunoni, A. R., Boggio, P. S., Ferrucci, R., Priori, A., & Fregni, F. (2013). Transcranial direct current stimulation: challenges, opportunities, and impact on psychiatry and neurorehabilitation. Frontiers in Psychiatry, 4, 19. https://doi.org/10.3389/fpsyt.2013.00019

      Cole, E. J., Stimpson, K. H., Bentzley, B. S., Gulser, M., Cherian, K., Tischler, C., Nejad, R., Pankow, H., Choi, E., Aaron, H., Espil, F. M., Pannu, J., Xiao, X., Duvio, D., Solvason, H. B., Hawkins, J., Guerra, A., Jo, B., Raj, K. S., Phillips, A. L., … Williams, N. R. (2020). Stanford accelerated intelligent neuromodulation therapy for treatment-resistant depression. The American Journal of Psychiatry, 177(8), 716–726. https://doi.org/10.1176/appi.ajp.2019.19070720

      Eisenberger, R. (1992). Learned industriousness. Psychological Review, 99(2), 248–267. https://doi.org/10.1037/0033-295X.99.2.248

      Hutton, T. M., Aaronson, S. T., Carpenter, L. L., Pages, K., Krantz, D., Lucas, L., Chen, B., & Sackeim, H. A. (2023). Dosing transcranial magnetic stimulation in major depressive disorder: Relations between number of treatment sessions and effectiveness in a large patient registry. Brain Stimulation, 16(5), 1510–1521. https://doi.org/10.1016/j.brs.2023.10.001

      Jannati, A., Oberman, L. M., Rotenberg, A., & Pascual-Leone, A. (2023). Assessing the mechanisms of brain plasticity by transcranial magnetic stimulation. Neuropsychopharmacology, 48(1), 191–208. https://doi.org/10.1038/s41386-022-01453-8

      Ke, Y., Liu, S., Chen, L., et al. (2023). Lasting enhancements in neural efficiency by multi-session transcranial direct current stimulation during working memory training. npj Science of Learning, 8(1), Article 23. https://doi.org/10.1038/s41539-023-00200-y

      Sabé, M., Hyde, J., Cramer, C., Eberhard, A., Crippa, A., Brunoni, A. R., Aleman, A., Kaiser, S., Baldwin, D. S., Garner, M., Sentissi, O., Fiedorowicz, J. G., Brandt, V., Cortese, S., & Solmi, M. (2024). Transcranial magnetic stimulation and transcranial direct current stimulation across mental disorders: A systematic review and dose-response meta-analysis. JAMA Network Open, 7(5), e2412616. https://doi.org/10.1001/jamanetworkopen.2024.12616

      Schulze, L., Feffer, K., Lozano, C., Giacobbe, P., Daskalakis, Z. J., Blumberger, D. M., & Downar, J. (2018). Number of pulses or number of sessions? An open-label study of trajectories of improvement for once- vs. twice-daily dorsomedial prefrontal rTMS in major depression. Brain Stimulation, 11(2), 327–336. https://doi.org/10.1016/j.brs.2017.11.002

      Woodham, R. D., Selvaraj, S., Lajmi, N., Hobday, H., Sheehan, G., Ghazi-Noori, A.-R., Lagerberg, P. J., Rizvi, M., Kwon, S. S., Orhii, P., Maislin, D., Hernandez, L., Machado-Vieira, R., Soares, J. C., Young, A. H., & Fu, C. H. Y. (2025). Home-based transcranial direct current stimulation treatment for major depressive disorder: a fully remote phase 2 randomized sham-controlled trial. Nature Medicine, 31(1), 87–95. https://doi.org/10.1038/s41591-024-03305-y

      Zhong, M., Cywiak, C., Metto, A. C., Liu, X., Qian, C., et al. (2021). Multi-session delivery of synchronous rTMS and sensory stimulation induces long-term plasticity. Brain Stimulation, 14(4), 884–894. https://doi.org/10.1016/j.brs.2021.05.003

    1. eLife Assessment

      This manuscript examines how overexpression of sphingosine 1-phosphate receptor 1 in a mouse model affects neutrophil distribution, metabolism and function. Given the significance of S1PR1 in inflammatory mechanisms, this work is valuable for uncovering novel aspects of S1PR1 signaling in neutrophil dynamics and function. The evidence provided is solid as it integrates single-cell transcriptomics, flow cytometry, and multiple in vivo studies. This work also raises key questions about distinct physiological and pathophysiological roles of S1PR1 signaling in neutrophils that could become therapeutic targets.

    2. Reviewer #1 (Public review):

      Summary:

      This interesting paper demonstrates that transgenic over-expression of sphingosine 1-phosphate receptor 1 (S1PR1) on neutrophils alters their phenotype, resulting in (1) accumulation of neutrophils in blood, spleen, lung, and liver; (2) a shift in homing receptor expression with reduced CXCR2 and elevated CXCR4; (3) altered transcriptional profile with an increase in "G5c" neutrophils and reduced "module scores" for apoptosis and inflammatory response; (4) reduced ROS production upon fLMP stimulation; and (5) altered responses to bacterial and viral infections of the lung. It raises many interesting questions about how S1P signaling regulates neutrophil biology, and hence will be the basis of future studies. These include: (1) What is the physiological role of S1PR1 signaling in neutrophils? Although there is no dramatic effect on numbers upon S1PR1 loss, is there an effect on any of the other parameters measured? (2) What is unique about the lung that S1PR1 over-expression is particularly impactful there? (3) What distinguishes the bacterial context in which S1PR1 over-expression is maladaptive from the viral context in which S1PR1 over-expression is protective? and (4) Can treatment with an S1PR1 agonist mimic S1PR1 over-expression? As a possibly related question, when in neutrophil development does S1PR1 signaling function to shift the phenotype?

      Strengths:

      (1) A comprehensive characterization of S1PR1-transgenic neutrophils.

      (2) Opens many interesting areas of investigation.

      Weaknesses:

      Although some characterization of the neutrophil-specific Mrp8-Cre is done, most of the experiments use the more widely expressed LysM-Cre. The redistribution phenotype is much stronger with LysM-Cre than with Mrp8-Cre, so it is unclear what effects are attributable to a cell-intrinsic role of S1PR1, even in studies of neutrophils analyzed ex vivo.

    3. Reviewer #2 (Public review):

      The authors have utilised two main models to assess the function of S1PR1 in neutrophils in mice. The knockout of this receptor shows no conclusive effect on neutrophil numbers or functions; it was only the overexpression that resulted in significant alterations. Therefore, often the conclusions do not describe normal or disease physiology but could be useful in a bioengineering context.

      Strengths:

      From a bioengineering standpoint, this seems like an important study - showing enforced expression of S1PR1 in neutrophils has improved outcomes for influenza infection (Figures 6 and 7).

      Weaknesses:

      Although the strength is the influenza model, genetic modification of human neutrophils cannot be a strategy, and therefore, is there any way to increase this receptor for mouse, or more importantly, human neutrophils? This study only looks at mice with a non-physiological model of overexpression. It does not offer a real therapeutic option, which drastically hinders the importance of the study. I have other concerns with the data analysis and interpretation, which I detail on a figure-by-figure basis (and how it relates to conclusions) below:

      Main specific issues:

      (1) Figure 2A+B: This is unconvincing; in the surface staining there seem to be real cells positive for the receptor (high staining in the histogram), but none of the transgenic protein is getting there? This undermines the idea that the effects of the transgene are related to S1P signalling. In the 'Total S1PR1' this is both underwhelming and misleading, as an isotype control (or better S1PR1 knockout) is missing, which would give a better representation of actual expression (flow cytometry autofluorescence famously increases in the red laser channels with fix/perm). The Imagestream chosen images are showing best-case scenarios - and aren't representative. What does the isotype/ KO look like here? All in all, the conclusion on receptor internalization is not well supported, especially when theoretically the TG overexpression should overload S1P availability. This also highlights the lack of another control - does overexpression of another random/non-functional protein have the same effect? To play devil's advocate, perhaps overloading of the ubiquitin-proteasome system is responsible?

      (2) Figures 2E-H: In the text, the authors should fix the statement 'Additionally, surface CXCR2 was downregulated and CXCR4 upregulated in LysM-S1pr1 TG neutrophils across bone marrow, spleen, and blood (Fig. 2, E and F)' to better reflect that there is no significant difference in the bone marrow regarding CXCR4. Of note, the total MFI from this data would also be informative, another noticeable absence being the gating strategies for much of the data. Also, alter the statement: 'CD62L expression was largely preserved across compartments, with only a modest reduction in bone marrow neutrophils (Fig. 2G)'. A 50% reduction in CD62L is not modest.

      (3) Supplemental Figure 3. A common theme: the wrong statistics have been used here, which has led to a false conclusion. Megakaryocyte/erythrocyte progenitors (MEPs) were only elevated in 2/3 TG mice, and the numbers are so small that this is not significant by any measure of the word. This is certainly not statistically significant if the correct test of (log-normalized) two-way ANOVA is performed (with Sidak's post hoc test). Another acceptable test would be Kruskal-Wallis with Dunn's post-test just for MEPs.

      (4) Starting at Figure 3, the authors refer to 'S1PR1hi neutrophil accumulation'. Crucially, the authors must here and throughout be explicitly clear in which cells they are referring to, as this can be misleading - particularly as there are real S1PR1-high cells identified in Figure 2A surface staining. It is my understanding that the authors here mean the transgenic artificially high mice - a very large distinction.

      (5) Figure 3A: It is difficult to interpret the figure with the necessary details about the experiment. For instance, there is no mention that this is sterile inflammation or what caused it.

      (6) Figure 3B and C: It should be made clear whether these splenic neutrophils are related to the time course of peritoneal inflammation in 3A. Why are there so many apoptotic neutrophils in the spleen? The low numbers here suggest a processing issue rather than real death in vivo (which usually is absent).

      (7) Figure 3D: This can also be misleading - the wrong statistics are again used. This should be a log-transformed two-way ANOVA. Regardless of this, the data is not strong enough to be conclusive, a minor effect at best that could also just be related to the type of cell tracker used.

      (8) Figure 5E: It is stated that 'LysM-S1pr1 TG mice exhibited a higher bacterial burden in the lungs than controls (Fig. 5E).' Again, misleading results, first the wrong statistical test was used (correct = log norm one-way ANOVA with Tukey's or Kruskal Wallis with Dunn's), secondly the only significance is between S1PR1(fsf) and the Mrp8-S1PR1, not with the LysM TG. 5F is also not strong, with only 2/7 values appearing outside the range of the control - P values can be misleading when poor statistics are used.

      (9) Figure 6G: Some discussion should be given for why Neutrophils are lower in BALF in the IAV model - even though higher in the lung in the non-IAC mice in Figure 1. In general, rather than focusing on the non-physiological differences, the discussion could better reflect the inconsistencies and more fully address the difference between the TG and KO and what this means going forward.

    4. Reviewer #3 (Public review):

      Summary:

      Using mice that overexpress S1PR1 in myeloid cells or specifically in neutrophils, the authors show that increased S1PR1 promotes neutrophil release from the bone marrow and accumulation in blood and peripheral tissues without causing baseline tissue injury. These cells acquire a CXCR4-high, CXCR2-low, CD101-low phenotype, survive longer, and display enhanced mitochondrial metabolism and mTOR signaling, together with reduced apoptotic, inflammatory, and ROS-related programs. Although phagocytosis is preserved, ROS production is markedly reduced. This is associated with impaired bacterial clearance in the lung but improved outcomes during influenza infection, including better survival, less weight loss, improved oxygenation, lower viral burden, and reduced lung inflammation. In contrast, myeloid S1PR1 deletion produces little detectable phenotype. The authors therefore propose that S1PR1 separates neutrophil persistence from inflammatory function, improving tolerance to viral lung injury at the expense of antibacterial defense.

      Strengths:

      This is a technically solid paper using novel mouse models to overexpress S1PR1 specifically in myeloid cells as well as neutrophils. The data are striking with respect to neutrophil expansion. The diverse roles of neutrophils and their population heterogeneity are an important scientific area that has led to many recent breakthroughs - PMC11785525; PMC12823425, thus this is a timely study.

      Weaknesses:

      The study mainly demonstrates what S1PR1 overexpression is sufficient to do, rather than establishing the physiological role of endogenous S1PR1. The conclusions should therefore be narrowed unless the authors provide stronger loss-of-function and physiological validation. As written, the abstract ("S1PR1 promotes mitochondrial fitness, enhances survival, and reduces inflammatory output") and the conclusion ("S1PR1 serves as a key regulatory axis") are sufficiency claims but should not be promoted as necessity claims. The honest sentence is: "Thus, a better conclusion would be that enforced S1PR1 expression is sufficient to reprogram neutrophils".

      The authors do not confirm efficient S1pr1 deletion in neutrophils. Furthermore, the knockout is examined only under steady-state conditions and limited in vitro stimulation, but not in the bacterial or influenza models where the transgenic phenotype is observed. Without these experiments, the study cannot establish whether endogenous S1PR1 is necessary for the reported functions.

      The degree of S1PR1 overexpression is not quantified relative to normal physiological levels. The authors should determine whether naturally occurring S1PR1-high neutrophils display the same survival, metabolic, trafficking, and inflammatory features observed in the transgenic cells.

      Analysis of relevant human or mouse datasets, including sepsis, ARDS, viral infection, cancer, or aging, would also help establish whether this neutrophil state exists physiologically.

      Surface S1PR1 expression appears similar between control and transgenic neutrophils, whereas total intracellular receptor is increased. This suggests that the phenotype may depend on receptor internalization or endosomal signaling. An internalization-deficient S1PR1 model, such as S1P1-S5A, would help distinguish sustained surface signaling from internalization-dependent signaling. The authors should also determine whether the phenotype requires ligand binding, Gi signaling, and mTOR activity.

      The reduction in CXCR2 and decreased neutrophil accumulation in the airways could alone explain the protection from influenza-induced lung injury. The current experiments do not clearly distinguish neutrophil reprogramming from defective migration into the alveolar space.

      Although this may be outside the scope of the current study, the authors should directly test whether CXCR2 inhibition reproduces the phenotype.

      The reported reduction in viral load should also be confirmed using plaque assay or TCID50, and the possible contribution of NET formation should be examined.

    5. Author response:

      Reviewer #1 (Public review):

      Summary:

      This interesting paper demonstrates that transgenic over-expression of sphingosine 1-phosphate receptor 1 (S1PR1) on neutrophils alters their phenotype, resulting in (1) accumulation of neutrophils in blood, spleen, lung, and liver; (2) a shift in homing receptor expression with reduced CXCR2 and elevated CXCR4; (3) altered transcriptional profile with an increase in "G5c" neutrophils and reduced "module scores" for apoptosis and inflammatory response; (4) reduced ROS production upon fLMP stimulation; and (5) altered responses to bacterial and viral infections of the lung. It raises many interesting questions about how S1P signaling regulates neutrophil biology, and hence will be the basis of future studies. These include: (1) What is the physiological role of S1PR1 signaling in neutrophils? Although there is no dramatic effect on numbers upon S1PR1 loss, is there an effect on any of the other parameters measured? (2) What is unique about the lung that S1PR1 over-expression is particularly impactful there? (3) What distinguishes the bacterial context in which S1PR1 over-expression is maladaptive from the viral context in which S1PR1 over-expression is protective? and (4) Can treatment with an S1PR1 agonist mimic S1PR1 over-expression? As a possibly related question, when in neutrophil development does S1PR1 signaling function to shift the phenotype?

      Strengths:

      (1) A comprehensive characterization of S1PR1-transgenic neutrophils.

      (2) Opens many interesting areas of investigation.

      We thank the reviewer for a very positive assessment of our work and for raising very interesting questions, which will be useful to extend this work in the future.

      Weaknesses:

      Although some characterization of the neutrophil-specific Mrp8-Cre is done, most of the experiments use the more widely expressed LysM-Cre. The redistribution phenotype is much stronger with LysM-Cre than with Mrp8-Cre, so it is unclear what effects are attributable to a cell-intrinsic role of S1PR1, even in studies of neutrophils analyzed ex vivo.

      We acknowledge that the data from Mrp8-Cre mouse strain are more limited than the LysM-Cre counterparts. This is because we initially characterized the LysM-Cre S1pr1 KO and TG strains and confirmed key findings relevant to neutrophils in the Mrp8-Cre counterparts. Going forward, more studies will be done in the Mrp8-Cre strain as suggested by the reviewer.

      Reviewer #2 (Public review):

      The authors have utilised two main models to assess the function of S1PR1 in neutrophils in mice. The knockout of this receptor shows no conclusive effect on neutrophil numbers or functions; it was only the overexpression that resulted in significant alterations. Therefore, often the conclusions do not describe normal or disease physiology but could be useful in a bioengineering context.

      We agree with the reviewer that some of the key findings described in our manuscript, for example, neutrophil survival, spleen size, and ROS reduction, etc., were not observed in the S1pr1 KO strains. Our interpretation is that other receptors, for example S1PR4, could be involved in compensating for the loss of S1PR1. We will explain this better in the revisions.

      Strengths:

      From a bioengineering standpoint, this seems like an important study - showing enforced expression of S1PR1 in neutrophils has improved outcomes for influenza infection (Figures 6 and 7).

      We agree with the reviewer that overexpression of S1PR1 could be useful from “bioengineering standpoint”.

      Weaknesses:

      Although the strength is the influenza model, genetic modification of human neutrophils cannot be a strategy, and therefore, is there any way to increase this receptor for mouse, or more importantly, human neutrophils? This study only looks at mice with a non-physiological model of overexpression. It does not offer a real therapeutic option, which drastically hinders the importance of the study. I have other concerns with the data analysis and interpretation, which I detail on a figure-by-figure basis (and how it relates to conclusions) below:

      The reviewer's comments are acknowledged. However, at this early stage of discovery, we feel that it would be premature to address the issue of “a real therapeutic option”. This can be addressed in the future.

      Main specific issues:

      (1) Figure 2A+B: This is unconvincing; in the surface staining there seem to be real cells positive for the receptor (high staining in the histogram), but none of the transgenic protein is getting there? This undermines the idea that the effects of the transgene are related to S1P signalling. In the 'Total S1PR1' this is both underwhelming and misleading, as an isotype control (or better S1PR1 knockout) is missing, which would give a better representation of actual expression (flow cytometry autofluorescence famously increases in the red laser channels with fix/perm). The Imagestream chosen images are showing best-case scenarios - and aren't representative. What does the isotype/ KO look like here? All in all, the conclusion on receptor internalization is not well supported, especially when theoretically the TG overexpression should overload S1P availability. This also highlights the lack of another control - does overexpression of another random/non-functional protein have the same effect? To play devil's advocate, perhaps overloading of the ubiquitin-proteasome system is responsible?

      We will conduct S1PR1 antibody staining with knockout neutrophils as requested by the reviewer. The results will be shown in the revision. The comment of the reviewer on “ubiquitin-proteosome” system is not relevant in our opinion and does not impact the validity of our findings. The approach that we used is standard mouse genetics, and the results should be interpreted from that perspective and not from “devil's advocate”.

      (2) Figures 2E-H: In the text, the authors should fix the statement 'Additionally, surface CXCR2 was downregulated and CXCR4 upregulated in LysM-S1pr1 TG neutrophils across bone marrow, spleen, and blood (Fig. 2, E and F)' to better reflect that there is no significant difference in the bone marrow regarding CXCR4. Of note, the total MFI from this data would also be informative, another noticeable absence being the gating strategies for much of the data. Also, alter the statement: 'CD62L expression was largely preserved across compartments, with only a modest reduction in bone marrow neutrophils (Fig. 2G)'. A 50% reduction in CD62L is not modest.

      These statements will be changed in the revised manuscript, which we hope to submit soon.

      (3) Supplemental Figure 3. A common theme: the wrong statistics have been used here, which has led to a false conclusion. Megakaryocyte/erythrocyte progenitors (MEPs) were only elevated in 2/3 TG mice, and the numbers are so small that this is not significant by any measure of the word. This is certainly not statistically significant if the correct test of (log-normalized) two-way ANOVA is performed (with Sidak's post hoc test). Another acceptable test would be Kruskal-Wallis with Dunn's post-test just for MEPs.

      We will reanalyze these data with different statistical tests and discuss these in the revision.

      (4) Starting at Figure 3, the authors refer to 'S1PR1hi neutrophil accumulation'. Crucially, the authors must here and throughout be explicitly clear in which cells they are referring to, as this can be misleading - particularly as there are real S1PR1-high cells identified in Figure 2A surface staining. It is my understanding that the authors here mean the transgenic artificially high mice - a very large distinction.

      We will clarify and edit this terminology throughout the manuscript in the revision. S1PR1<sup>hi</sup> does not strictly imply receptor on the cell surface. It also includes internalized S1PR1 receptors.

      (5) Figure 3A: It is difficult to interpret the figure with the necessary details about the experiment. For instance, there is no mention that this is sterile inflammation or what caused it.

      This will be edited in the revision.

      (6) Figure 3B and C: It should be made clear whether these splenic neutrophils are related to the time course of peritoneal inflammation in 3A. Why are there so many apoptotic neutrophils in the spleen? The low numbers here suggest a processing issue rather than real death in vivo (which usually is absent).

      This will be addressed in the revision.

      (7) Figure 3D: This can also be misleading - the wrong statistics are again used. This should be a log-transformed two-way ANOVA. Regardless of this, the data is not strong enough to be conclusive, a minor effect at best that could also just be related to the type of cell tracker used.

      Same as #3 above.

      (8) Figure 5E: It is stated that 'LysM-S1pr1 TG mice exhibited a higher bacterial burden in the lungs than controls (Fig. 5E).' Again, misleading results, first the wrong statistical test was used (correct = log norm one-way ANOVA with Tukey's or Kruskal Wallis with Dunn's), secondly the only significance is between S1PR1(fsf) and the Mrp8-S1PR1, not with the LysM TG. 5F is also not strong, with only 2/7 values appearing outside the range of the control - P values can be misleading when poor statistics are used.

      Same as #3 above.

      (9) Figure 6G: Some discussion should be given for why Neutrophils are lower in BALF in the IAV model - even though higher in the lung in the non-IAC mice in Figure 1. In general, rather than focusing on the non-physiological differences, the discussion could better reflect the inconsistencies and more fully address the difference between the TG and KO and what this means going forward.

      Same as #5 above.

      Reviewer #3 (Public review):

      Summary:

      Using mice that overexpress S1PR1 in myeloid cells or specifically in neutrophils, the authors show that increased S1PR1 promotes neutrophil release from the bone marrow and accumulation in blood and peripheral tissues without causing baseline tissue injury. These cells acquire a CXCR4-high, CXCR2-low, CD101-low phenotype, survive longer, and display enhanced mitochondrial metabolism and mTOR signaling, together with reduced apoptotic, inflammatory, and ROS-related programs. Although phagocytosis is preserved, ROS production is markedly reduced. This is associated with impaired bacterial clearance in the lung but improved outcomes during influenza infection, including better survival, less weight loss, improved oxygenation, lower viral burden, and reduced lung inflammation. In contrast, myeloid S1PR1 deletion produces little detectable phenotype. The authors therefore propose that S1PR1 separates neutrophil persistence from inflammatory function, improving tolerance to viral lung injury at the expense of antibacterial defense.

      Strengths:

      This is a technically solid paper using novel mouse models to overexpress S1PR1 specifically in myeloid cells as well as neutrophils. The data are striking with respect to neutrophil expansion. The diverse roles of neutrophils and their population heterogeneity are an important scientific area that has led to many recent breakthroughs - PMC11785525; PMC12823425, thus this is a timely study.

      We thank the reviewer for positive comments.

      Weaknesses:

      The study mainly demonstrates what S1PR1 overexpression is sufficient to do, rather than establishing the physiological role of endogenous S1PR1. The conclusions should therefore be narrowed unless the authors provide stronger loss-of-function and physiological validation. As written, the abstract ("S1PR1 promotes mitochondrial fitness, enhances survival, and reduces inflammatory output") and the conclusion ("S1PR1 serves as a key regulatory axis") are sufficiency claims but should not be promoted as necessity claims. The honest sentence is: "Thus, a better conclusion would be that enforced S1PR1 expression is sufficient to reprogram neutrophils".

      We will edit the text to better reflect sufficiency versus necessity terms.

      The authors do not confirm efficient S1pr1 deletion in neutrophils. Furthermore, the knockout is examined only under steady-state conditions and limited in vitro stimulation, but not in the bacterial or influenza models where the transgenic phenotype is observed. Without these experiments, the study cannot establish whether endogenous S1PR1 is necessary for the reported functions.

      We will provide qRT-PCR data (that we have done already but did not include in the original version) confirming efficient deletion of the S1pr1 gene.

      The degree of S1PR1 overexpression is not quantified relative to normal physiological levels. The authors should determine whether naturally occurring S1PR1-high neutrophils display the same survival, metabolic, trafficking, and inflammatory features observed in the transgenic cells.

      We agree that the level of S1PR1 overexpression in our transgenic neutrophils should be quantitatively compared with endogenous S1PR1 expression. We will provide the quantitative result in the revision. Our analyses of independent human (GSE216009) and mouse (GSE243466 and GSE266518) single-cell RNA-seq datasets indicate that endogenous mRNA expression for this receptor is higher in specific neutrophil states associated with inflammatory, immature, and tissue-associated populations. Notably, these endogenous S1PR1-expressing neutrophils show transcriptional features involving altered oxidative/inflammatory programs, chemotaxis, and mitochondrial/metabolic regulation that partially overlap with the phenotype of S1PR1-transgenic neutrophils. We will provide additional quantitative and dataset analyses in the revised manuscript.

      Analysis of relevant human or mouse datasets, including sepsis, ARDS, viral infection, cancer, or aging, would also help establish whether this neutrophil state exists physiologically.

      As described above, we have analyzed independent human and mouse single-cell RNA-seq datasets from sepsis, cancer, and other inflammatory conditions to determine whether endogenous S1PR1-expressing neutrophil states occur naturally across biological and disease contexts. These analyses support the presence of distinct endogenous S1PR1-expressing neutrophil populations across multiple biological contexts. We will provide the expanded analyses and corresponding data in the revised manuscript.

      Surface S1PR1 expression appears similar between control and transgenic neutrophils, whereas total intracellular receptor is increased. This suggests that the phenotype may depend on receptor internalization or endosomal signaling. An internalization-deficient S1PR1 model, such as S1P1-S5A, would help distinguish sustained surface signaling from internalization-dependent signaling. The authors should also determine whether the phenotype requires ligand binding, Gi signaling, and mTOR activity.

      We agree that future experiments will use internalization-defective S1PR1 S5A knock-in neutrophils.

      The reduction in CXCR2 and decreased neutrophil accumulation in the airways could alone explain the protection from influenza-induced lung injury. The current experiments do not clearly distinguish neutrophil reprogramming from defective migration into the alveolar space.

      We agree. We did not examine the role of CXCR2 in the influenza experiments. Our interpretation is that reduced ROS from transgenic neutrophils reduced lung injury.

      Although this may be outside the scope of the current study, the authors should directly test whether CXCR2 inhibition reproduces the phenotype.

      This is a good point that the reviewer brought up. This can be addressed in a future study.

      The reported reduction in viral load should also be confirmed using plaque assay or TCID50, and the possible contribution of NET formation should be examined.

      This is again a good suggestion that can be addressed in a future study.

    1. eLife Assessment

      This paper demonstrates that relocating the MreB-Rod complex to the poles of E. coli can reprogram cell-wall synthesis, causing cells to elongate through peptidoglycan incorporation at the poles rather than along the cell length. This is a valuable finding for understanding how bacterial growth can be spatially reprogrammed and provides insight into how distinct modes of bacterial growth might arise. However, the evidence is currently incomplete, particularly regarding the nature of the polar structures, their relationship to cell shape and growth, and the requirement for a pre-existing pole.

    2. Reviewer #1 (Public review):

      This is an interesting paper, with the primary finding being that localizing MreB or PBP2 to the cell poles in E. coli is primarily demonstrated via an aggregate formed by expressing M. xanthus MreB.

      I have 2 main concerns:

      (1) First, the authors should clarify and adjust their interpretation of FDAA incorporation: As written, the authors interpret FDAA incorporation as being caused by incorporation during PG polymerization. However, in E. coli, FDAA incorporation does not result from the elongation of PG strands or their initial 4-3 crosslinking by DD transpeptidases following polymerization, but rather by the remodeling of L,D-transpeptidases.

      Thus, it is not accurate to refer to the FDAA incorporation as PG elongation, but rather the modification of crosslinks from 4-3 to 3-3 crosslinks at that location. To claim a link to PG polymerization, other experiments or substantial explanations are needed.

      (E. coli Cells Incorporate FDAAs by L,D-TPases in a Growth Independent Manner) - https://doi.org/10.1021/acschembio.

      (2) Second, the evidence provided that this is polar elongation is not sufficient to prove elongation. In all images claiming polar growth, the FDAA focus appears as a single spot, which corresponds to the MreB aggregate visible in bright-field. That might indicate incorporation, but it does not demonstrate polar elongation. To prove this, the authors should do different-length pulses of FDAAs and demonstrate an increasing length of labeled PG along the cell. Cells with one focus should show increasing length; polar foci should elongate from both ends. I find the current 2-color labeling insufficient, as BADA labels the entire cell.

      Small points:

      (3) Lines 350- 302: "In this case, as the nonpolar region is no longer the growth zone, the established cylindrical PG structure is sufficient to maintain cell width, where MreB filaments become nonessential." Does polar elongation give robustness to rod shape? The authors should include an analysis of cell width and its variation within a cell and between cells.

      (4) In the abstract: "This reprogrammed growth mode bypasses the requirement for MreB filaments, highlighting a plasticity of the Rod system that suggests polar elongation may have emerged through the evolutionary loss of MreB." This argument should not be made without evolutionary analysis or reference to such work indicating this is the case.

      (5) 365-367 "PG-depleted spheroplasts can spontaneously regenerate rod shape through curvature-dependent localization of MreB filaments [21, 57]". This should be amended. The Billing paper did indeed study spheroplasts, but the Hussain paper used teichoic acid-depleted cells that still had a cell wall.

    3. Reviewer #2 (Public review):

      Summary:

      Based on observations of localisation of MreBEc at the poles within an aggregate-like structure, upon heterologous expression of MreBMx, the authors set out to investigate how this non-canonical localisation of MreBs leads to a reprogramming of peptidoglycan synthesis to the poles. This is analogous to the polar growth observed in phyla which are not dependent on dispersed growth of PG, but only at the poles, and are MreB independent.

      The authors proceed to establish that PG synthesis is MreB-dependent, Rod enzyme-dependent, and requires the prior establishment of a pole.

      Strengths:

      (1) It is a very interesting idea to design experiments to demonstrate reprogramming of non-polar to polar growth based on the observation of localisation of a heterologously expressed MreB.

      (2) The experiments to demonstrate the factors that determine polar growth and the observation of the PG in each of these experimental situations are convincing.

      (3) I find the observation of an extra layer of PG in the heterologously expressed system very intriguing. It will be interesting to see if this layer merges with the other PG layer at some stage or branches from the non-polar growth near the poles.

      Weaknesses:

      (1) It is not clear what exactly the identity of the polar aggregates is and how much of this activity is an artefact of partially functional MreBs.

      (2) I find it intriguing that the localisation and growth are predominantly at one pole only. It is unclear to me how this can be reconciled with growth and shape maintenance, and an increase in length and width. Is the increase in length and width a consequence of misshapen cells that are bulged in the absence of a normal PG layer?

      (3) The authors do not follow up on the observations in the first figure on the length and width changes and the extra peptidoglycan layer (which I feel are the most interesting aspects), and how this can be connected to the polar growth observed in the later sections of the manuscript.

      (4) The claim that this could be a precursor of an MreB-independent polar growth mechanism appears to be a bit far-fetched, because the system is still dependent on having an established pole for PG synthesis to occur in the new place.

    4. Reviewer #3 (Public review):

      Summary:

      Most rod-shaped bacteria grow by one of two mechanisms: growth from the pole or growth from the midcell. It is rare for a single species to utilize both modes of growth, although a few examples do exist. Here, the authors have artificially induced E. coli cells to grow from the poles, either by expressing mreB from Myxoccocus xanthus in E. coli, leading to the mislocalization of MreB to the poles in large aggregates, or by forcing the localization of major cell wall synthesis proteins to the cell pole. The fact that cells switched modes of growth suggests an evolutionary pathway from midcell to polar growing cells as well as suggests that there might be unknown conditions in nature when cells may switch growth modes.

      Strengths:

      (1) The authors use a strain that has replaced mreB with a functional fluorescent version at the native site. This eliminates any effects of having two copies of mreB. Because MreB is fluorescently tagged, they can monitor its localization when mreB from M. xanthus is expressed in E. coli. They notice that MreBec now forms bright polar foci and that there appear to be changes to the cell wall at the pole.

      (2) D-amino acids are specific to the cell wall, and fluorescent versions (FDAA) have been used to mark sites of new cell wall insertion. The authors use these FDAAs to determine how cell wall synthesis correlates to MreB and if that changes when MreBmx is expressed. Again, there is pretty clear evidence that cell wall synthesis follows MreB localization to the pole.

      (3) MreB itself does not synthesize the cell wall, but localizes the proteins, such as PBP2, that do. Using a published method to force proteins to the pole, the authors show that when they target PBP2 to the pole, they can phenocopy the polar growth seen when MreB is polar. Interestingly, these cells become resistant to A22, a drug that targets MreB, suggesting that localized growth at the pole does not require MreB and is sufficient to maintain rod shape.

      Weaknesses:

      (1) The authors do not show what the poles of control cells look like, making it difficult to determine if there is a change when MreBmx is expressed. However, the localization of both MreBec and MreBmx clearly forms bright foci at the pole.

      (2) While more quantification is needed, the authors show some evidence that RodZ, an MreB interaction partner, is needed for this polar growth, as cells lacking rodZ still form foci at pole-like regions when MreBmx is in the cell; however, these cells remain spherical and do not elongate from these foci.

      When MreB is deleted, and cells become spherical, the authors were unable to cause the polar growth mode. They suggest that this is due to the lack of a preexisting pole; however, experimental evidence to test this is missing.

      Conclusion:

      Overall, the authors do a good job of showing that E. coli can grow with a polar method rather than a midcell method of cell wall insertion. It is unclear why MreBec forms at poles when MreBmx is present and even if this MreB is functional. The foci look similar to inclusion bodies, which are normally aggregates of misfolded proteins that migrate to the poles. Past work has shown that when MreB is more polarly localized, branches form, which is not seen here. Importantly, the authors also show that there is feedback between the localization of MreB and PBP2 as both appear to regulate the localization of the other.

    5. Author response:

      Public Reviews:

      Reviewer #1 (Public review):

      This is an interesting paper, with the primary finding being that localizing MreB or PBP2 to the cell poles in E. coli is primarily demonstrated via an aggregate formed by expressing M. xanthus MreB.

      I have 2 main concerns:

      (1) First, the authors should clarify and adjust their interpretation of FDAA incorporation: As written, the authors interpret FDAA incorporation as being caused by incorporation during PG polymerization. However, in E. coli, FDAA incorporation does not result from the elongation of PG strands or their initial 4-3 crosslinking by DD transpeptidases following polymerization, but rather by the remodeling of L,D-transpeptidases.

      Thus, it is not accurate to refer to the FDAA incorporation as PG elongation, but rather the modification of crosslinks from 4-3 to 3-3 crosslinks at that location. To claim a link to PG polymerization, other experiments or substantial explanations are needed.

      (E. coli Cells Incorporate FDAAs by L,D-TPases in a Growth Independent Manner) - https://doi.org/10.1021/acschembio.

      We appreciate the reviewer for bringing up this important question. We will address this concern by further explaining our results in the revised manuscript.

      However, we respectfully disagree with the reviewer for two reasons:

      First, the paper by Kuru et al. (mentioned by the reviewer) revealed that (exact quote): “Our in vitro and in vivo data unequivocally demonstrate that these bacteria incorporate FDAAs using two extra cytoplasmic pathways: through activity of their D, D-transpeptidases, and, if present, by their L, D-transpeptidases…These mechanistic findings enabled development of a new, FDAA-based, in vitro labelling approach that reports on subcellular distribution of muropeptides, an especially important attribute to enable the study of bacteria with poorly defined growth modes” (Kuru et al., 2019).

      Thus, the paper by Kuru et al. identified two FDAA labeling patterns, a D, D-transpeptidase-dependent, concentrated labeling for PG growth and an L, D-transpeptidase-dependent, growth-independent labeling for PG modification along the entire cell envelope (Kuru et al., 2019). The polar FDAA foci we presented do not match the reported pattern of L, D-transpeptidase-dependent incorporation.

      Second, we provided the evidence in Fig. 3b that the polar FDAA labeling is due to the activity of PBP2, a D, D-transpeptidase, because mecillinam that inhibits PBP2 is sufficient to abolish polar FDAA incorporation.

      (2) Second, the evidence provided that this is polar elongation is not sufficient to prove elongation. In all images claiming polar growth, the FDAA focus appears as a single spot, which corresponds to the MreB aggregate visible in bright-field. That might indicate incorporation, but it does not demonstrate polar elongation. To prove this, the authors should do different-length pulses of FDAAs and demonstrate an increasing length of labeled PG along the cell. Cells with one focus should show increasing length; polar foci should elongate from both ends. I find the current 2-color labeling insufficient, as BADA labels the entire cell.

      We appreciate this comment and totally agree with the reviewer. The wide BADA labeling band could indeed come from L, D-transpeptidases. We will follow the reviewer’s recommendation to address this comment with additional staining experiments.

      Small points:

      (3) Lines 350- 302: "In this case, as the nonpolar region is no longer the growth zone, the established cylindrical PG structure is sufficient to maintain cell width, where MreB filaments become nonessential." Does polar elongation give robustness to rod shape? The authors should include an analysis of cell width and its variation within a cell and between cells.

      We appreciate this comment. Judging from the bright-field images we presented, we believe that polar elongation does give robustness to rod shape. We will follow the reviewer’s recommendation and provide the said analysis.

      (4) In the abstract: "This reprogrammed growth mode bypasses the requirement for MreB filaments, highlighting a plasticity of the Rod system that suggests polar elongation may have emerged through the evolutionary loss of MreB." This argument should not be made without evolutionary analysis or reference to such work indicating this is the case.

      We appreciate this comment and will remove this statement from our revised manuscript.

      (5) 365-367 "PG-depleted spheroplasts can spontaneously regenerate rod shape through curvature-dependent localization of MreB filaments [21, 57]". This should be amended. The Billing paper did indeed study spheroplasts, but the Hussain paper used teichoic acid-depleted cells that still had a cell wall.

      We thank the reviewer for pointing out this mistake. We will amend this statement in our revised manuscript.

      Reviewer #2 (Public review):

      Summary:

      Based on observations of localisation of MreBEc at the poles within an aggregate-like structure, upon heterologous expression of MreBMx, the authors set out to investigate how this non-canonical localisation of MreBs leads to a reprogramming of peptidoglycan synthesis to the poles. This is analogous to the polar growth observed in phyla which are not dependent on dispersed growth of PG, but only at the poles, and are MreB independent.

      The authors proceed to establish that PG synthesis is MreB-dependent, Rod enzyme-dependent, and requires the prior establishment of a pole.

      Strengths:

      (1) It is a very interesting idea to design experiments to demonstrate reprogramming of non-polar to polar growth based on the observation of localisation of a heterologously expressed MreB.

      (2) The experiments to demonstrate the factors that determine polar growth and the observation of the PG in each of these experimental situations are convincing.

      (3) I find the observation of an extra layer of PG in the heterologously expressed system very intriguing. It will be interesting to see if this layer merges with the other PG layer at some stage or branches from the non-polar growth near the poles.

      Weaknesses:

      (1) It is not clear what exactly the identity of the polar aggregates is and how much of this activity is an artefact of partially functional MreBs.

      We appreciate this comment and will address it by further explaining our results in the revised manuscript. While we can only say that the polar aggregates resemble inclusion bodies, we do believe that they cause polar PG growth because in the cells that express MreB<sub>Mx</sub>, polar PG growth does not occur at the poles that lack MreB aggregates.

      (2) I find it intriguing that the localisation and growth are predominantly at one pole only. It is unclear to me how this can be reconciled with growth and shape maintenance, and an increase in length and width. Is the increase in length and width a consequence of misshapen cells that are bulged in the absence of a normal PG layer?

      We appreciate this comment. Judging from the bright-field images we presented, we believe that polar elongation does not generate bulges and is thus sufficient for maintaining rod shape. We will follow the reviewer’s recommendation to clarify this.

      (3) The authors do not follow up on the observations in the first figure on the length and width changes and the extra peptidoglycan layer (which I feel are the most interesting aspects), and how this can be connected to the polar growth observed in the later sections of the manuscript.

      We appreciate this comment. We believe that the thickened PG patches are integral parts of the polar PG, rather than an extra layer, which is, however, technically challenging to prove. Thus, we will relay on fluorescence microscopy to visualize polar PG growth.

      (4) The claim that this could be a precursor of an MreB-independent polar growth mechanism appears to be a bit far-fetched, because the system is still dependent on having an established pole for PG synthesis to occur in the new place.

      We appreciate this comment and will remove such speculations from our revised manuscript.

      Reviewer #3 (Public review):

      Summary:

      Most rod-shaped bacteria grow by one of two mechanisms: growth from the pole or growth from the midcell. It is rare for a single species to utilize both modes of growth, although a few examples do exist. Here, the authors have artificially induced E. coli cells to grow from the poles, either by expressing mreB from Myxoccocus xanthus in E. coli, leading to the mislocalization of MreB to the poles in large aggregates, or by forcing the localization of major cell wall synthesis proteins to the cell pole. The fact that cells switched modes of growth suggests an evolutionary pathway from midcell to polar growing cells as well as suggests that there might be unknown conditions in nature when cells may switch growth modes.

      Strengths:

      (1) The authors use a strain that has replaced mreB with a functional fluorescent version at the native site. This eliminates any effects of having two copies of mreB. Because MreB is fluorescently tagged, they can monitor its localization when mreB from M. xanthus is expressed in E. coli. They notice that MreBec now forms bright polar foci and that there appear to be changes to the cell wall at the pole.

      (2) D-amino acids are specific to the cell wall, and fluorescent versions (FDAA) have been used to mark sites of new cell wall insertion. The authors use these FDAAs to determine how cell wall synthesis correlates to MreB and if that changes when MreBmx is expressed. Again, there is pretty clear evidence that cell wall synthesis follows MreB localization to the pole.

      (3) MreB itself does not synthesize the cell wall, but localizes the proteins, such as PBP2, that do. Using a published method to force proteins to the pole, the authors show that when they target PBP2 to the pole, they can phenocopy the polar growth seen when MreB is polar. Interestingly, these cells become resistant to A22, a drug that targets MreB, suggesting that localized growth at the pole does not require MreB and is sufficient to maintain rod shape.

      Weaknesses:

      (1) The authors do not show what the poles of control cells look like, making it difficult to determine if there is a change when MreBmx is expressed. However, the localization of both MreBec and MreBmx clearly forms bright foci at the pole.

      We appreciate this comment and totally agree with the reviewer. We will provide reference images in the revised manuscript.

      (2) While more quantification is needed, the authors show some evidence that RodZ, an MreB interaction partner, is needed for this polar growth, as cells lacking rodZ still form foci at pole-like regions when MreBmx is in the cell; however, these cells remain spherical and do not elongate from these foci.

      These results strongly support our conclusions. In the cells that express MreB<sub>Mx</sub>, MreB<sub>Mx</sub> causes MreB<sub>Ec</sub> to mislocalize, and mislocalized MreB<sub>Ec</sub> recruits Rod enzymes (RodA and PBP2) to poles through the connector protein RodZ. Here when we delete rodZ, while MreB<sub>Ec</sub> still forms aggregates, it is unable to recruit Rod enzymes to those aggregates.

      In contrast, when we directly relocalize PBP2 to cell poles through the PopZ tag, polar PG growth can bypass the requirement for MreB<sub>Ec</sub> filaments (Fig. 4).

      When MreB is deleted, and cells become spherical, the authors were unable to cause the polar growth mode. They suggest that this is due to the lack of a preexisting pole; however, experimental evidence to test this is missing.

      We appreciate this comment. We plan to localize PBP2 to cell poles in an mreB depletion strain to test if cells still remain rods when mreB is depleted.

      Conclusion:

      Overall, the authors do a good job of showing that E. coli can grow with a polar method rather than a midcell method of cell wall insertion. It is unclear why MreBec forms at poles when MreBmx is present and even if this MreB is functional. The foci look similar to inclusion bodies, which are normally aggregates of misfolded proteins that migrate to the poles. Past work has shown that when MreB is more polarly localized, branches form, which is not seen here. Importantly, the authors also show that there is feedback between the localization of MreB and PBP2 as both appear to regulate the localization of the other.

      Because CFP-labeled MreB<sub>Ec</sub> is still fluorescent in the aggregates, we believe that some MreB<sub>Ec</sub> molecules are still correctly folded there, at least on the surfaces of those aggregates.

      Reference

      Kuru, E., Radkov, A., Meng, X., Egan, A., Alvarez, L., Dowson, A., Booher, G., Breukink, E., Roper, D.I., Cava, F., Vollmer, W., Brun, Y., and VanNieuwenhze, M.S. (2019) Mechanisms of incorporation for D-amino acid probes that target peptidoglycan biosynthesis. ACS Chem Biol.

    1. eLife Assessment

      This important study leverages a large global dataset of tens of thousands of tuberculosis samples to place recurrent protein-coding mutations into their three-dimensional structural context, offering an expanded view of how antibiotic resistance emerges compared to traditional genetic analyses alone. The strength of evidence is compelling, supported by the scale and breadth of the dataset and the systematic structural analysis, although some of the assumptions made in the modeling approach are only partially supported. Overall, the work will be of broad interest to researchers studying microbial evolution, antibiotic resistance, and structure-function relationships in pathogens.

    2. Reviewer #1 (Public review):

      Summary:

      In this manuscript, Green et al. attempt to use large-scale protein structure analysis to find signals of selection and clustering related to antibiotic resistance. This was applied to the whole proteome of Mycobacterium tuberculosis, with a specific focus on the smaller set of known antibiotic-resistance-related proteins.

      Strengths:

      The use of geospatial analysis to detect signals of selection and clustering on the structural level is really intriguing. This could have a wider use beyond the AMR-focussed work here and could be applied to a more general evolutionary analysis context. Much of the strength of this work lies in breaking ground into this structural evolution space, something rarely seen in such pathogen data. Additional further research can be done to build on this foundation, and the work presented here will be important for the field.

      The size of the dataset and use of protein structure prediction via AlphaFold, giving such a consistent signal within the dataset, is also of great interest and shows the power of these approaches to allow us to integrate protein structure more confidently into evolution and selection analyses.

      Comments on revised version.

      All my comments from the previous round of reviews have been addressed.

    3. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      This valuable study leverages a large global dataset of tens of thousands of tuberculosis samples to place recurrent protein-coding mutations into their three-dimensional structural context, offering an expanded view of how antibiotic resistance emerges compared to traditional genetic analyses alone. The strength of evidence is convincing, supported by the scale and breadth of the dataset and the systematic structural analysis, although some of the assumptions made in the the modeling approach are only partially supported. Overall, the work will be of broad interest to researchers studying microbial evolution, antibiotic resistance, and structure-function relationships in pathogens.

      We thank the reviewers and editors for their careful critique of our work. We believe the work has been strengthened by addressing the comments and are delighted to submit a revised version. This version has a detailed discussion of prior literature on structure analysis of antibiotic resistance variants in Mycobacterium tuberculosis, more details on dataset origins and data processing to improve reproducibility, better explanations of the evolutionary assumptions underlying the scoring method we developed, and more discussion of proteins that show homoplasic and clustered mutational signals that are not known to confer antibiotic resistance. We include detailed responses to the comments below.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      In this manuscript, Green et al. attempt to use large-scale protein structure analysis to find signals of selection and clustering related to antibiotic resistance. This was applied to the whole proteome of Mycobacterium tuberculosis, with a specific focus on the smaller set of known antibiotic-resistance-related proteins.

      Strengths:

      The use of geospatial analysis to detect signals of selection and clustering on the structural level is really intriguing. This could have a wider use beyond the AMR focussed work here and could be applied to a more general evolutionary analysis context. Much of the strength of this work lies in breaking ground into this structural evolution space, something rarely seen in such pathogen data. Additional further research can be done to build on this foundation, and the work presented here will be important for the field.

      The size of the dataset and use of protein structure prediction via AlphaFold, giving such a consistent signal within the dataset, is also of great interest and shows the power of these approaches to allow us to integrate protein structure more confidently into evolution and selection analyses.

      Weaknesses:

      There are several issues with the evolutionary analysis and assumptions made in the paper, which perhaps overstate the findings, or require refining to take into account other factors that may be at play.

      (1) The focus on antimicrobial resistance (AMR) throughout the paper contains the findings within that lens. This results in a few different weaknesses:

      (a) While the large size of the analysis is highlighted in the abstract and elsewhere, in reality, only a few proteins are studied in depth. These are proteins already associated with AMR by many other studies, somewhat retreading old ground and reducing the novelty.

      (b) Beyond the AMR-associated proteins, the proteome work is of great interest, but only casually interrogated and only in the context of AMR. There appears to be an assumption that all signals of positive selection detected are related to AMR, whereas something like cas10 is part of the CRISPR machinery, a set of proteins often under positive selection, and thus unlikely to be AMR-related.

      We agree that environmental pressures beyond AMR may impose positive selection. In response to the reviewer’s comment we have now included more results and a supplementary figure of the findings about Cas10. We have expanded our results about the proteins found with significant clustering that are not known AMR-associated proteins, and clarified in the discussion that we don’t believe positive selection is caused only by AMR.

      We note that data from homoplasic substitutions in proteins (Figure 1) does indeed support that AMR is the strongest driver of positive selection in Mtb. A challenge is that knowledge about protein function varies in depth by protein, and in Mtb the AMR proteins are among the most well-studied in the proteome. Thus, explanations from literature are most readily available for AMR proteins. Moreover, proteins may exhibit positive selection for more than one reason – for example, we find significant clustering in the proteins GlmM, GlmS, GlmU, and MurA, all of which are involved in amino sugar metabolism, a key component of the cell wall. These proteins could plausibly have a role in AMR via cell wall permeability mechanisms, and could plausibly have a role in adaptation to host environment via the same (or other) mechanisms.

      (2) The strength of the signal from the structural information and the novelty of the structural incorporation into prediction are perhaps overstated.

      (a) A drop of 13% in F1 for a gain of 2% in PPV is quite the trade-off. This is not as indicative of a strong predictor that could be used as the abstract claims. While the approach is novel and this is a good finding for a first attempt at such complex analysis, this is perhaps not as significant as the authors claim.

      (b) In relation to this, there is a lack of situating these findings within the wider research landscape. For instance, the use of structure for predicting resistance has been done, for example, in PncA (https://academic.oup.com/jacamr/article/6/2/dlae037/7630603, https://www.sciencedirect.com/science/article/pii/S1476927125003664, https://www.nature.com/articles/s41598-020-58635-x) and in RpoB (https://www.nature.com/articles/s41598-020-74648-y). These, and other such works, should be acknowledged as the novelty of this work is perhaps not as stark as the authors present it to be.

      We appreciate this comment and have made efforts to better situate our work in the wider context of structure-based prediction of antibiotic resistance. We have included description of and citation to these works and others in the introduction, results, and discussion. A differentiator between our work and previous is that we have trained a predictor across all proteins in the WHO catalogue of known resistance variants, not limited ourselves to a single protein at a time. We feel this makes the case that structure (and specifically proximity to known resistance-conferring variants) is a universally useful feature for resistance mutation prediction. We have also updated our abstract to reflect the exact performance of our method.

      Introduction: “Protein three-dimensional structure has shown utility as an input feature for identifying resistance-conferring variants in known resistance-conferring proteins such as RpoB,[26,27] PncA [28–30], and AtpE.[31]While past work has sought to reannotate parts of the M. tuberculosis proteome with computationally predicted protein structures using older structure prediction methods,[32] we can now infer a protein structure for nearly every protein in the proteome using AlphaFold,[33] leading to new works examining the 3D location of mutations in known and suspected resistance-conferring proteins.[34,35]”

      (3) The authors postulate that neutral AA substitutions would be randomly distributed in the protein structure and thus use random mutations as a negative control to simulate this neutral evolution. However, I am unsure if this is a true negative control for neutral evolution. The vast majority of residues would be under purifying selection, not neutral selection, especially in core proteins like rpoB and gyrA. Therefore, most of these residues would never be mutated in a real-world dataset. Therefore, you are not testing positive selection against neutral selection; you are testing positive against purifying, which will have a much stronger signal. This is likely to, in turn, overestimate the signal of positive selection. This would be better accounted for using a model of neutral evolution, although this is complex and perhaps outside the scope. Still, it needs to be made clear that these negative controls are not representative of neutral evolution.

      The goal of our negative control was to simulate the random accumulation of amino acid substitutions without the effects of selection, which we had originally referred to as “neutral evolution” but is better described as “randomly accumulating substitutions.” We agree that in the absence of antibiotics, essential proteins like RpoB and GyrA are probably under purifying selection, and thus will have depletion of mutations in their hydrophobic cores. We have revised our wording in the results section “A protein-level statistic to test for mutational clustering” to make it clear that our negative control is that of randomly accumulating substitutions. We have included possible extension to more realistic evolutionary scenarios in the discussion.

      As a side note, if we were to compare the observed mutation 3-D pattern to a purifying selection model, that may overestimate the signal compared with the randomly accumulating substitution model that we currently present in the manuscript. Purifying selection would tend to result in slower evolutionary rates than positive selection or randomly accumulation substitutions. So, a control based on purifying selection would have to have fewer mutations to account for the same evolutionary time, and this could lead to underestimating clustering.

      (4) In a similar vein, the use of 15 Å as a cut-off for stating co-localisation feels quite arbitrary. The average radius of a globular protein is about 20 Å, so this could be quite a

      large patch of a protein. I think it may be good to situate the cut-off for a 'single location' within a size estimator of the entire protein, as 15 Å could be a neighbourhood in a large protein, but be the whole protein for smaller ones.

      We interrogated the use of 15 Å as a cutoff and found that it is indeed not very stringent, and functions more as a filter to remove the most egregious examples of proteins lacking single-location clustering. We include a new supplementary figure showing the number of significant hits as the cutoff is varied from 2 Å to 40 Å, and summarize these results in the main text. We note that we in fact find a weak negative relationship between protein length and the distance between the top two residues with highest G-score (R<sup>2</sup> = 0.008, b = -3.1806, p-value = 0.048), the opposite of what would be expected under a scenario the distance between residues is simply driven by protein size and not a signal for clustering.

      Reviewer #2 (Public review):

      Summary:

      This is an important study that, for the first time, systematically places the homoplastic genetic variation observed in the coding regions in a large collection of >31,000 M. tuberculosis samples into the protein structural context. This should be much more informative when, e.g. predicting antimicrobial resistance. The authors imaginatively apply the Getis-Ord score, which originated in geographical spatial analysis but has also been used in human disease to demonstrate that missense mutations in M. tuberculosis known to be associated with antimicrobial resistance are clustered in space. That they are able to consider almost all of the proteome using a large dataset of 31,000 M. tuberculosis complex clinical samples, which makes the evidence convincing.

      Strengths:

      To my knowledge, this is the first study to place the homoplastic missense mutations from a large clinical dataset into their protein structural context and attempt to look for clustering in space, which could be indicative of a recent evolutionary pressure, such as the use of antibiotics. The field usually only views resistance through the genetic paradigm, so it is delightful to see a structural paradigm being brought to bear, as this should, in theory, be much more informative, as protein structure is much closer to function. In addition, the dataset used is large (>31,000 clinical M. tuberculosis samples), and the authors are able to consider almost all of the ORFs (3,687/3,996) in the M. tuberculosis reference, and hence the analysis is comprehensive.

      Weaknesses:

      It is not apparent at the time of this review if the study could be reproduced by other researchers as e.g. whilst the authors state that the raw sequencing files (FASTQ) underpinning the dataset of 31,428 M. tuberculosis isolates can be downloaded the table in the Supplement containing the sample and accession identifiers contains rows that do not contain NCBI accessions e.g. '01R0685' or 'IDR 1600023875' or '1479144813357T181715lib5022nextseqn0035151bp' instead of the expected form e.g. 'SAMEA1016138'. I have searched the NCBI SRA using these terms and got no results, so they cannot be used to download any FASTQ files. There is also no information in the preprint on how the reads were processed (which is a complex process) and the dataset of SNPs subsequently built. One can trace back through the references, but I cannot find anywhere where one can download the SNP dataset, which would permit researchers to reproduce at least the latter stages of the work -- one obvious option would be to make the SNP dataset available. Likewise, the authors have constructed a "M. tuberculosis structureome", which would be very useful for the community but does not appear to be publicly available. At the time of the review, not all the GitHub repositories were public, so these points may have been rectified when that was corrected.

      We have made a number of changes to improve reproducibility of the manuscript.

      First, we have updated Table 1 with additional information to reflect the dataset of origin. While most of the isolates used are available from NCBI (94.5%), the remainder are from other sources. An additional 4.8% are exclusively from PATRIC (now the BV-BRC) and 0.4% are from ENA. Some of the isolates were originally named by their internal identifiers, not their NCBI BioSamples, which has been rectified. Of the three identifiers the reviewer cites, two were originally from Reseq-TB and are now listed with their NCBI BioSamples, and the third, 01R0685, corresponds to one isolate deposited with others in a single BioProject (https://www.ncbi.nlm.nih.gov/bioproject/?term=PRJEB26000; individual isolate available here: https://www.ncbi.nlm.nih.gov/biosample/10125872). We have clarified in the table metadata that the relevant search identifier.

      Second, we have added a new section to our methods about the origins of the Mtb genomic data used in our manuscript, and describing the process of variant calling and SNP dataset construction.

      Third, we have provided the data in a zenodo repository along with instructions for using the protein distance map files provided: 10.5281/zenodo.20766453

      Lastly, we have ensured that the github is publicly available.

      The authors correctly point out in the Introduction that supervised methods like GWAS or ML need datasets with matching genetic and phenotypic drug susceptibility data, which are much difficult/expensive to obtain, but don't then close the loop by comparing their results back to such supervised methods. They pick out RnJ as having previously been identified by a GWAS, but it would have provided a useful validation of their method to e.g. demonstrating that X% of the genes they identify were also identified by GWAS/ML studies, and therefore their method can achieve similar results but without having to collect pDST data.

      We agree that this is a compelling possible extension of the work but due to time constraints have chosen not to pursue it for this manuscript.

      Whilst the authors acknowledge that assuming all sites are equally likely to mutate in their random shuffling procedure is a shortcoming, a bigger weakness is, I suspect, that one should also only consider which amino acids could arise at each codon due to a SNP. Shuffling assumes any amino acid can arise at any codon which is only possible with multiple nucleotide changes, which is possible but highly unlikely.

      Our approach is based on analyzing, in the wild-type protein structure, the 3D location at which mutations occur. In this calculation we do not consider the identity of the amino acid change per se. We have now clarified in the methods that we are not explicitly simulating biochemical change of the wild-type amino acid to any given mutant, rather we are analyzing the wild type amino acid in its structural context. We have added the following text in the methods section: In Computing inter-residue distances, “The EVcouplings Python package was used to compute the distance between wild-type amino acid residues in all protein structures”; in Computing the Getis-Ord score for clustering of homoplastic mutations, “The two values input to the Getis-Ord statistic computation are a per-residue score x, here the per-amino acid homoplasy score, and a weight matrix W that contains the inverse of the inter-residue distances computed from the wildtype amino acids”; and in Preparing GeO score calibration data, “Note that we do not recompute inter-residue distances when simulating mutations in an amino acid, as the distances used as input to GeO score are the wild-type inter-residue distances.” We hope this addresses the reviewer’s concern.

      Finally, the authors implicitly assume that the mutations do not perturb the structure of the proteins, which is likely to be generally true for essential genes but less likely to be true for non-essential genes. This assumption underpins their entire approach and should be borne in mind when evaluating the results.

      Our approach is based on analyzing, in the wild-type protein structure, the 3D location at which missense mutations occur and are observable in a naturally evolving population. It is true that we have not undertaken an analysis of whether any given mutation does or does not perturb the protein structure in which it occurs. However, location alone is a useful piece of information to analyze, as the location of naturally occurring mutations gives a readout of what types of mutations are allowed to persist under natural selection. Among our findings is that mutations display significant clustering even in non-essential genes, which we address in our discussion, “for proteins where mutations that lead to loss of function are known to cause resistance, such as PncA and RsmG (GidB), it is not necessarily expected to find clustering of mutations. We suspect that the observed clustering is due to mutations in a certain region of the protein being more likely to cause loss of function.”

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) It would be good if a more detailed description of the sequencing dataset origins were presented. The Supplementary Table 1, which is meant to hold these data, points to another paper, which in turn points to another paper which does not detail collection strategies for this data. This needs to be clear so that any bias in collection, which could over-inflate selection signals, can be assessed.

      [From public reviews]

      First, we have updated Table 1 with additional information to reflect the dataset of origin. While most of the isolates used are available from the NCBI (94.5%), the remainder are from other sources. An additional 4.8% are exclusively from PATRIC (now the BV-BRC) and 0.4% are from ENA. Some of the isolates were originally named by their internal identifiers, not their NCBI BioSamples, which has been rectified. Of the three identifiers the reviewer cites, two were originally from Reseq-TB and are now listed with their NCBI BioSamples, and the third, 01R0685, corresponds to one isolate deposited with others in a single BioProject (https://www.ncbi.nlm.nih.gov/bioproject/?term=PRJEB26000; individual isolate available here: https://www.ncbi.nlm.nih.gov/biosample/10125872). We have clarified in the table metadata that the relevant search identifier.

      (2) Line 351: The exact download date is missing here.

      We have added the exact download date

      (3) Line 363: Did you mean lowest e-value? Higher would be worse.

      Thank you for catching this, it is indeed the lowest e-value (confusion stemmed from looking at highest negative log e-value).

      Reviewer #2 (Recommendations for the authors):

      (1) The GitHub repo* was not public at the time of review (nor was it listed under user aggreen), so I could not check how reproducible the results are -- please make it public.

      Absolutely, this has been addressed.

      (2) The authors say "we envision structure being added as an additional feature in future work to predict resistance phenotypes from sequences" - this is not true, as some work has already been published predicting resistance in MBTC going back to 2019**

      We have addressed this with wording changes and citations to the mentioned work (see public responses)

      (3) Throughout the term 'non-synonymous' is used; that would include premature stop codons. Would 'missense' be more appropriate?

      You’re correct in pointing out that the term “non-synonymous” is too general for what we mean in this paper. Our analysis included missense mutations and in-frame indels, but did not include premature stop (nonsense) mutations or frameshift mutations. So, the term missense (alone) is narrower than what we wish to convey. We have made clarifications throughout.

      (4) The aminoglycosides are an important, albeit less used, class of antibiotics, and mutations arise in the ribosomal genes, e.g. rrs (which hence do not encode protein). They have therefore been excluded for obvious reasons, but it would help a reader from the tuberculosis field if this were acknowledged. Likewise, a reader might wonder why Rv0678 isn't in Figure 1 - I suspect it is because most of the samples were sequenced before the introduction of bedaquiline, but again, it would help if this were explained.

      See response to (5)

      (5) On a related note, it is not surprising that rpoC appears in Figure 1 due to its role in compensating for the fitness cost that arises when a rifamipicin-resistance mutation occurs in rpoB: obviously not central but a nice "oh yeah that makes sense" point for the reader if it were briefly mentioned.

      These are both great points about the relevance of our results to the Mtb community. We have added an additional paragraph interpreting the results of Figure 1 that mentions the reason for the appearance of RpoC and Cas10, and non-appearance of non-coding genes and genes relevant to resistance to newly introduced and repurposed drugs.

      (6) How is the "minimum coordinate difference" calculated? I assume all the structures are missing hydrogens as usual, so for two amino acids A and B, is it the smallest distance between any pair of heavy atoms from A and B? That would, I assume, introduce some bias for larger amino acids like Trp, or did you calculate from shared atoms like the backbone C_alpha atoms? That in turn will tend to make the distances a bit larger. It would be useful to know, as you explicitly mention a 1.5 nm threshold.

      We have explained this in a new section of the results, “The EVcouplings Python package was used to compute the distance between amino acid residues in all protein structures [49]. The package calculates the distance between all heavy (non-hydrogen) atoms in residue i and residue j, then returns the minimum of those distances.”

      (7) Given the reference used (H37Rv) is Lineage 4, one wonders about deeprooted/phylogenetic mutations, but then I suspect this sentence is doing a lot of that heavy-lifting: "We performed ancestral sequence reconstruction to determine the number of independent arisals of each mutation (homoplasy)". For the more general reader, it would be useful to touch on exactly what you mean and the importance of only considering homoplastic mutations.

      We have expanded our explanation in this section to better make the case for the use of homoplastic variants in our analysis, “Because analyzing the frequency of alleles in a population can be biased by oversampling of particular lineages, and by evolutionary recency, we chose to analyze the number of independent arrivals of each mutation (homoplasy) rather than their population-level frequency. This ensures that more recent evolutionary events are not underrepresented due to lack of time to spread in the population. To accomplish this, we used a previously compiled a dataset of genomes of 31,428 isolates from the Mycobacterium tuberculosis complex (MTBC), with ancestral sequence reconstruction to determine the number of independent arrivals of each mutation (Supplementary Data 1).”

      (8) Whilst this is true: "Evolution-based approaches are an alternative for finding variants associated with antibiotic resistance without requiring resistance phenotype data", the dataset used to, e.g. build the second edition of the WHO catalogue of resistanceassociated variants has >50,000 samples and therefore is larger than the dataset you have analysed here. The last time I looked, there were >100k M. tuberculosis samples in NCBI, and therefore, to be valid, your approach should really use more samples than are available with WGS and pDST data. I appreciate, however, that this will not be possible for this manuscript, but it is an obvious criticism.

      We appreciate this critique and acknowledge that the number of isolates with both WGS and pDST has increased rapidly in recent years. The first edition of the WHO catalogue (2021) used 38,215 isolates, which increased to over 50k in the second edition. The dataset of homoplastic mutations on which we based this paper was originally published in 2021.

      To address your comment, we have softened our assertions in the introduction about the utility of evolution-based approaches, and emphasize instead the different nature of the underlying signal, “Evolution-based approaches are an alternative for finding variants associated with antibiotic resistance by analyzing their mutational frequency and phylogenetic distribution”

      (9) Minor point, but the second sentence in the Introduction ignores that one can diagnose MDR-TB using phenotypic methods as well as genetic methods.

      We have added an additional citation and mention of laboratory phenotypic methods.

      (10) There are a few typos: "genic" "G-sore"

      Addressed.

    1. eLife Assessment

      This study modeled vocal learning in zebra finches with a network of three components: a pathway with delayed/slow Hebbian learning that mimics the HVC-RA projection, a pathway with reinforcement learning that mimics the HVC-BG-RA projection, and a motor unit that mimics the syrinx and produces song output. The model convincingly reproduces key features of song learning and is a valuable step towards understanding how vocal learning is substantiated in the song system. The generic and simplified nature of the model limits its generalization to other forms of motor learning. Further examination of model assumptions and presentation of testable predictions would increase the impact of the study.

    2. Reviewer #1 (Public review):

      The authors sought to devise a model of the song system that captures the essential features of that neural circuitry and couple it to a behavioral model that captures the challenges of motor learning while remaining tractable. They seek to use this model to explain known features of song learning and relate them to the general problems associated with learning via gradient ascent. Their syrinx model uses two control parameters-air sac pressure and syringeal labial tension-to generate birdsong-like spectrograms. Normalized spectrograms generated by the syrinx model are compared to a target spectrogram by computing Pearson's correlation coefficient. This correlation coefficient quantifies the performance of the model, and the goal of learning is to maximize it. The syringeal model, while simplified, is complex enough to generate multiple local maxima in the correlation coefficient within the 2D control space with wide variation in the magnitude of the maxima, making it challenging to find a good optimum via gradient ascent. They also use a more abstract motor model where local maxima are generated by randomly placing Gaussians within the 2D control space. These two approaches to modeling motor space are a major strength of this work.

      The neural model is extremely generic and consists of units (each representing a population of excitatory and inhibitory neurons) with continuous-valued outputs ("firing rate") ranging from -1 to +1 with sigmoidal activation functions. Premotor HVC simply generates a fixed temporal sequence that drives activity in RA (analogous to the primary motor cortex) that constitutes commands to the motor controller that translates RA output into 2D control signals for the syrinx. A second, indirect, pathway from HVC to RA is represented by a single node ("BG"); in songbirds this pathway consists of 3 structures (one of them quite complex and heterogeneous) with recurrent connections (i.e., from LMAN back to Area X). Learning is primarily driven by performance-modulated Hebbian plasticity in HVC-BG connection weights. There is also Hebbian plasticity in HVC-RA weights that are in effect trained by the RA activity patterns driven by BG-RA connections.

      I worry that this model is too simplified to capture essential features of the song system (and cortico-basal ganglia circuits more generally). That problem is most acute in the way the authors model (or fail to model) the anterior forebrain pathway (AFP) through X, DLM, and LMAN; their model in effect reduces the AFP to just LMAN. I do not think that invalidates this study, but there is a significant danger that this model will end up missing the mark in some important way relative to a more realistic model. However, there is another flaw that comes close to doing that-the connection weights between the units are allowed to vary from -1 to +1. That means, for example, that the connections between HVC and RA units can be inhibitory and can flip between excitation and inhibition during learning. I can imagine some hand-wavy justifications for this (e.g., the connection becomes "inhibitory" because excitation to inhibitory neurons becomes stronger than that to excitatory neurons within the unit), but I can't easily imagine one that I would find persuasive.

      A key feature of this model is "synaptic volatility" in the connections between HVC and BG. In songbirds, performance often improves steadily throughout the day, then deteriorates after sleep, a feature which the authors suggest is an important method for avoiding getting stuck on relatively low-performing local optima. I find this suggestion to be intriguing and reasonably persuasive. To implement this in their model, after a simulated "day" of practice, the HVC-BG connection weights are partially randomized (synaptic volatility). There is nothing intrinsically wrong with this idea, but the authors imply that there is experimental support for this phenomenon, which they relate to "continuous remodeling of the cortico-striatal synapses with volatility over hours to days." None of the papers cited really support this kind of synaptic randomization. However, given the simplicity of the authors' neural model, the HVC-BG weights are probably the only place this randomization can be implemented. The authors' implementation does not just add noise to HVC-BG weights "overnight"; it is also scaled inversely by the magnitude of accumulated weight change through the day. The authors do not appear to provide a justification for this aspect of their synaptic volatility.

      If we accept the authors' model as detailed and accurate enough for their purposes, it does explain several features of song learning and relates them to solving general problems of learning through gradient ascent. This is particularly true of the daily deterioration of performance and how that relates to escaping local maxima. They show that their model reproduces some of the known effects of lesioning the motor (HVC-RA) and anterior forebrain (HVC-BG-RA) pathways and how those effects depend on the current stage of song learning. They also explore the implications of delayed maturation of the HVC-RA pathway, exemplified by a gradual increase in HVC-RA weights. However, it is not clear that this actually happens; the one paper they cite in support of this contention (Mooney and Rao, 1994) shows no such thing (it does show that HVC axons enter RA later in development than LMAN axons do). Moreover, it's not clear how essential this could be given that some songbird species continue to show profound vocal plasticity throughout their lives and presumably long after their HVC-RA pathway has fully matured. The authors compare their dual pathway model to a single pathway model (an AFP-only model, in effect); a very illuminating comparison that demonstrates the advantage of the dual pathway. They close by showing that dual pathway performance is robust under variation of key parameters.

    3. Reviewer #2 (Public review):

      Summary:

      The authors describe a computational model for the acquisition of sensorimotor skills and explore these dynamics using vocal learning in the zebra finch, a system rich in experimental data, to describe the developmental trajectory of vocal imitation by trial and error. They set up their model as a dual-pathway system, with a cortical pathway that drives the vocal effector and a basal ganglia (BG) pathway that uses dopamine-mediated reinforcement learning (RL) by gradient descent to optimize the vocal imitation process.

      Strengths:

      A key strength of the model, due in part to the fact that his model was generated by a computational laboratory that has also contributed significantly to the collection of experiment-driven empirical data, is that the model is biologically constrained and incorporates a considerable amount of experimental data, including some of the latest findings in the field. In addition to providing a compelling model for the acquisition of vocal learning, this biologically based RL model outperforms many current models. A key feature of this model, which makes it unique, is the implementation of a synaptic volatility variable within the BG pathway that aims to mimic published work showing that juvenile birds exhibit post-sleep deterioration. The model uses a motor output to drive a biophysical model of the avian vocal organ (syrinx) and explores not only the ability to copy song acoustic units (syllables) but also the underlying neural dynamics in both the cortical and BG pathways, showing that each converges onto the types of neural activity patterns that are observed experimentally. Because of the richness of experimental data in this system, the authors can perform "computational experiments" where they can block sleep-driven synaptic volatility or lesion various pathways to replicate experimental observations.

      In addition to providing important computational insights to our understanding of vocal learning in the songbird, this study provides key insights into the general architectures that are optimal for RL by gradient descent. These include the conclusion that effective RL requires adaptive regulation of exploration and exploitation, that cortical consolidation must occur at a slower timescale than BG-driven exploration, and intriguingly that the introduction of synaptic volatility prevents RL models of incomplete learning by getting "stuck" in local minima.

      Weaknesses:

      In the methods section, the authors state "... HVC and RA layers are fully connected, as are the HVC and BG layers. Synaptic weights in these pathways are plastic, reflecting activity-dependent plasticity at RA and BG synapses." Unless I missed it, it is unclear how much the authors consider the synaptic differences between HVC and BG inputs to RA. This seems like an important feature to highlight, especially given that HVC-RA connections are primarily AMPA-mediated whereas those from LMAN are predominantly NMDA. The authors should be clearer about how they model these synapses and better highlight (and describe) the importance of these synaptic differences in their modeling efforts. Ideally, they should evaluate whether the differences in synapse type influence the outcome of their model. It would be interesting, for example, to test the effect on learning of synaptic conductance substitution (i.e., replacing NMADA with AMPA) on the LMAN-RA synapse.

      The model focuses exclusively on the interaction of two converging pathways, and learning is based purely on acoustic feature properties of what seem like four independent syllables of similar or identical duration. For this model, this is fine. But it would be helpful for the authors to state more clearly that they are not modeling respiratory influences on syllable production, which include amplitude modulation of the syllables and expiratory pulse duration. It should be noted that the authors do not (unless I missed it) mention the existence or role of recurrent loops in song initiation (and possibly syllable sequencing). They should at least mention this in the discussion, perhaps as a limitation and item for future versions of the model.

    4. Reviewer #3 (Public review):

      Summary:

      This study modeled vocal learning in zebra finches with a network of three components: a pathway with delayed/slow Hebbian learning that mimics the HVC-RA projection, a pathway with reinforcement learning that mimics the HVC-BG-RA projection, and a motor unit that mimics the syrinx and produces song output. The model convincingly reproduces key features of song learning and is a valuable step towards understanding how vocal learning is substantiated in the song system. Further examination of model assumptions and presentation of testable predictions would increase the impact of the study.

      Strengths:

      (1) The model reproduces several key features of song learning, including learning in a non-convex performance landscape, decreasing motor variability during learning, the relative importance of the HVC-RA and HVC-BG-RA pathways at different learning stages, and sleep-related deterioration.

      (2) The model incorporates several key physiological properties of the song system, including performance-dependent dopamine signals to the BG, the neural variability in the system, and multiple global/local optima of the motor production landscape.

      (3) The study convincingly demonstrates the advantages of a dual-pathway network over a single-pathway one.

      (4) The model demonstrates the counterintuitive benefit of sleep-related performance deterioration for facilitating the escape from local optima to reach the global optimum.

      (5) The model seems robust to some variations in model parameters and task structure.

      Weaknesses:

      (1) The study could be more impactful if the model can generate new, testable predictions. The predictions provided in the Discussion are not well justified. For example, with the HVC spine turnover, it seems unlikely that BG lesions would abolish overnight performance deterioration. Because the volatility term is inversely related to learning during the day, the changes in LMAN and RA during sleep are not necessarily larger than those during the day.

      (2) The section "Neural activity patterns in the model parallels song system neurophysiology" seems fully anticipated because the model construction is based on the known neural activity patterns. Are there new testable predictions from the model?

      (3) Certain model assumptions lack explanation or justification.<br /> a) The authors treat "the delayed maturation of the cortical pathway" as an important component. However, it is unclear if/how this component was implemented in the model. If it was not included in the model, the authors should remove statements related to the delayed maturation idea.<br /> b) How critical is the inverse relationship between learning and sleep-deterioration? If it is known that sleep-related deterioration is inversely related to learning during the previous day, a citation should be added. Similarly, it should be clarified if the spine turnover observed in HVC depends on previous plastic changes, like the assumed inverse relationship in the model.<br /> c) Spine turnover has been demonstrated in HVC and not yet in Area X, but the model implements volatility only in the HVC-BG pathway. It would be important to compare the effects of sleep-related volatility in the HVC-RA and HVC-BG-RA projections.<br /> d) In Table 1/Figure 8, the learning rate for HVC-BG is 10^4 times bigger than the learning rate for HVC-RA. What is the biological justification for this difference?

      (4) Related to #3, it is unclear how changing those assumptions would affect model performance.

      (5) It is not explained/shown why cross-day exploration (with sleep, sporadic) is better than continuous, non-sleep-related ones. Figure 7B presents results with different noise levels, which may approximate non-sleep-related, continuous volatility, but only for the single-pathway model. Comparable simulations by adding continuous volatility in the dual-pathway model would be helpful. For example, would just a bigger intrinsic noise in either pathway confer the same benefit in escaping local optima, e.g., epsilon-greedy exploration? If the authors can establish the advantages of sleep-specific volatility in its particular form (based on day learning) and relate it to other sensorimotor learning behaviors, it could increase the general impact of the study.

      (6) The model description can be improved.<br /> a) What are mBG and mRA in Eqs 6/7?<br /> b) Where is the learning rate specified in Equations 1-8?<br /> c) It is unexplained why Equations 1-3 are not in the same form: Equations 1 and 2 normalize by activity, but Equation 3 normalizes by weight (Equation 8 also normalizes by weight). The authors should confirm that these equations are correct.<br /> d) How J_HVC is generated should be defined.<br /> e) How is R calculated in Equation 10? Does this model maintain a PPE for each syllable or for the overall performance? Would it make any difference in learning?<br /> f) Are the weight updates in Equations 9 and 13 added at different time points (e.g., immediately after each spike versus after a syllable). This should be clarified (perhaps a diagram would help).<br /> g) How are the optimal threshold and slope determined for Equation 14?<br /> h) Parameters in Equation 20 are not defined.

    5. Author response:

      We thank all three reviewers for their detailed and constructive reviews, which will help us improve the manuscript. Concerning the simplicity of the model, we will elaborate on the limitations that arise from the modelling choices and better justify these choices. With that in mind, we believe this level of abstraction is well suited to the study’s objective. As a rate-coded model, it is designed to capture interactions between multiple circuit pathways, and lays the foundation to delve into neuronal dynamics in the future for further insight, once the network-level questions we address here - regarding the exploration-exploitation tradeoff and evasion of sub-optimal performance convergence - are well examined. For this purpose, we believe the model’s simplified network-level perspective is not a bug, but a feature. We will also explore in depth why and how the dual pathway model performs better in non-convex sensorimotor landscapes, building on prior analytical work that speaks to the question (Sankar, Leblois & Rougier, ICDL 2022). We will further substantiate and discuss our choice concerning the potential mechanisms underlying the proposed overnight synaptic volatility. We will include comparable approaches (including relevant machine learning analogues) in the introduction. Finally, we will clarify our interpretation of the sign of synaptic weights. We hope to incorporate all valuable feedback by the reviewers in the revised manuscript.

      References

      Remya Sankar, Arthur Leblois, Nicolas P. Rougier (2022). Dual pathway architecture underlying vocal learning in songbirds. In IEEE International Conference on Development and Learning, ICDL 2022, London, United Kingdom, September 12-15, 2022. pages 265-271, IEEE, 2022.

    1. eLife Assessment

      The authors investigate how dominance hierarchy shapes defensive strategies in mice under two naturalistic threats: a transient visual looming stimulus and a sustained live rat. This study provides important insights into how social context and dominance hierarchy modulate innate defensive behaviors across distinct naturalistic threats. The strength of evidence is convincing, with detailed classification and analysis of behaviors.

    2. Reviewer #1 (Public review):

      Summary:

      This study presents an interesting behavioral paradigm and reveals interactive effects of social hierarchy and threat type on defensive behaviors. However, addressing the aforementioned points regarding methodological detail, rigor in behavioral classification, depth of result interpretation, and focus of the discussion is essential to strengthen the reliability and impact of the conclusions in a revised manuscript.

      Strengths:

      The paper is logically sound, featuring detailed classification and analysis of behaviors, with a focus on behavioral categories and transitions, thereby establishing a relatively robust research framework.

      Comments on revised version.

      I think the authors have addressed all of my comments.

    3. Reviewer #2 (Public review):

      Summary

      The authors examine how dominance hierarchy modulates defensive strategies in mice exposed to two naturalistic threats: a transient visual looming stimulus and a sustained live rat. By comparing single versus paired testing conditions, they demonstrate that social presence attenuates fear responses, and that dominant and subordinate mice display distinct behavioral and social patterns depending on threat type. The study offers a rich behavioral dataset and a potentially valuable framework for investigating hierarchical influences on innate fear.

      Strengths

      (1) The use of two ecologically relevant threat paradigms allows for meaningful comparisons across transient and sustained contexts.

      (2) Behavioral quantification is thorough, incorporating manual annotation of multiple behavior types and transition‑matrix analyses.

      (3) The comparison between dominant and subordinate pairs is novel within the innate‑fear literature.

      (5) The manuscript is well structured and clearly written, with figures that are visually informative and effectively support the main conclusions.

      Weaknesses

      The investigation of neural mechanisms underlying the observed behavioral effects remains limited.

    4. Reviewer #3 (Public review):

      Summary:

      This study examines how dominance hierarchy influences innate defensive behaviors in pair-housed male mice exposed to two types of naturalistic threats: a transient looming stimulus and a sustained live rat. The authors show that social presence reduces fear-related behaviors and promotes active defense, with dominant mice benefiting more prominently. They also demonstrate that threat exposure reinforces social roles and increases group cohesion. The work highlights the bidirectional interaction between social structure and defensive behavior.

      Strengths:

      This study makes a valuable contribution to behavioral neuroscience through its well-designed examination of socially modulated fear. A key strength is the use of two ethologically relevant threat paradigms - a transient looming stimulus and a sustained live predator, enabling a nuanced comparison of defensive behaviors. The experimental design is robust, systematically comparing animals tested alone versus with their cage mate to cleanly isolate social effects. The behavioral analysis is sophisticated, employing detailed transition maps that reveal how social context reshapes behavioral sequences, going beyond simple duration measurements. The finding that social modulation is rank-dependent adds significant depth, linking social hierarchy to adaptive defense strategies. Furthermore, the demonstration that threat exposure reciprocally enhances social cohesion provides a compelling systems-level perspective. Together, these elements establish a strong behavioral framework for future investigations into the neural circuits underlying socially modulated innate fear.

      Comments on revised version.

      The authors have addressed the initial major criticism regarding the lack of causal evidence for neural mechanisms, which has alleviated our concerns. This provides a more solid behavioral foundation for future investigations into neural circuit mechanisms.

    5. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This study presents an interesting behavioral paradigm and reveals interactive effects of social hierarchy and threat type on defensive behaviors. However, addressing the aforementioned points regarding methodological detail, rigor in behavioral classification, depth of result interpretation, and focus of the discussion is essential to strengthen the reliability and impact of the conclusions in a revised manuscript.

      Strengths:

      The paper is logically sound, featuring detailed classification and analysis of behaviors, with a focus on behavioral categories and transitions, thereby establishing a relatively robust research framework.

      Weaknesses:

      Several points require clarification or further revision.

      (1) Methods and Terminology Regarding Social Hierarchy:

      The study uses the tube test to determine subordinate status, but the methodological description is quite brief. Please provide a more detailed account of the experimental procedure and the criteria used for determination.

      We have included more details about how the tube test was performed in the revised manuscript. Social rank within each mouse pair was determined using a standard tube test paradigm. To minimize stress, mice were pair-housed for at least two weeks with a 15-cm tube placed in their home cage to allow voluntary exploration. Prior to rank assessment, mice were trained to traverse a 30-cm tube over two consecutive days (10 trials per day), with alternating entry from either end to prevent side bias. On the test day, each mouse pair underwent up to seven competitive trials in the same 30-cm tube. In each trial, the two mice were simultaneously released from opposite ends of the tube. A “win” was defined as one mouse successfully advancing through the tube while the opponent retreated completely out of the tube (all four paws outside) for at least 5 seconds. The first mouse to achieve four wins was designated as the dominant individual, whereas the opponent was classified as subordinate. Social rank stability was reassessed one day after threat exposure using the same criteria. Only pairs with consistent ranks were included in subsequent analyses.

      The dominance hierarchy is established based on pairs of mice. However, the use of terms like "group cohesion" - typically applied to larger groups - to describe dyadic interactions seems overstated. Please revise the terminology to more accurately reflect the pairwise experimental setup.

      Thanks for the comment. We have replaced the term “group cohesion” with “social engagement”.

      (2) Criteria and Validity of Behavioral Classification:

      The criteria for classifying mouse behaviors (e.g., passive defense, active defense) are not sufficiently clear. Please explicitly state the operational definitions and distinguishing features for each behavioral category.

      Passive defense was defined as an immobility-based defensive strategy characterized by suppression of locomotor activity, including freezing and tail rattling. Active defense was defined as movement- or posture-dependent defensive strategy, including approach, investigation, withdrawal, and stretch-attend. We have clarified these in the revised manuscript.

      How was the meaningfulness and distinctness of these behavioral categories ensured to avoid overlap? For instance, based on Figure 3E, is "active defense" synonymous with "investigative defense," involving movement to the near region followed by return to the far region? This requires clearer delineation.

      Defensive behaviors in the rat exposure paradigm were grouped into two categories: passive and active defense, each comprising distinct behaviors. All the manually annotated behaviors were mutually exclusive; that is, each video frame was assigned a single behavioral label to avoid overlap across behaviors. Active defense includes four behaviors: approach, investigation, withdrawal, and stretch-attend. We have clarified these points in the revised manuscript.

      The current analysis focuses on a few core behaviors, while other recorded behaviors appear less relevant. Please clarify the principles for selecting or categorizing all recorded behaviors.

      Thank you for pointing this out. In the current study, we focused primarily on defensive and social behaviors. We also included several neutral solitary behaviors related to anxiety and defensive state, such as sniffing, grooming, and rearing, which were consistently expressed across animals and closely linked to our main findings. We have clarified these in the revised manuscript.

      (3) Interpretation of Key Findings and Mechanistic Insights:

      Looming exposure increased the proportion of proactive bouts in the dominant zone but decreased it in the subordinate zone (Figure 4G), with a similar trend during rat exposure. Please provide a potential explanation for this consistent pattern. Does this consistency arise from shared neural mechanisms, or do different behavioral strategies converge to produce similar outputs under both threats?

      Thanks for bringing up this important question. The consistent increase in proactive bouts in dominant mice across both paradigms suggests a consistent rank-dependent reorganization of dyadic interaction under threats. We propose that this convergence reflect a shared neural mechanism that links defensive state with social-rank information, potentially involving top-down regulation from the mPFC to threat-specific midbrain and hypothalamic defensive circuits. We have expanded the discussion to incorporate this explanation.

      (4) Support for Claims and Study Limitations:

      The manuscript states that this work addresses a gap by showing defensive responses are jointly shaped by threat type and social rank, emphasizing survival-critical behaviors over fear or stress alone. However, it is possible that the behavioral differences stem from varying degrees of danger perception rather than purely strategic choices. This warrants a clear description and a deeper discussion to address this possibility.

      We thank the reviewer for this insightful comment. We agree that, in principle, behavioral differences could arise from variations in perceived danger rather than strategic choice. In humans, decisions can sometimes reflect value-based strategies that override perceived danger. In contrast, under naturalistic threat conditions, mice likely rely predominantly on danger perception to make behavioral decisions, and such responses are expected to be consistent with value-based strategies shaped by natural selection. In the revised manuscript, we have expanded the Discussion to address the role of threat perception and its relationship to decision-making in our behavioral paradigms.

      The Discussion section proposes numerous brain regions potentially involved in fear and social regulation. As this is a behavioral study, the extensive speculation on specific neural circuitry involvement, without supporting neuroscience data, appears insufficiently grounded and somewhat vague. It is recommended to focus the discussion more on the implications of the behavioral findings themselves or to explicitly frame these neural hypotheses as directions for future research.

      We have revised the Discussion to focus more directly on behavioral findings and added explicit neural hypotheses as potential future directions.

      Reviewer #2 (Public review):

      Summary:

      The authors investigate how dominance hierarchy shapes defensive strategies in mice under two naturalistic threats: a transient visual looming stimulus and a sustained live rat. By comparing single versus paired testing, they report that social presence attenuates fear and that dominant and subordinate mice exhibit different patterns of defensive and social behaviors depending on threat type. The work provides a rich behavioral dataset and a potentially useful framework for studying hierarchical modulation of innate fear.

      Strengths:

      (1) The study uses two ecologically meaningful threat paradigms, allowing comparison across transient and sustained threat contexts.

      (2) Behavioral quantification is detailed, with manual annotation of multiple behavior types and transition-matrix level analysis.

      (3) The comparison of dominant versus subordinate pairs is novel in the context of innate fear.

      (4) The manuscript is well-organized and clearly written.

      (5) Figures are visually informative and support major claims.

      Weaknesses:

      Lack of neural mechanism insights.

      The current study focused on behavior. In the revised manuscript, we have incorporated a discussion of potential neural mechanisms and highlight this as an important direction for future work.

      Reviewer #3 (Public review):

      Summary:

      This study examines how dominance hierarchy influences innate defensive behaviors in pair-housed male mice exposed to two types of naturalistic threats: a transient looming stimulus and a sustained live rat. The authors show that social presence reduces fear-related behaviors and promotes active defense, with dominant mice benefiting more prominently. They also demonstrate that threat exposure reinforces social roles and increases group cohesion. The work highlights the bidirectional interaction between social structure and defensive behavior.

      Strengths:

      This study makes a valuable contribution to behavioral neuroscience through its well-designed examination of socially modulated fear. A key strength is the use of two ethologically relevant threat paradigms - a transient looming stimulus and a sustained live predator, enabling a nuanced comparison of defensive behaviors. The experimental design is robust, systematically comparing animals tested alone versus with their cage mate to cleanly isolate social effects. The behavioral analysis is sophisticated, employing detailed transition maps that reveal how social context reshapes behavioral sequences, going beyond simple duration measurements. The finding that social modulation is rank-dependent adds significant depth, linking social hierarchy to adaptive defense strategies. Furthermore, the demonstration that threat exposure reciprocally enhances social cohesion provides a compelling systems-level perspective. Together, these elements establish a strong behavioral framework for future investigations into the neural circuits underlying socially modulated innate fear.

      Weaknesses:

      The study exhibits several limitations. The neural mechanism proposed is speculative, as the study provides no causal evidence.

      Establishing causal evidence for neural mechanisms is beyond the scope of the current behavioral study. We highlight this as an important direction for future work in the revised manuscript.

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      (1) Clarify the definitions of all behavioral categories (escape, assessment, passive defense, active defense, etc.) in detail.

      We have clarified the definitions of all behavioral categories in detail in the revised manuscript.

      (2) Reduce speculative statements about SC-VMHdm-mPFC circuitry, and provide the data about this circuit, if possible.

      We have reduced the speculative statements about this circuitry in the revised manuscript.

      Reviewer #3 (Recommendations for the authors):

      We commend the authors on a carefully executed and conceptually clear study that makes a valuable contribution to behavioral neuroscience. Below are some issues and suggestions for this manuscript:

      (1) Please add the following to the discussion: the reason why, compared to looming exposure, social behaviors were more often followed by defensive behavior during rat exposure.

      We have added this discussion to the revised manuscript. This difference reflects the distinct temporal characteristics of the two threats. Following the transient looming stimulus, the threat rapidly ceases once the stimulus ends, reducing the need for sustained defensive behavior. Consequently, social interactions primarily occur after threat termination and are less frequently interleaved with subsequent defensive behaviors. Consistent with this interpretation, looming exposure selectively increased the duration of social behavior in subordinate mice. In contrast, rat exposure represents a sustained multisensory threat that maintains a persistently elevated defensive state. Under these conditions, both the frequency and duration of social interactions increased in dominant and subordinate mice. Moreover, huddling emerged as the predominant social behavior, and the frequent transitions between freezing and huddling suggest that social interactions become integrated with ongoing defensive responses, potentially serving as a safety-seeking or cohesive defense during sustained threat.

      (2) Figures 1B and 1C showed that Grooming has significantly decreased. Please verify whether the statistical methods and results are correct.

      The grooming data in the original Figure 1C did not distinguish dominant and subordinate mice and did not show significant decrease. We have removed it in the revised manuscript. In original Figure 2I (current Figure 1P), the social modulation on grooming behavior was observed only in dominant mice (Two-way ANOVA with post hoc Tukey’s range test).

      (3) To investigate how social context modulates the expression and progression of defensive responses, the authors analyzed behaviors in two time windows: the early phase (0-5 seconds after stimulus onset) and the late phase (20-60 seconds after onset) in Figure2. What is the rationale for selecting 5 seconds as the cutoff for the early-phase behavioral analysis? Were the behaviors of mice between 5 and 20 seconds also analyzed?

      The reason to select 5 seconds as the cutoff for early-phase behavioral analysis is that most behavioral decisions are made within this time window. We also analyzed the behaviors of mice between 5 and 20 seconds and have integrated them into the revised manuscript.

      (4) Figure 2 demonstrated that the social context attenuates looming-evoked defensive behavior in a rank-dependent manner. Then, I would like to ask if the influence of the social context on rat-evoked defensive behavior is also present in a rank-dependent manner? Discuss the similarities and differences between the social context's effect on looming-evoked defensive behavior and rat behavior.

      The influence of the social context on rat-evoked defensive behavior is also present in a rank-dependent manner. Specifically, total stretch-attend (SA) time and SA frequency were increased only in dominant mice (revised Figures 2Q and 2R), whereas average SA duration was decreased only in subordinate mice (revised Figure 2S). Approach-investigation-withdraw (AIW) frequency was increased only in dominant mice while approach speed was increased only in subordinate mice (revised Figures 2U and 2V).

      Social context exerts a broad protective influence across threat types, but the form of this modulation differs depending on the nature of the threat. Similarity: social presence consistently alleviates threat-induced stress and reshapes defensive behavior in a rank-dependent manner, with dominant benefiting more strongly under both threats, suggesting higher social rank may associate with greater flexibility to integrate social modulation. Difference: the behavioral outcomes of social regulation are distinct across threat types. During looming, social presence primarily suppresses immediate defensive responses and alleviates post-looming anxiety, suggesting under transient and unpredictable threat, social context dampens acute defense and facilitates behavioral recovery. In contrast, during the sustained rat exposure, social presence promotes a shift in defensive strategy from passive to active defense, rather than mere suppression of defensive output. These results suggest that social context flexibly adjusts defensive behavior according to ecological demands by reducing excessive defense and anxiety in response to transient looming and facilitating active coping when threatened by a sustained live predator. These discussions have been included to the revised manuscript.

      (5) In Figure 4, looming exposure increased these two measures only in subordinate individuals, whereas rat exposure affected both ranks, indicating threat-specific modulation of social behaviors. The visual looming paradigm primarily simulates visual stimuli triggered by aerial predators, while the rat exposure stimulus involves not only visual cues but also other sensory inputs, such as olfactory signals for the experimental animals. This may lead to differences in the defensive behaviors of mice, particularly increasing the proportion of proactive behaviors in dominant individuals. If olfactory information transmission is blocked during rat exposure, can the behavioral phenotypes shown in Figure 4 still be observed?

      The rationale for incorporating both looming and rat exposure paradigms was to mimic the distinct, ethologically relevant predator encounters by rodents in natural environments. As the reviewer pointed out, the multimodality nature of the rat threat may contribute to differences in the defensive behaviors displayed by mice. This point is also acknowledged in our manuscript, where we state that rat exposure “imposes prolonged stress and elicits a broader repertoire of defensive behaviors”.

      We agree that systematically dissecting the contribution of specific sensory modalities (e.g., olfactory, visual, or auditory) to defensive behaviors during rat exposure is an interesting and important question. Based on prior literature, we speculate that olfactory cues likely play a major role. However, the main aim of the present study is to investigate how social regulation of defensive behaviors depends on dominance hierarchy and threat type, rather than isolating modality-specific sensory mechanisms. Within this framework, we preserved the multisensory features of the rat exposure to maintain its ecological validity and to emphasize its distinction from the visual-only looming threat. We thank the reviewer for raising this insightful question, but it is beyond the scope of current study. We will consider it as an important direction for future investigations.

      (6) Discuss the potential mechanisms for the similarities and differences in the impact of the looming threat and the rat threat on cohesive behavior?

      We have expanded the Discussion to address the potential neural mechanisms underlying both the similarities and differences in the effects of looming and rat threats on social cohesion. As discussed in response to issue (1), we propose that the behavioral differences between the two threats arise from their distinct nature. Here, we further discuss a potential circuit mechanism. Specifically, we propose that the mPFC serves as a common hub integrating social context and dominance-related information with threat processing, thereby contributing to the shared enhancement of social engagement under both threats. We further speculate that the distinct patterns of social behavior may arise from partially distinct defensive circuits. SC-centered visual threat circuits may facilitate rapid post-threat social engagement by recruiting vigilance-related networks, whereas VMHdm-centered predator-defense circuits may promote sustained social cohesion by engaging neural circuits that support coordinated coping during persistent defensive states. We emphasize that these are hypotheses, and that future studies combining circuit-level recordings and causal perturbations within the behavioral framework established here will be required to test them.

    1. eLife Assessment

      This fundamental work substantially advances our understanding of short-term plasticity mechanisms by providing evidence for release-independent low-frequency synaptic depression that reflects a model that proposes a redistribution of vesicles within the readily releasable pool, via a reduction in docking site occupancy due to vesicle undocking. The evidence supporting this model is convincing, with rigorous modeling, electrophysiological data and computational analysis. The work will be of broad interest to cellular neuroscientists and synaptic physiologists.

    2. Reviewer #1 (Public review):

      Summary:

      In this work, the authors investigate the mechanisms of low-frequency synaptic depression at cerebellar parallel fiber to interneuron synapses using unitary recordings that allow direct quantification of synaptic vesicle release. They show that sparse stimulation can induce robust synaptic depression even in the absence of substantial vesicle consumption, and that this depressed state is rapidly reversed when stimulation frequency is increased. To account for these observations, the authors propose a model in which low-frequency depression reflects a redistribution of vesicles within the readily releasable pool, in particular a reduction in docking site occupancy due to vesicle undocking.

      Strengths:

      I found the experimental work to be of high quality throughout. The use of simple synapse recordings to count individual vesicle release events is particularly powerful in this context and allows questions to be addressed that are difficult to approach with more conventional approaches. The demonstration that low-frequency depression can occur independently of prior vesicle release, together with the rapid recovery observed during high-frequency stimulation, places strong constraints on possible underlying mechanisms and represents a clear strength of the study.

      The modeling framework is clearly laid out and helps organize a broad set of observations across stimulation frequencies. Several of the experimental tests appear well motivated by the model, including the recovery train experiments, the analysis of failures, and the use of doublet stimulation. Taken together, the data provide a coherent phenomenological description of low-frequency depression and its relationship to vesicle availability within the readily releasable pool.

      Weaknesses:

      The major concerns raised during the initial review have been successfully addressed through substantial revisions of both the manuscript and the presentation of the model. The distinction between experimental observations and model-based interpretation is now considerably clearer, and the discussion more appropriately reflects the explanatory nature of the modeling framework.

    3. Reviewer #2 (Public review):

      Summary:

      Silva and co-workers exploit their previously established methods of analyzing release events at single parallel fiber to molecular layer interneuron synapses. They observed synaptic depression at low transmission frequencies (< 5 Hz) which rapidly recovers during high-frequency transmission. Analysis of the time course of low-frequency depression revealed an initial rapid and a slow linearly increasing time course. Strikingly, the initial depression occurred even in the absence of proceeding release arguing against vesicle depletion as the underlying mechanism.

      Strengths:

      The main strength of the study is the careful demonstration of an interesting synaptic phenomenon challenging the classical vesicle-centered interpretation of synaptic depression.

      Weaknesses:

      There are no weaknesses.

    4. Reviewer #3 (Public review):

      Summary:

      The manuscript builds on the observation that, at some synapses, low-frequency stimulation causes synaptic depression which can be reversed by subsequent high-frequency stimulation. Such low-frequency depression (LFD) cannot be easily explained by the depletion of a single vesicle pool. Here, Silva and colleagues propose a model of activity-dependent vesicle trafficking to explain LFD at synapses between cerebellar granule cells and molecular layer interneurons.

      Strengths:

      Overall, LFD is interesting and worthy of examination, and the authors provide new experimental results that are of the high quality expected from this group.

      Weaknesses:

      The study proposes a novel model of vesicle trafficking that is not explained by known biological mechanisms, and the manuscript does not adequately compare or discuss alternative models.

      I have several concerns about how the authors interpret the data. First, the manuscript's primary conceptual advance is the idea that LFD involves vesicle undocking, rather than depletion. However, most experiments were performed under conditions that promote vesicle depletion (3 mM extracellular Ca2+). When experiments were repeated in physiological Ca2+, there appeared to be little or no LFD (stats are not provided). Second, the RS/DS/DU/undocking model, though not outside the realm of possibility, is not readily explained by known mechanisms and is only loosely supported by experimental findings. Third, when simulating LFD, the authors do not compare alternative models and use inappropriate language to imply that a model fit represents the truth (e.g. "the finding of identical experimental and simulated values confirms that the undocking mechanism accounts for LFD"). Finally, the model is presented in an overly complicated manner. The sheer amount of terms and nomenclature makes the manuscript confusing and difficult to read. Overall, the manuscript would benefit from added experiments and more statistics, a better justification and evaluation of the model, and more nuanced language.

      Comments on revised version.

      I appreciate the authors' detailed responses to my initial review of the manuscript. My suggestions reflect a sincere attempt to improve the clarity and accuracy of the paper. I disagree with some of the authors conclusions and their responses to my comments, but do not wish to make any further suggestions. The authors are entitled to their own views.

      A final note: If the authors want to refrain from making definitive statements on underlying cellular mechanisms, they should consider amending the title of the manuscript.

    5. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      In this work, the authors investigate the mechanisms of low-frequency synaptic depression at cerebellar parallel fiber to interneuron synapses using unitary recordings that allow direct quantification of synaptic vesicle release. They show that sparse stimulation can induce robust synaptic depression even in the absence of substantial vesicle consumption, and that this depressed state is rapidly reversed when stimulation frequency is increased. To account for these observations, the authors propose a model in which low-frequency depression reflects a redistribution of vesicles within the readily releasable pool, in particular, a reduction in docking site occupancy due to vesicle undocking.

      Strengths:

      I found the experimental work to be of high quality throughout. The use of simple synapse recordings to count individual vesicle release events is particularly powerful in this context and allows questions to be addressed that are difficult to approach with more conventional approaches. The demonstration that low-frequency depression can occur independently of prior vesicle release, together with the rapid recovery observed during high-frequency stimulation, places strong constraints on possible underlying mechanisms and represents a clear strength of the study.

      The modelling framework is clearly laid out and helps organize a broad set of observations across stimulation frequencies. Several of the experimental tests appear well-motivated by the model, including the recovery train experiments, the analysis of failures, and the use of doublet stimulation. Taken together, the data provide a coherent phenomenological description of low-frequency depression and its relationship to vesicle availability within the readily releasable pool.

      We thank the Reviewer for his positive assessment of our work.

      Weaknesses:

      While the experimental results are strong, the manuscript would benefit from rebalancing the strength of the mechanistic conclusions drawn from the modelling in light of its limitations. The framework is clearly useful and provides a coherent interpretation of the data, but it is not uniquely constrained by the experimental observations, and alternative models or interpretations could plausibly account for the findings. The use of different model regimes concatenated across time, with substantially different parameter values, highlights the abstract nature of the approach. For these reasons, the model seems best presented as one plausible explanatory framework rather than a definitive biological mechanism. Clarifying the distinction between data-driven observations and model-based inferences would help readers assess which conclusions are strongly supported and which remain more speculative.

      The interpretation of the Ca2+-related experiments would benefit from more cautious wording. The absence of detectable changes in presynaptic Ca2+ signals does not exclude more localized or subtle Ca2+-dependent mechanisms, and conclusions regarding Ca2+ independence should therefore be framed accordingly. In addition, while low-frequency depression is still observed at reduced extracellular Ca2+, these experiments appear less diagnostic of the specific model-derived mechanism emphasized elsewhere in the manuscript - namely, a selective reduction in docking-site occupancy - and should be discussed with appropriate qualification in the text.

      Concerning Ca2+ signals, the Reviewer is right. While we found no change in Ca2+ signalling apart from a slow Ca2+ accumulation during long trains at 1 Hz, the possibility of an undetected change cannot be excluded. We have added a word of caution in this direction on p. 11. Concerning the 1.5 mM Ca2+ experiments, the Reviewer presumably alludes to the first recovery train (yellow) point in Supplementary Fig. 2C. This is also the last point (s11) of the slow train at 0.5 Hz because no delay at all was interposed between the slow train and the recovery train. We have now included one more experiment (with a present total number n = 6), and we have corrected Fig. S2C accordingly. In the new version the depression measured for s4-s10 vs s1 during the 0.5 Hz trains is 0.69 +/- 0.05 (p = 0.00058, paired one-tail t-test). The ratio of the s1 value of the recovery train compared to control s1 is 0.83 +/- 0.08 (p = 0.028, paired one-tail t-test).

      Major points:

      (1) Clarify and qualify mechanistic claims derived from the model.

      Throughout the manuscript, changes in model parameters are at times described as if they directly reflected underlying physiological mechanisms. As a result, the conceptual distinction between experimentally observed phenomena, model-derived variables, and biological interpretation is not always clear. Several conclusions in the Results and Discussion are phrased as mechanistic statements, although they rest on assumptions intrinsic to the modelling framework. The authors should systematically review the text and explicitly distinguish between (i) experimentally observed changes in synaptic responses and (ii) inferences about vesicle docking states or transitions within the model.

      In particular, statements implying that vesicle undocking is the mechanism underlying low-frequency depression should be rephrased to reflect that this is an interpretation within the proposed framework rather than a uniquely demonstrated biological process. For example, statements such as "Low-frequency depression is caused by synaptic vesicle undocking" should be replaced with formulations such as "Within the framework of our model, low-frequency depression is accounted for by a redistribution of synaptic vesicles away from docking sites" or "Our results are consistent with a model in which changes in vesicle docking-state occupancy contribute to low-frequency depression."

      A particularly problematic example is the statement that "these experiments further confirm that LFD only involves a decrease in δ, without accompanying changes in ρ or IP size." Here, an experimentally defined phenomenon (LFD) is directly equated with changes in model-derived variables. Such statements should be revised to make clear that δ, ρ, and IP size are inferred quantities within the model, and that the experimental data are interpreted through this framework rather than directly confirming changes in these parameters. Similarly, overgeneralizing statements such as "Undocking therefore represents the key mechanism controlling short-term depression across stimulation frequencies" should be softened to reflect that this conclusion emerges from the model rather than from direct experimental evidence.

      As suggested, we clarify the distinction in the revised version between experimental data and modelling, and we refrain from making definitive statements on underlying cellular mechanisms.

      (2) Address the biological interpretation of time-dependent model regimes.

      The model relies on distinct parameter regimes applied at different time points, with some transitions effectively suppressed in certain regimes. While this approach captures the data well, its biological interpretation remains unclear. The authors should either (i) expand the discussion to outline plausible biological processes that could give rise to such regime changes (for example, calcium-dependent modulation of transition rates or activity-dependent changes in vesicle state stability), or (ii) more explicitly frame this aspect of the model as a descriptive abstraction rather than a mechanistic proposal. This further underscores the need to clearly separate the descriptive role of the model from claims about underlying biological mechanisms.

      We thank the Reviewer for drawing our attention to this important point. Below 10 ms, rate constants are largely determined by the large-amplitude, fast-decaying Ca2+ signal occurring near voltage-dependent Ca2+ channels (‘Ca2+ nanodomain’). After 10 ms, the rate constants depend on the low-amplitude, slowly decaying Ca2+ signals averaged over the entire varicosity (‘volume-averaged Ca2+’). We explain this better in the revised version (Materials and Methods, p. 21).

      (3) Reframe conclusions drawn from calcium-related experiments.

      The calcium imaging data demonstrate no detectable changes in the measured presynaptic calcium signals under the tested conditions, but they do not rule out that calcium signals contribute in ways undetectable by the assay. Conclusions should therefore be revised to reflect this limitation, avoiding statements that exclude a role for calcium-dependent mechanisms. Wording such as "we did not detect evidence for..." would be more appropriate than conclusions implying the absence of an effect.

      Similarly, while low-frequency depression is still observed at reduced extracellular calcium (1.5 mM Ca<sup>2+</sup>), the specific mechanistic signature emphasized elsewhere in the manuscript - namely a selectively reduced first response during a high-frequency recovery train - is no longer apparent. These experiments should therefore be discussed as consistent with the proposed framework, but not as providing independent support for a selective reduction in docking-site occupancy. Explicitly acknowledging this limitation would improve clarity and avoid overinterpreting these data.

      This has been discussed above (‘weaknesses’).

      (4) Soften interpretations based on non-significant comparisons.

      In several places, comparisons that do not reach statistical significance are used to argue for equivalence between conditions (for example, comparisons involving failure versus non-failure trials or different LFD conditions). These conclusions should be revised to emphasize the limits of statistical power and framed as a lack of evidence for a difference rather than evidence of independence.

      We have amended this point in the revised version.

      Reviewer #2 (Public review):

      Summary:

      Silva and co-workers exploit their previously established methods of analyzing release events at single parallel fiber to molecular layer interneuron synapses. They observed synaptic depression at low transmission frequencies (< 5 Hz), which rapidly recovers during high-frequency transmission. Analysis of the time course of low-frequency depression revealed an initial rapid and a slow linearly increasing time course. Strikingly, the initial depression occurred even in the absence of preceding release, arguing against vesicle depletion as the underlying mechanism.

      Strengths:

      The main strength of the study is the careful demonstration of an interesting synaptic phenomenon challenging the classical vesicle-centered interpretation of synaptic depression.

      We thank the Reviewer for his positive assessment of our work.

      Weaknesses:

      No major weaknesses were identified by this reviewer.

      The finding of release-independent synaptic depression is important and would have widespread implications. Therefore, some more analyses to increase the confidence in these findings could be performed.

      My concern is whether rundown could explain the findings. If the rate of failures in s1 increases and at the same time the amplitude decreases during the experiments, an apparent depression in s2 could arise. The Supplementary Figure 5A addresses run-down, but the figure is not easy to understand, and, as far as I understood, it does not address the question of whether the release-independent depression could be caused by a rundown. To address this, the analysis of Figure 5 could be repeated by investigating the failure rate and amplitude separately or by analyzing the 1st and 2nd half of the recordings separately.

      The Reviewer makes a very important point that had escaped our attention. If the responses were declining over the course of an experiment, near the end of the recordings, a high proportion of failures would be associated with a weak response to the second AP. This could distort the relation between initial failures and amount of LFD, perhaps to the point of indicating LFD after failures when there were none. As suggested by the Reviewer, we tested this possibility by examining the stability of the synaptic responses during experiments. We found a mean s<sub>1</sub> value of 0.87 ± 0.13 for the first half of the experiments used in Fig. 5, and of 1.10 ± 0.17 for the second half (p > 0.05, n = 10). This analysis shows that there was no rundown during these experiments. We show in Author response image 1 a plot of s1 as a function of train number in these experiments, for two examples as well as for the average across frequencies and across experiments. These plots do not suggest any artefactual correlation between failures, mean s1, and rundown.

      Author response image 1.

      Plot of s1 as a function of train number for the experiments of Fig. 5

      In response to a request of Reviewer 2, Author response image 1 illustrates the evolution of s1 values as a function of train number for the experiments used to produce Figure 5. In each experiment, about 20 s1 values were obtained at two ISIs (either 10 ms and 500 ms, or 800 ms and 1600 ms). Author response image 1 shows two examples of s1 values as a function of train number (these values fluctuate widely between 0 and 3), and the average across cells and ISI values. There is no indication of a rundown of s1 values as a function of train number.

      Reviewer #3 (Public review):

      Summary:

      The manuscript builds on the observation that, at some synapses, low-frequency stimulation causes synaptic depression, which can be reversed by subsequent high-frequency stimulation. Such low-frequency depression (LFD) cannot be easily explained by the depletion of a single vesicle pool. Here, Silva and colleagues propose a model of activity-dependent vesicle trafficking to explain LFD at synapses between cerebellar granule cells and molecular layer interneurons.

      Strengths:

      Overall, LFD is interesting and worthy of examination, and the authors provide new experimental results that are of the high quality expected from this group.

      Weaknesses:

      The study proposes a novel model of vesicle trafficking that is not explained by known biological mechanisms, and the manuscript does not adequately compare or discuss alternative models.

      I have several concerns about how the authors interpret the data. First, the manuscript's primary conceptual advance is the idea that LFD involves vesicle undocking, rather than depletion. However, most experiments were performed under conditions that promote vesicle depletion (3 mM extracellular Ca2+). When experiments were repeated in physiological Ca2+, there appeared to be little or no LFD (stats are not provided). Second, the RS/DS/DU/undocking model, though not outside the realm of possibility, is not readily explained by known mechanisms and is only loosely supported by experimental findings. Third, when simulating LFD, the authors do not compare alternative models and use inappropriate language to imply that a model fit represents the truth (e.g., "the finding of identical experimental and simulated values confirms that the undocking mechanism accounts for LFD"). Finally, the model is presented in an overly complicated manner. The sheer amount of terms and nomenclature makes the manuscript confusing and difficult to read. Overall, the manuscript would benefit from added experiments and more statistics, a better justification and evaluation of the model, and more nuanced language.

      We respectfully disagree with these sweeping criticisms, as described in more detail below.

      Major concerns:

      (1) Most experiments were performed under conditions that exacerbate depletion

      In order to attribute LFD to vesicle undocking rather than depletion, it is important to show LFD under conditions where depletion is minimal. As mentioned above, the authors only report significant LFD in elevated extracellular Ca2+. In a small number of experiments performed in more physiological Ca2+ (1.5 mM), there is no depression after a single stimulus, and it is not clear that there was statistically significant depression during a low-frequency train. Several studies cited in support of LFD share this problem:

      Abrahamsson et al., (2007) recorded from Schaffer collaterals in 4 mM Ca, 3-4X physiological Ca2+.

      Doussau et al., (2010) recorded from Aplysia synapses in 3X Ca compared to seawater.

      Rudolph et al., (2011) is cited as an example of LFD. However, this study performed experiments at high release probability cerebellar climbing fibers, and reported depression that increased monotonically with stimulation frequency, so it does not resemble the phenomenon studied in this paper. Lin et al., (2022) also largely describe monotonic depression at the calyx.

      The Reviewer suggests that LFD may only occur under non-physiological conditions, if the release probability has been increased by artificially elevating the extracellular Ca2+.

      The implication is that LFD is at best a curiosity with little or no significance for brain signalling. We disagree with this point of view for several reasons.

      Concerning the statement ‘In order to attribute LFD to vesicle undocking rather than depletion, it is important to show LFD under conditions where depletion is minimal’: This is the purpose of the analysis shown in Fig. 5.

      The statement ‘the authors only report significant LFD in elevated extracellular Ca2+’ is inaccurate. Fig. S2C shows a clear LFD in 1.5 mM Ca2+, as acknowledged by Reviewer 1 (‘low-frequency depression is still observed at reduced extracellular Ca2+’). However, we failed to provide a p-value for the depression in the initial version of the paper (p was 0.004, n = 5 with this data set; paired t-test, one-tail). In the revised version, we document the 1.5 mM results more extensively, including the incorporation of the results of an additional experiment, and an explicit statistical analysis of the data (p = 0.00058, n = 6; paired t-test, one-tail).

      Concerning the statement ‘there is no depression after a single stimulus’: We find that the onset kinetics of LFD is slower in 1.5 Ca2+ than in 3 Ca2+ (respectively 1.8 ISI and 0.51 ISI, Fig. 2C and Fig. S2C). This explains that the PPR is not significantly <1 in 1.5 Ca2+ without implying any weakening of the extent of LFD at steady state.

      As explained in the manuscript (p. 5), in a previous work, we developed a method to ascribe changes in SV pools, within the RS/DS model, with specific modifications of s1, s2 and s5-s8 during test 100 Hz trains (Tran et al., 2022). This method was developed in 3 mM Ca2+ conditions, and for this reason, we performed most experiments for the present work in 3 mM Ca2+.

      Chiu and Carter (2024) demonstrated LFD in neocortical synapses; they performed their study in 1.2 mM Ca2+, not in elevated Ca2+.

      Rudolph et al., (2011) showed low-frequency depression not only in elevated external Ca2+, but also in 0.5 mM Ca2+. While Rudolph et al., (2011) did not make an explicit link between their observations and LFD, there is no reason to doubt that these observations are an example of LFD. They showed a biphasic depression when switching the stimulation frequency from 0.05 Hz to 2 Hz. In one of the founding papers of LFD, Doussau et al., (2010) describe a biphasic depression when switching the stimulation frequency from 0.025 Hz to 1 Hz; Fig. 1 of the two papers (Rudolph 2011 and Doussau 2010) are strikingly similar.

      Lin et al., (2022) would probably not agree with the statement that the depression at the calyx is ‘largely monotonic’, as they stress the finding of quasi-constant depression between 5 and 50 Hz.

      The authors note that their results differ from those of Atluri and Regehr, but do not mention that a possible reason for the difference is the increased release probability in their experiments.

      In fact, we clearly listed the difference in external Ca2+ as a likely source of the discrepancy by saying ‘This discrepancy presumably stems from differences in experimental conditions (room temperature, stimulation of multiple presynaptic PFs and 2 mM external Ca<sup>2+</sup> concentration in the previous work, vs. near-physiological temperature, single presynaptic stimulation and 3 mM external Ca<sup>2+</sup> here)’.

      The authors should provide statistics for the data obtained in 1.5 mM Ca, and discuss why LFD is increased in conditions that also elevate vesicle release probability.

      See our comments above: the revised version includes the requested statistics. On p. 6 of the manuscript, we do provide an explanation for the apparent lack of LFD at 1.5 Ca2+ and 2 Hz, namely a superimposition of LFD with facilitation. At 1.5 Ca2+ and 0.5 Hz, our LFD numbers are not weaker than at 3 mM Ca2+ and 0.5 Hz of 1 Hz.

      Altogether, it is correct that many LFD experiments have been carried out in high release probability synapses and/or under conditions of elevated Ca2+. However, the reasons underlying these choices are diverse (in our case, to build on the previous SV pool analysis developed in Tran et al., 2022 in 3 Ca2+ conditions) and do not imply a limitation to the phenomenon. LFD is present in physiological conditions for low-to-moderate release probability synapses (as shown in our work), and altogether, there is no reason to dismiss LFD as non-physiological.

      (2) Lack of biological mechanisms supporting the model

      The model is presented without compelling biological support. The evidence in support of vesicle undocking comes from experiments by the Watanabe lab, which showed fewer-than expected docked vesicles under EM when cultured synapses were stimulated immediately prior to high-pressure freezing. Kusick et al were careful to note that these vesicles may have been lost to fusion.

      The Watanabe lab showed an SV deficit at docking sites at times ranging from about 100 ms to several seconds (Kusick et al., 2020, their Fig. 5E). This corresponds to the ISI values where we see paired-pulse depression. In their Summary, Kusick et al., raise the possibility of SV fusion as an alternative to undocking at the 100 ms time point. But the same issue had previously been considered in Miki et al., 2018 with other techniques (their Fig. 2d), where it was shown that the SV deficit seen in paired-pulse experiments could not be explained by fusion. This leaves undocking as the most likely explanation, at least in our preparation. We have added a new paragraph on p. 14 to clarify this point.

      The putative undocking Kusick describes is immediate (< 5 ms after stimulation), and it was not shown to be Ca2+ sensitive. This manuscript describes "calcium-dependent undocking" that proceeds from 10 ms - 200 ms. Multiple studies from the Watanabe lab show that a single stimulus lowers the number of docked vesicles, and subsequently, there is a transient redocking of vesicles that can be blocked by EGTA or Syt7 knockout.

      This is not an accurate description of the Kusick results or of our results. In the Kusick paper, the SV deficit seen at <5 ms after stimulation is attributed to exocytosis, not to undocking. Clearly, it is Ca2+-dependent. Our manuscript describes potential calcium-dependent undocking not during the time 10 ms- 150 ms, during which our undocking rate is assumed to be calcium-independent, but starting at 150 ms and lasting a few hundred ms thereafter.

      I also question the rationale for the authors' model that 2 vesicles are coupled in series to a single release site. Previous papers from this lab cited EM studies from frog and neuromuscular that showed filamentous connections between vesicles (do these synapses show LFD?). Here, the authors primarily cite their previous models to support their arguments. I encourage them to continue searching for ultrastructural evidence for 2-vesicle-docking-units and to cite such studies.

      It is important to remember that our sequential two-step model was not based on EM data, but on a series of functional data including variance-mean analysis of summed SV release numbers; covariance analysis among subsequent SV release numbers; analysis of release latencies as a function of stimulus number during an AP train; analysis of SV release numbers under conditions of very high release probability. We note that the phenomenon of Ca2+-dependent docking that we proposed based on these observations has been consistent with flash-and-freeze or zap-and-freeze results from several laboratories. Concerning potential filamentous connections between SVs and the AZ plasma membrane at a distance of several 10s of nm, this has been seen not only in frog or mice neuromuscular junctions, but also at brain synapses (ex: Siksou et al., Journal of Neuroscience 2007; Cole et al., Journal of Neuroscience 2016; Fernandez-Busnadiego, Journal of Cell Biology 2010; 2013).

      (3) Comparison to other vesicle models

      The authors use overly assertive language to suggest that the model proves a mechanism. "Altogether, these results indicate that the slow phase of LFD ... reflects a δ decrease without significant changes in pr, in ρ or in IP size". Simulating data does not conclusively "indicate" the underlying mechanism, but the authors could state their data can be "explained by a model where..".

      Please see our response above to a similar point by Reviewer 1.

      However, LFD does not require activity-dependent undocking. Instead, the phenomenon has been explained by high-release probability, paired with an activity-dependent increase in either docking or release probability (Chiu and Carter, 2024; Doussau et al., 2017). Does the new model do a better job of replicating some facet of the data? If multiple models can explain the same data, how can we determine which model is correct? The "Alternative Presynaptic Depression Mechanisms" should be expanded to discuss these issues.

      We could not find statements in the Chiu and Carter paper or in the Doussau et al., paper explaining LFD ‘by high-release probability, paired with an activity-dependent increase in either docking or release probability’. As far as we can see, Chiu and Carter do not propose any specific mechanism for LFD, beyond saying that depression and facilitation must be separate. Doussau et al., (their Fig. 6) clearly frame their interpretation in a sequential two-step model. As in the preceding Miki et al., paper (which they cite extensively), they assume a rapid (a few ms), Ca-dependent transition between their ‘reluctant pool’ and their ‘fully-releasable pool’, respectively homologous to RS and DS. Thus, the Doussau et al., interpretation is close to that presented in our present work, even though significant differences exist. An important difference is that Doussau et al., did not use simple synapses, so that they did not have access to key synaptic parameters such as the number of docking sites or the release probability per docking site. Consequently, the model in Doussau et al., does not have the same level of detail as ours. The revised version explains better the differences and similarity between the models of Doussau et al., and that exposed in our work (new paragraph on p. 14).

      Recommendations for the authors:

      Reviewing Editor Comments:

      Three reviewers have seen the paper. They agree that the evidence for low-frequency depression (LFD) is solid, but they all suggest that the data with 1.5 mM extracellular calcium should be analyzed and discussed in more detail. The finding of release-independent depression is surprising, but further data analysis to investigate whether it could be caused by run-down is required. The reviewers think the study will be more complete if the authors add N's and stats for LFD in 1.5 mM Ca. In addition, the most important change would be adding clarity in the text about the limitations and simplifications of the model. Please edit the text to clearly state that the presented model provides one possible framework to account for the data, and that other models might also be plausible, and clearly describe the limitations and simplifications of the model.

      Thank you for your positive assessment of our work. As explained in more detail below, we have added the requested statistics for 1.5 mM Ca data. We have also thoroughly revised our text to separate more clearly facts and interpretation. Finally, we now explain in more detail the limitations of our model.

      Reviewer #1 (Recommendations for the authors):

      Minor points

      (1) Statistical analysis and reporting.

      Please justify the use of one-tailed statistical tests and consider whether two-tailed tests would be more appropriate in some cases.

      In Materials and Methods (p. 17), we explain that one-tailed statistics were used when the sign of any possible deviation was predicted; otherwise two-tailed statistics were used.

      Exact p-values should be reported consistently rather than using threshold notation.

      We have corrected this as far as we could. Exact p-values that are still missing will be provided in the version of record.

      For analyses using rank-based statistics (e.g., calcium imaging and recovery experiments), please clarify which comparisons were paired and which were unpaired, and ensure that this is clearly indicated in the Methods and figure legends.

      We have now clearly indicated whether these statistics were paired or unpaired.

      (2) Figure clarity and presentation.

      In Figure 1, clarify whether the responses shown in panel C represent measured EPSCs or derived quantities, as this is not immediately clear from the labelling.

      Thank you for spotting this, the figure has been corrected.

      In Figure 2, please show the variability of control responses used for normalization.

      The variability of control responses used for normalization is now clearly explained in Materials and Methods (p. 18).

      In Figure 3, the use of similar color schemes for different model states and for data-model comparisons is confusing and should be revised.

      Thank you for spotting this; has been corrected.

      In Figures 4 and 5, consider reintroducing or clarifying schematic representations of the model states referred to in the text to aid reader comprehension.

      Figure 4 had already such representations. We have now modified these representations to make them simpler and clearer. We considered adding some model to Fig. 5 but could not come to any satisfactory option.

      (3) Controls and interpretation of pharmacological experiments.

      For experiments involving intracellular BAPTA, it would strengthen the interpretation to either include or explicitly discuss positive controls demonstrating that the manipulation was effective under the recording conditions. Alternatively, loosen the interpretation of those data.

      We have modified the text as suggested by the Reviewer.

      (4) Terminology and consistency.

      Ensure consistent use of terminology for vesicle states and model variables throughout the text and figures, and clarify whether different labels (e.g., RS/DS versus alternative notations) refer to equivalent states.

      We have revised our manuscript as suggested.

      (5) Model fit at short inter-stimulus intervals.

      In Figure 3B, the model appears to capture the overall trend of facilitation but shows a noticeable deviation from the data at the shortest inter-stimulus intervals. It would be helpful to clarify whether this reflects limitations of the current parameterization, known simplifications in the model (e.g., assumptions about fast calcium-dependent processes), or variability across recordings. Briefly commenting on this mismatch would help readers interpret the strengths and limits of the model in this regime.

      Because of the statistical nature of data (quantal fluctuations) and simulations (Monte Carlo trials), some differences between data and simulations are unavoidable. The low simulation point at 10 ms ISI in Fig. 3B may in addition result from an imperfect transition between the two time domains (time domains 1 and 2) of the simulation. We have added a sentence in Materials and Methods (p. 22) to explain this.

      (6) Model parameterization and fitting.

      Please clarify how model parameters were estimated, including the fitting procedure, cost function, and any constraints applied. It would be helpful to report confidence intervals or other measures of uncertainty for key fixed parameters, and to indicate how variable these parameters are across cells or recordings. In addition, a brief discussion of how sensitive the main conclusions are to variation in key parameters (such as release probability or transition rates) would help readers assess the robustness and identifiability of the model-derived interpretations.

      In addition, it is not entirely clear how model parameters and transitions differ across the concatenated temporal regimes used in the simulations, nor which parameters are held fixed versus altered or effectively suppressed. Summarizing, for each regime, the parameter values (or ranges) and active transitions-either in a table or schematic-would greatly improve transparency and help readers assess parameter identifiability and the robustness of the model-derived conclusions.

      For each group of experiments, results under control conditions (8-APs at 100 Hz) were averaged together, leading to the determination of one set of parameters (N, ⍴, δ, r, s, p<sup>r</sup>). Next the parameters for the second time domain were obtained by trial and error to optimize the fit for Fig. 1 (PPR) data. The parameters were constrained such the PPR recovered to 1 with a 10-second ISI. Unfortunately, the parameter domain was too complex to make an unsupervised fitting procedure practical. We have now added a paragraph on p. 22 to indicate which model parameters were heavily constrained by data and which ones were less constrained. Individual synapses were not simulated. Regarding the different time domains, the probabilities rb and sb are introduced for the second time domain and rf and sf are changed as shown in Fig. Supp. 3. The kinetic parameters are reset at the beginning of each simulated AP.

      (7) Clarify the scope of the model with respect to the slow component of low-frequency depression.

      The manuscript identifies a slow component of low-frequency depression that persists over long timescales. While this is an interesting observation, it is not fully clear whether this component is intended to be captured by the current model or represents an additional process outside its scope. Making this distinction explicit would help readers better grasp the scope and limitations of the model.

      As stated in our manuscript, modelling the slow component of LFD falls outside the scope of the present work. The set of rate constants in Fig. S3 does not produce any slow component, and we failed to find a combination of rate parameters that would produce a slow component. This is now clearly stated in the Materials and Methods section.

      Reviewer #2 (Recommendations for the authors):

      I have only the following suggestions to further improve the manuscript.

      (1) It is argued that LFD is alleviated when using doublets rather than singlets (Figure 8), but is the steady-state depression in s1 between singlets and doublets at the end of the experiment statistically different?

      It is unclear what should be considered steady state here. Average s<sub>1</sub> values are smaller for singlets than for doublets over the last 25 points (0.58 +/- 0.05 and 0.74 +/- 0.05, p = 0.02, unpaired t-test).

      (2) I do not understand Supplementary Figure 4. Does the value of zero in this plot indicate that the stimulation does not evoke any release anymore? How can this be differentiated from rundown? Can the LFD be reversed by a high-frequency stimulation?

      We have rewritten the description of this figure. We have also modified the figure labelling. Finally, we have changed the size and color of the symbol at zero delay to stress that the linear fit is constrained to include this point. Altogether we trust that these changes have clarified the presentation of these results.

      (3) The findings are related to Kusick et al., (2020). However, I think that in Kusick et al., (2020) as well as in Watanabe et al., (2013; doi:10.1038/nature12809), all tested time points ranging from a few milliseconds to several seconds after a stimulus show vesicle depletion except the time point of 14 ms in Kusick et al., (2020). I think these findings are, therefore, inconsistent with the much slower observed LFD.

      The increase in docked SV# at 14 ms is reported not only in Kusick et al., 2020 but also in Wu et al., 2023 as well as in Ogunmowo et al., 2025. Later, at 100 ms, the number of docked SVs drops again (Kusick et al., 2020, their Fig. 5 d-e, and Ogunmowo et al., 2025), to finally recover on a time scale of seconds (Kusick et al., 2020). These data fit with our observations if LFD is due to undocking, as we propose.

      (4) Page 13, 4th paragraph:

      Eshra et al., (2021) provide experimental evidence for "a Ca2+ -dependent rate of SV entry into the RRP".

      The Reviewer is right: Eshra et al., actually reported a weak Ca2+-dependence of the replenishment. We have therefore eliminated the Eshra reference in this sentence (note that this paper is mentioned elsewhere in the manuscript).

      Reviewer #3 (Recommendations for the authors):

      (1) Complicated terminology: I had a hard time understanding the model due to terms like δ, ρ, IP, DS, RS, RS gate, etc. Anything the authors can do to reduce the number of acronyms and simplify their description/depiction of the model will be helpful for readers.

      This resembles recommendation (4) of Reviewer 1. As stated in our response above, we have entirely revised the manuscript with these issues in mind. In addition, the abbreviation list at the onset of the manuscript should help the readers.

      (2) Figure 5 may represent the strongest evidence in support of an undocking model. This is the smallest figure, and I found it harder to understand. For example, I was initially confused because the blue "fail S1" traces seemed to suggest a very large PPR value. I suggest vastly expanding the size of this figure and the associated results section.

      Thank you for your interest in Fig. 5. We have explained the normalization procedure in more detail in the revised version. Also, we have modified the axis label of Fig. 5B to improve clarity. Finally, we have added a paragraph to explain the reason why the RS/DS model accounts for the results of Fig. 5.

      (3) Given the similar results but differing mechanistic conclusions of Doussau et al., (2017), that study merits more discussion in the manuscript.

      See our response to this point above (main comment (3) by the same Reviewer).

      (4) Several studies of LFD synapses have found a role for kinase activity (Silverman-Gavrila et al., 2005; Doussau et al,. 2010). This mechanism could be discussed or experimentally tested here.

      As mentioned in the Discussion section, the nature of the proteins involved in LFD, or in docking/undocking, remains uncertain. It is not surprising that broad-spectrum kinases and phosphatases affect LFD, as reported in the papers quoted by the Reviewer, but the full implications of these findings will only become clear

    1. eLife Assessment

      This valuable study compares auditory cortex responses to sounds and cochlear implant stimulation measured with surface electrode grids in rats. Beyond the reduced frequency resolution of cochlear implants observed previously, this study suggests further discrepancies between neuronal representations of cochlear stimulations and natural sounds. The evidence for this result is solid. This study is of interest to researchers in the auditory neuroscience field and clinicians implementing treatments with cochlear implants.

    2. Reviewer #1 (Public review):

      Summary

      This manuscript addresses an important question in auditory neuroscience and neuroprosthetics: whether cortical responses to cochlear-implant stimulation resemble those evoked by natural acoustic stimulation, or whether electrical stimulation engages a distinct cortical population representation. The authors use high-density intracranial EEG recordings in rats to compare responses to pure tones in normal-hearing animals with responses to single-channel cochlear-implant stimulation in deafened animals. They combine analyses of event-related potentials, high-gamma activity, trial-by-trial variability, PCA/TCA-based dimensionality reduction, and decoder-based measures of stimulus information.

      Strengths

      A major strength of the study is the question it addresses. Understanding how electrical cochlear stimulation is represented centrally is highly relevant for cochlear-implant design, fitting strategies, rehabilitation, and broader theories of sensory neuroprosthetics. The comparison between acoustic and electrical stimulation, including within-animal comparisons in a subset of cases, is valuable because it directly asks whether implant-evoked cortical activity can be interpreted within the same framework as normal acoustic responses.

      The methodological approach is also a strength. Dense cortical surface recordings provide simultaneous access to spatial and temporal features of auditory cortical responses. The combination of PCA, TCA, and decoder analyses gives complementary views of the data, and the information-transfer analysis provides an interesting way to test whether representations learned from acoustic stimulation generalise to electrical stimulation.

      The revision has improved the manuscript considerably. The use of mixed-effects models better matches the partially paired experimental design. The expanded Methods improve reproducibility. The revised cohort descriptions and figure legends make the experimental design easier to follow. The clarification of Figure 8 strengthens the interpretation of the poor cross-modal transfer result, and the more cautious framing of spatial organisation better reflects the data.

      Weaknesses

      The main remaining limitation concerns experimental validation. The authors now clearly state that deafening was not verified with ABRs, hair-cell counts, or behavioural confirmation in the animals used for the main iEEG dataset, and that support for deafening efficacy in this cohort relies on prior validation of the same procedure in separate cohorts. This transparency is welcome and improves the manuscript, but the lack of direct validation in the main cohort remains a limitation, particularly given the importance of complete deafening and reliable cochlear implantation for interpreting implant-evoked cortical responses.

      Appraisal

      This does not undermine the main conclusions, but it should be kept in mind when interpreting the strength of the cochlear-implant comparisons. Overall, the revised manuscript is much stronger and more appropriately framed than the previous version. The study addresses an important problem, uses useful analytical approaches, and provides convincing evidence that acoustic and acute cochlear-implant stimulation evoke cortical responses with poor representational transfer.

      Likely impact of the work on the field

      The work is likely to be of interest to auditory neuroscientists, cochlear-implant researchers, and neuroengineers. Even where some conclusions require caution, the dataset and analytical framework may be useful for future studies aiming to relate central neural responses to implant programming, perceptual learning, or closed-loop neuroprosthetic strategies.

      Comments on revised version.

      The revised manuscript is substantially improved. The authors have clarified the methodological details, improved the statistical treatment of partially paired data, provided a clearer account of the animal cohorts, clarified the Figure 8 cross-modal decoding framework, and moderated several claims about spatial organisation and perceptual interpretation. Importantly, the manuscript now distinguishes more clearly between non-random spatial organisation, coarse cochleotopic structure, and sharply graded tonotopy or cochleotopy.

      The results support the conclusion that acoustic and cochlear-implant stimulation evoke cortical responses with different properties. In particular, acoustic responses support better single-trial stimulus decoding than cochlear-implant responses, and decoders trained on acoustic responses generalise poorly to implant-evoked responses. The evidence for spatial organisation is more nuanced: the cochlear-implant condition shows evidence of coarse, non-random spatial structure, but not strong evidence for consistent, sharply graded cochleotopy. Overall, the revised manuscript makes a valuable contribution, provided that the conclusions remain framed around acute cortical responses and poor representational transfer rather than definitive claims about long-term perceptual experience.

    3. Reviewer #2 (Public review):

      Summary:

      This article reports measurements of iEEG signals on the rat auditory cortex during cochlear implant or sound stimulation in separate groups of rats. The observations indicate some spatial organization of cochlear implant stimuli, but that is very different from cochlear implants.

      Strengths:

      The study includes some interesting analyses of the sound and cochlear implant representation structure based on decoders.

      Weaknesses:

      The observation that responses to cochlear implant stimulation (stimulation) is spatially organized but not exactly as sound-driven responses is not new.

      The analyses in Fig. 8 supporting the claim that there is a mismatch between cochlear implant and normal sound representations remain hard to evaluate. The shuffle control now provided by the authors indicates that the information transfers between normal hearing representations is at chance level and between cochlear implant and normal hearing is below chance (Fig 8H). This clearly indicates, unlike the authors suggest in their response, that the analysis used to make this claim is not sensitive enough. Therefore, the claim does not seem to be supported.

    4. Reviewer #3 (Public review):

      Summary:

      Through micro-electroencephalography, Hight and colleagues studied how the auditory cortex in its ensemble respond to cochlear implant stimulation compared to the classic pure tones. Taking advantage of a double implanted rat model (Micro-ECoG and Cochlear Implant), they tracked and analyzed changes happening in the temporal and spatial aspects of the cortical evoked responses in both normal hearing and cochlear-implanted animals. After establishing that single trial responses were sufficient to encode the stimuli properties, the authors then explored several decoder architectures to study the cortex ability to encode each stimuli modality in a similar or different manner. They conclude that a) intracranial EEG evoked responses can be accurately recorded and did not differed between normal hearing and cochlear-implanted rats; b) Although coarsely spatially organized, CI-evoked responses had higher trial-by-trial variability than pure tones; c) Stimulus identity is independently represented by temporal and spatial aspect of cortical representations and can be accurately decoded by various means from single trials; d) and that Pure tones trained decoder can't decode CI-stimulus identity accurately.

      Strength:

      The model combining micro-eCoG and cochlear implantation and the methodology to extract both the Event Related Potentials (ERPs) and High-Gammas (HGs) is well designed and appropriately analyzed. Likewise, the PCA-LDA and TCA-LDA are powerful tools that take full advantage of the information provided by the cortical ensembles.

      The overall structure of the paper, with a paced and exhaustive progress through each step and evolution of the decoder is very appreciable and easy to follow. The exploration of single trial encoding and stimulus identity through temporal and spatial domains is providing new avenues to characterize the cortical responses to CI stimulations and their central representation. The fact that single trials suffice to decode the stimulus identity regardless of their modality is of great interest and noteworthy. Although the authors confirm that iEEG remains difficult to transpose in clinic, the insights provided by the study confirm the potential benefit of using central decoders to help in clinic settings.

      Weakness:

      The conclusion of the paper, especially the concept of distinct cortical encoding for each modality, is unfortunately only partially supported by the results. Although acknowledged by the authors, fundamental limitations related to CI stimulation might have weaken the results.

      The authors stimulated in a Monopolar mode which, albeit being clinically relevant, notoriously generates a high current spread in rodent models. Thus, it seems possible that current spread ended stimulating indistinctly higher turns of the cochlea or even the modiolus in a non-specific manner, greatly reducing (or smearing) the place-coding/frequency resolution of each electrode, which in turn could explain the coarse topographic (or non-random) organization of the cortical responses.

      Although the authors acknowledge that post-lingual CI users always have an adaptation period, their conclusion is based on measurements that are relatively "early" in the CI-use timeline so to speak since iEEG were collected a) acutely right after mono-aural implantation and stimulation, b) under anesthesia, c) using unmodulated pulse train fixed at 900pps regardless of the electrode used and thus lacking any temporal information shifts in relationship to electrode cochleotopic placement. Basically, all CI electrodes had the same rate whereas you would expect basal CI electrodes to be amplitude modulated at higher frequencies than apical electrodes.

      Nevertheless, the reviewer wants to reiterate that the study proposed by Hight et al. is well constructed, relevant to the field and that the overall proposal of improving patient performances and help their adaptation in the first months of CI use by studying central responses should be pursued as it might help establish new guidelines or create new clinical tools.

      Comments on revised version.

      The reviewer would like to thank the Authors for their work on this new version. The reviewer is satisfied with the current state of manuscript and its associated public review and has no further comment.

    5. Author response:

      The following is the authors’ response to the previous reviews

      Summary of revision for all reviewers:

      We are encouraged that the reviewers recognized the importance of the central question addressed by this study, the value of comparing acoustic- and cochlear-implant-evoked cortical responses within a common framework, and the potential relevance of these analyses for auditory neuroscience and neuroprosthetic design.

      At the same time, the reviewers made clear that the manuscript would be strengthened by:

      (1) Clearer calibration of claims regarding spatial organization 

      (2) Clarified explanation of the Figure 8 cross-modal analyses

      (3) More explicit discussion of the methodological and interpretive limitations of the present dataset, particularly the acute, anesthetized, and partially indirectly validated nature of the experiments

      In response, we substantially revised the manuscript to improve methodological clarity, narrow claims where appropriate, and align the title, abstract, results, discussion, figures, and legends with the precision of the data.

      Summary of major changes to revised manuscript:

      We revised the manuscript throughout to distinguish non-random spatial organization, coarse topographic/cochleotopic structure, and locally graded tonotopy/cochleotopy, and we softened claims accordingly, especially for the cochlear implant and TCA-derived analyses.

      We substantially clarified the Figure 8 cross-modal decoding framework, including:

      - The within-modality normal-hearing control

      - The normal-hearing-trained / cochlear-implant-tested analysis

      - The shuffled baseline used to estimate chance-level information transfer

      We expanded the methods and discussion to make the scope and limitations more explicit, including:

      - The acute and anesthetized nature of recordings

      - The use of monopolar stimulation

      - The lack of direct deafening validation in the main iEEG cohort

      - That ECAP forward-masking measurements were obtained in a separate acute cohort

      - The possibility that the observed acoustic–electrical mismatch is transient rather than fixed

      We also improved the consistency of our terminology, panel labeling, figure legends, and cohort descriptions.

      Public Reviews:

      Reviewer #1 (Public review):

      We thank the reviewer for recognizing the importance of the question addressed by this study and for pushing us to sharpen the distinction between non-random spatial structure, coarse topography, and graded cochleotopy. These comments also helped us better calibrate our claims about TCA-derived maps and cross-modal generalization, and prompted clearer discussion of the limits of the present dataset.

      (1) The main weakness is that the evidence for spatial organization remains difficult to interpret. In Figure 2, the authors argue that both tone-evoked and cochlear implant-evoked responses are spatially organized, but the slope analyses are not significant for the cochlear implant condition. The revised vector-strength analysis supports the presence of non-random spatial structure, but this is not the same as demonstrating a clear graded cochleotopic organization. The manuscript would be strongest if it consistently distinguished between non-random spatial structure, coarse topography, and true graded tonotopy or cochleotopy.

      We agree. We revised the manuscript to distinguish these levels of interpretation more carefully and consistently. Specifically, we now reserve “non-random spatial organization” for results supported by the vector strength/shuffle analyses, and describe cochlear-implant-evoked maps as showing coarse and variable spatial structure rather than uniformly robust graded cochleotopy.

      We now also emphasize more explicitly that implanted animals varied considerably: some showed clear nonrandom organization, whereas others exhibited effectively random global maps. At the same time, the aggregate implant-evoked data remained non-random and showed decreasing spatial correlation with increasing electrode separation, which we interpret as evidence for coarse cochleotopic structure at the population level, rather than strong local graded cochleotopy in every animal.

      Accordingly, we revised the manuscript to distinguish:

      “Non-random spatial organization”

      “Coarse topographic/cochleotopic structure”

      “Locally graded tonotopy/cochleotopy”

      (2) A related issue is that some figure titles and interpretive statements still appear stronger than the data justify. For example, the TCA results in Figure 7 are described as revealing topographically organized latent spatial factors, but the statistical support appears strongest for normal-hearing high-gamma responses, with weaker or non-significant results in other conditions. These data remain interesting, but they would be better framed as evidence for weak or coarse spatial structure rather than robust topographic organization across all modalities.

      Now in our updated manuscript we revised the Figure 7 title, legend, and results to avoid implying robust topographic organization across all conditions. The manuscript now describes the TCA-derived spatial factors as showing coarse, non-random spatial structure, with the strongest support for local tonotopy in the normal hearing high-gamma condition and weaker or non-significant support for robust local topography in the other conditions.

      (3) The decoder analyses are improved, especially with the added tone-to-tone control. This control supports the conclusion that poor acoustic-to-CI transfer is not simply a failure of the TCA/LDA pipeline. However, the analysis remains model-dependent, and the absolute information transfer values are low. It would be helpful either to include an analogous analysis using raw ERP/high-gamma features or to explain more explicitly why the TCA-based approach is the appropriate primary test. The data support poor generalization between acoustic and implant-evoked cortical responses, but claims about perceptual qualities should remain speculative because perception is not directly measured in these experiments.

      Thanks, we now revised the Figure 8 results section to first explain that TCA is especially appropriate here because it constrains the decomposition into separable spatial, temporal, and trial factors. This makes it possible to learn spatial and temporal factors in the normal-hearing condition, hold those fixed, and re-optimize only trial factors in withheld normal-hearing data or cochlear-implant-evoked data, providing an interpretable test of within-modality recovery and cross-modal generalization.

      Thus, our emphasis on TCA is primarily conceptual, not merely computational. Raw-feature decoding could also be informative, but for the specific cross-modal question addressed here, TCA provides the more interpretable framework.

      We also now state more explicitly that the absolute information-transfer values are low, and we correspondingly limit our conclusion: normal-hearing-trained representations generalize poorly to acute cochlear-implant-evoked responses in this framework. We further and thoroughly revised the abstract, results, and discussion so that any perceptual implications remain explicitly speculative, since perception was not measured directly in these experiments.

      (4) Finally, although methodological reporting is much improved, some verification remains indirect. The authors provide useful implantation criteria and cite prior validation of their deafening approach, but the manuscript would be clearer if it explicitly distinguished between validation performed in the present animals and validation based on previous cohorts. This distinction is important because surgical variability, implantation efficacy, and deafening completeness can influence the interpretation of cochlear implant experiments.

      We revised the manuscript to make this distinction explicit. In the methods, we now state clearly that we did not obtain ABR measurements, hair-cell counts, or behavioral confirmation of deafening in the animals used for the present iEEG dataset, and that support for deafening efficacy in this cohort therefore relies on prior validation of the same procedure in separate cohorts.

      We also now clarify that the ECAP forward-masking measurements were obtained in a separate acute cohort (N=3) and were not recorded from the animals used in the main iEEG dataset.

      Our goal in these revisions was to make the provenance of each validation step fully explicit and to avoid implying that all validation measures were obtained in the same animals used for cortical recording.

      Reviewer #2 (Public review):

      We thank the reviewer for highlighting the value of the decoder-based analyses while also challenging us to frame the study’s novelty more precisely and to make Figure 8 substantially clearer. In response, we narrowed our novelty claims, improved the explanation of the cross-modal decoding framework, and clarified the logic of the shuffled baseline.

      (1) The observation that responses to cochlear implant stimulation (stimulation) is spatially organized is not new (e.g. Adenis et al. 2024)

      Thanks, good point. We revised the Introduction to make clear that prior studies have already demonstrated spatial organization of cochlear-implant-evoked cortical responses, including prior animal work cited in the manuscript. Our intended contribution is therefore not the demonstration of spatial organization per se, but rather the use of a shared analytical framework to test whether acoustic and electrical stimulation produce overlapping or transferable spatiotemporal cortical population representations, including single-trial decoding and cross-modal generalization analyses.

      (2) The claim that spatial and temporal dimensions contribute information about the sound is also not new there is a large literature on this topic.

      We revised the manuscript to avoid implying novelty for this general principle. Our intended point is narrower: in the present work, spatial and temporal dimensions were recorded simultaneously across a 60-channel cortical surface array and integrated into trial-by-trial population analyses that could be directly compared across normal-hearing and cochlear-implant conditions within a shared framework.

      (3) The analyses supporting the claim that there is a mismatch between cochlear implant and sound representation are still unclear, particularly in Fig. 8.

      We revised both the Figure 8 legend and the results section to explain the analysis step-by-step. In particular, we now distinguish clearly among:

      The normal-hearing to normal-hearing re-optimization control, which tests whether the TCA/LDA framework can recover stimulus-related structure when modality is unchanged;

      The normal-hearing-trained / cochlear-implant-tested analysis, which tests cross-modal generalization; the shuffled baseline, which estimates chance-level information transfer by destroying any structured relationship between predicted labels and actual implant channels while preserving matrix dimensions.

      We now also explain explicitly why the shuffled baseline can equal or slightly exceed the measured normal hearing-to-cochlear-implant transfer. Because observed cross-modal transfer was extremely small, shuffling does not restore meaningful structure; rather, it shows that the measured transfer lies at or below chance level. We believe that this substantially improves the clarity and interpretability of Figure 8.

      Reviewer #3 (Public review):

      We thank the reviewer for recognizing the strengths of the study design, the single-trial decoding analyses, and the potential clinical relevance of this approach, while also pressing us to better address the limitations of monopolar stimulation, possible place-frequency mismatch, and the acute nature of the recordings. These comments prompted important revisions to both interpretation and discussion.

      (1a) The conclusion of the paper, especially the concept of distinct cortical encoding for each modality, is unfortunately partially supported by the results as the authors ignored fundamental limitations of CI related stimulation. First, the authors stimulated in a Monopolar mode which, albeit being clinically relevant, notoriously generates a high current spread in rodent models.

      We agree that monopolar stimulation is an important limitation, especially in the rodent cochlea, where current spread may be broader than in human subjects. We now acknowledge this more explicitly in the discussion.

      At the same time, our ECAP forward-masking measurements in a separate acute cohort provided supportive evidence for spatially and temporally tuned peripheral activation under the stimulus intensities used here. We therefore interpret the data as indicating that some peripheral selectivity was present, while also acknowledging that broader current spread under monopolar stimulation may have contributed to the coarse cortical organization observed in the implant condition.

      We also now state explicitly that the observed coarse cortical organization likely reflects a combination of peripheral current spread, downstream cortical pooling, and the spatial resolution limits of mesoscale surface iEEG, rather than any single factor alone.

      (1b) Comparing the averaged BF maps for iEEG (Fig-2A, C), BFs ranged from 4 to 16kHz with a predominance of 4kHz BFs. The lack of BFs at higher frequencies might reveal a potential location mismatch between the frequency range sampled at the level of the cortex (low to medium frequencies) and the frequency range covered by the CI inserted mostly in the first turn-and-a-half of the cochlea (high to medium frequencies). Looking at Fig2F (and to some extend 2A) most of CI electrodes elicited responses around the 4kHz regions and averaged maps show a predominance of CI-3-4 across cortex (Fig-2C, H and Sup Fig. 3) from areas with 4kHz BF to areas with 16kHz BF. It is doubtful that CI-3-4 are located near the 4kHz region based on Müller's work (1991) on the frequency representation in the rat cochlea.

      We appreciate this point and agree that a precise one-to-one alignment between cortical best-frequency maps and intracochlear electrode position cannot be established from the present data. We now clarify this limitation in the manuscript, noting that the iEEG-based maps are relatively coarse and appear to overrepresent mid-frequency regions, limiting direct inference from implant electrode number to a precise acoustic-frequency equivalent.

      We also emphasize that our central conclusion does not depend on assigning each implant electrode a specific best frequency, but rather on comparing the spatial organization and cross-modal generalization of cortical population responses.

      (1c) Moreover, Supplemental figure 3 shows that only a couple of CI electrodes are predominately represented at the level of the cortex. Thus, it seems possible that current spread ended stimulating indistinctly higher turns of the cochlea or even the modiolus in a non-specific manner, greatly reducing (or smearing) the placecoding/frequency resolution of each electrode, which in turn could explain the coarse topographic (or coarsely tonotopic according to the manuscript) organization of the cortical responses.

      We agree that the predominance of only a subset of implant electrodes in the cortical maps is consistent with relatively coarse peripheral and/or cortical place coding. We revised the discussion to acknowledge more explicitly that possible current spread, broad peripheral excitation, and the limited spatial resolution of surface iEEG could all contribute to the coarse spatial organization observed in the implant condition.

      We also emphasize in the revised text that this possibility does not negate the evidence for non-random organization, but it does limit the strength of any claim about finely graded cochleotopy.

      (2) Second, although the authors acknowledge that post-lingual CI users always have an adaptation period, their conclusion is based on measurements that are relatively "early" in the CI-use timeline so to speak since iEEG were collected a) acutely right after mono-aural implantation and stimulation, b) under anesthesia, c) using unmodulated pulse train fixed at 900pps regardless of the electrode used and thus lacking any temporal information shifts in relationship to electrode cochleotopic placement. Basically, all CI electrodes had the same rate whereas you would expect basal CI electrodes to be amplitude modulated at higher frequencies than apical electrodes.

      We agree. We now emphasize more clearly that the present experiments probe only an early and simplified stage of cochlear implant use. The recordings were acute, performed under anesthesia, and used short constant-rate pulse trains chosen to isolate fundamental organizing features of primary cortical responses rather than to reproduce the full complexity of clinical stimulation.

      We therefore frame our conclusions specifically in terms of acute cortical encoding under these conditions, and now state explicitly that these data do not address how chronic use, wakefulness, adaptation, or more complex modulation-rich stimuli may alter more global cortical representations.

      (3) As much as the reviewer likes the overall approach with the use of PCA-LDA and TCA, and agrees that information transfer seems inexistant at time of measurement, authors should be more careful in their strong conclusion that two distinct encoding exist. The non-overlapping between sound and electric stimulation representations might exist only transiently and this should be acknowledged a bit more in the discussion. Without repetition of iEEG measurement at later period with chronic use of the CI, it is not possible to definitively claim that two distinct, non-overlapping coding co-exist at all times.

      We agree and appreciate this point. We revised the Discussion to make this much more explicit. Our data support poor overlap or poor cross-modal generalization at the time of acute implant activation, but they do not establish that this relationship is fixed over chronic implant use.

      We now state directly that the observed mismatch may be transient, and that future longitudinal studies will be required to determine whether cortical representations become more aligned with experience, whether downstream readout adapts to a novel code, or whether both processes contribute.

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      (1) The authors have provided new analyses to support the claim that CI and NH representations do not match. However, the analyses performed are presented in an unclear manner. Most particularly, it is very hard to understand what has been done in Fig. 8D and Fig. 8G. The legend in D is incomplete. The legend for G does not explain why there are lines with different color and what are the different lines. The reasoning behind the shuffled CI dataset is not explained. Why does shuffling restore information transfer, this is very counter intuitive and raises again the question of the value of this analysis.

      Agreed, we have revised the Figure 8 legend and rewritten the Figure 8 results section to initially explain the analysis workflow more clearly.

      We now explain that the shuffled condition is used to estimate a chance-level baseline for information transfer in the normal-hearing-trained / cochlear-implant-tested confusion matrix. Specifically, we compute mutual information after randomizing the predicted labels in the confusion matrix, thereby destroying any structured relationship between predicted tone labels and actual implant channels while preserving matrix dimensions and marginal structure (results, Fig. 8 section).

      We also now state explicitly that while this shuffled condition can yield slightly higher mutual information than the non-shuffled NH→CI condition, the point is not that shuffling restores meaningful decoding. Rather, the point is that the observed NH→CI transfer is so low that it does not exceed this chance-level baseline (results, Fig. 8 section).

      (2) Note that the legends of fig. 8 are mislabel (goes up to H although the last panel is G, a legend for E is missing).

      Thank you for catching this error. We have corrected the Figure 8 legend so that all panels are accurately labeled and described.

      Reviewer #3 (Recommendations for the authors):

      (1) Fig. 2C and 2H are from 2 to 16kHz (as announced in the previous responses) but then Sup. Fig. 3 goes from 2 to 32kHz. Still on Sup. Fig. 3, color code for CI electrodes is inverted compared to the rest of the MS.

      Thanks, good catches. We have updated Figures 2C and 2H to represent the full tested frequency range, and we have corrected the inverted color code for cochlear-implant electrodes in Supplemental Figure 3.

      (2) Fig. 2A. says n=1, Fig. 2F should say the same. Fig. 2C and 2H should also have n=1 on bottom map. Fig. 2D and 2I should mention n=1. Same thing with Fig. 3A/3C, Fig. 3B/3D (n=7), Fig. 7A/7B/7C.

      We have revised the relevant figure legends to make clear that data are from a single animal unless otherwise noted, while minimizing visual crowding in the figure panels.

      (3) The reviewer once again thinks it would be better to not truncate the axis of Figs. 4C, 6C, 7D as it creates confusion with Figs. 5D and 8G, where suddenly, there are 8 electrodes. The legend justification isn't enough.

      We appreciate this concern and have clarified the issue in the revised manuscript. In some animals (N=3), some electrodes in the 8-channel array were non-functional prior to implantation, so those animals contributed six rather than eight usable implant channels. We now make this explicit in the manuscript and supplemental figures, including the number of channels used in the respective animals in the methods under “Cochlear implant programming” and in Supplemental Figure 3. 

      (4) Although the reviewer appreciated the justification of 15 PCA for their analysis, this justification should be provided as is in the methods, especially as the Sup. Fig. 4 does not even explain the significance of the linear regression of components 16 to 30. Some other reviewers would say that 10 components were more than enough without this justification.

      We agree and have provided further justification of the 15 PCA components to provide justification as is in the Methods “Principal component analysis” and in the results describing Figure 4. 

      (5) Sup. Fig. 1 and the eCAP measurements are a nice and needed addition to the MS but the authors should make it clearer that it came from 3 animals that are not part of the cohort presented in the rest of the paper / in Sup. Fig. 2. The method section certainly not make that distinction, giving the impression that all animals have been tested for channel interactions and temporal recovery.

      We agree and have revised the methods section to state clearly that the ECAP forward-masking measurements were obtained in a separate group of acutely implanted rats (N=3), rather than the main iEEG cohort in Supplemental Figure 1.

      (6) As said in the public review, the authors should acknowledge more the potential transiency of the nonoverlapping representation in the discussion, especially since most of the recordings are acute.

      We agree, and we have revised the discussion accordingly. As described in our public response above, we now state directly that the poor overlap between acoustic and electrical representations was observed under acute recording conditions and may not persist unchanged with chronic implant use or behavioral adaptation.

    1. eLife Assessment

      These are valuable findings for those interested in how neural signals reflect auditory speech streams, and in understanding the roles of prediction, attention, and eye movements in this tracking. The evidence is solid, as the revised analyses broadly support the main claims and address the key methodological concerns raised during review. Remaining limitations concern the interpretation of some sensor-level and eye-movement effects, but these do not substantially undermine the overall conclusions.

    2. Reviewer #1 (Public review):

      (Editors' note: this version has been assessed by the Reviewing Editor without further input from the original reviewers. The authors have addressed the comments raised in the previous round of review.)

      Summary:

      This study aimed at replicating two previous findings that showed (1) a link between prediction tendencies and neural speech tracking, and (2) that eye movements track speech. The main findings were replicated which supports the robustness of these results. The authors also investigated interactions between prediction tendencies and ocular speech tracking, but the data did not reveal clear relationships. The authors propose a framework that integrates the findings of the study and proposes how eye movements and prediction tendencies shape perception.

      Strengths:

      This is a well-written paper that addresses interesting research questions, bringing together two subfields that are usually studied in separation: auditory speech and eye movements. The authors aimed at replicating findings from two of their previous studies, which was overall successful and speaks for the robustness of the findings. The overall approach is convincing, methods and analyses appear to be thorough, and results are compelling.

      Weaknesses:

      Eye movement behavior could have presented in more detail and the authors could have attempted to understand whether there is a particular component in eye movement behavior (e.g., blinks, microsaccades) that drives the observed effects.

    3. Reviewer #2 (Public review):

      Summary

      Schubert et al. recorded MEG and eye tracking activity while participants were listening to stories in single-speaker or multi-speaker speech. In a separate task, MEG was recorded while the same participants were listening to four types of pure tones in either structured (75% predictable) or random (25%) sequences. The MEG data from this task was used to quantify individual 'prediction tendency': the amount by which the neural signal is modulated by whether or not a repeated tone was (un)predictable, given the context. In a replication of earlier work, this prediction tendency was found to correlate with 'neural speech tracking' during the main task. Neural speech tracking is quantified as the multivariate relationship between MEG activity and speech amplitude envelope. Prediction tendency did not correlate with 'ocular speech tracking' during the main task. Neural speech tracking was further modulated by local semantic violations in the speech material and by whether or not a distracting speaker was present. The authors suggest that part of the neural speech tracking is mediated by ocular speech tracking. Story comprehension was negatively related with ocular speech tracking.

      Strengths

      This is an ambitious study, and the authors' attempt to integrate the many reported findings related to prediction and attention in one framework is laudable. The data acquisition and analyses appear to be done with great attention to methodological detail. Furthermore, the experimental paradigm used is more naturalistic than was previously done in similar setups (i.e.: stories instead of sentences).

      Weaknesses

      While the analysis pipeline is outlined in much detail, some analysis choices appear ad-hoc and could have been more uniform and/or better motivated (other than this is what was done before).

    4. Reviewer #3 (Public review):

      I thank the authors for their extensive revision of this paper, and I found some elements greatly improved.

      In particular, the authors do embrace a somewhat more speculative tone in the current version, which I think is fitting for this work, as the data seem (to me) to be not fully conclusive. The data set collected here is clearly valuable and unique (and I would encourage the authors to make it publicly available!), however, my overall impression is that the specific analyses reported here might not fully.

      Despite the revised description of methods, results and figures, I still have trouble understanding many of the results and the authors conclusive interpretation of them. These are my main reservations:

      (1) Regarding "individual prediction tendency" - thank you for adding clarifying methodological details and showing the data in a new Figure (#2). Honestly, however, I still can't say that I fully understand the result. For example, why is there also a significant response in the random condition as well? And how do you interpret the interesting time-course (with a peak ~200ms prior to the stimulus, and a reduction overtime from there?<br /> Also (I may have missed this, but...) what neural data was used to train the classifier and derive the "prediction tendency" index? Was it just the broadband neural response? Is there a way to know which sensors contributed to this metric (e.g., are they predominantly auditory? Frontal?)? And is there a way to establish the statistical significance of this metric (e.g., how good the decoder actually was in predicting behavioral sensitivity?). I don't see any statistics in the results section describing the individual prediction tendency.

      (2) Regarding the TRF analysis - Thanks for clarifying the approach used to obtain 2-second long "segments" of speech tracking. This is an interesting approach, however I think quite new(?) , and for me it raises a whole new set of questions, as well as additional controls and data that I would have liked to see, to be convinced that results are significant. I will elaborate:

      - Do I understand correctly that you segment the real and predicted neural response into 2-second-long segments and then calculate the Pearsons' correlation between them to assess the goodness of the model? This is very unclear, since in the methods section you state only that "the same" analysis was performed as for the full data - but what exactly? Clearly, values will be very different when using such short segments. I feel that additional details are still required (and perhaps data shown) to fully understand the "semantic violation" analysis of TRFs.

      - I would like to reiterate my previous comment regarding the use of permutation tests to verify the validity of TRF-based measures derived. This would be especially important when using new approaches (such as the segmentation used here). The authors argue that this is not needed since this was not done in their previously published study. However, this sounds a bit like "two wrongs make a right" argument... why not just do it, and let us know that this 2-second segmentation approach allows estimating reliable speech tracking?

      - Following up on my previous comment that defining "clusters" as at least two neighboring channels (Figure 3) - the fact that this is a default in Fieldtrip is by no means sufficient justification! This seems quite liberal to me, especially given the many comparisons performed. Here too, permutations can help to determine the necessary data-driven threshold for corrections. This is of course critical for interpreting the result shown in Figures 3E&G that are critical "take home messages" of the paper - i.e., that the prediction-index from the first part of the experiment is related to speech tracking in the second part of the experiment. To my eyes, this does not look extremely convincing, but perhaps the authors can show more conclusive data to support this (e.g., scatter plots of the betas across participant?).<br /> - A similar point can be made for the effect of semantic violations (though here the scalp-level result is somewhat more clustered). The authors point out that the semantic effect is a "replication" of their result reported in Schubert et al. 2023, but if I am not mistaken the results there were somewhat different (as was the manipulation). It would be nice to explicitly discuss the similarity/difference between these effects.

      (3) Regarding the ocular-TRFs -

      - Maybe this is just me, but I believe that effects that are robust should be clearly visible in the data, without the need for fancy "black-box" statistical models. In the case of the ocular TRFs, it is hard for me to see how these time-courses are not just noise (and, again, a permutation test would have helped to convince me...). The inconsistent results for horizontal and vertical eye-movements vis a vis the experimental conditions (single vs. multi-speaker conditions) don't help either, despite the authors argument that these are "independent" - but why should this be the case, especially if there is nothing really to look at in this task?<br /> - I remain with this scepticism for the mediation-portion of the analysis as well... But perhaps replications from other groups or making the data public will help shed further light on this in the future.

      Minor<br /> - Thanks for adding information about the creation of semantic-violation stimuli. Since the violations and lexical-controls were taken from different audio recordings, it would have been nice to verify that differences between neural responses cannot be attributed to differences in articulations (e.g., by comparing their spectro-temporal properties).

    5. Author response:

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This study aimed at replicating two previous findings that showed (1) a link between prediction tendencies and neural speech tracking, and (2) that eye movements track speech. The main findings were replicated which supports the robustness of these results. The authors also investigated interactions between prediction tendencies and ocular speech tracking, but the data did not reveal clear relationships. The authors propose a framework that integrates the findings of the study and proposes how eye movements and prediction tendencies shape perception.

      Strengths:

      This is a well-written paper that addresses interesting research questions, bringing together two subfields that are usually studied in separation: auditory speech and eye movements. The authors aimed at replicating findings from two of their previous studies, which was overall successful and speaks for the robustness of the findings. The overall approach is convincing, methods and analyses appear to be thorough, and results are compelling.

      Weaknesses:

      Eye movement behavior could have presented in more detail and the authors could have attempted to understand whether there is a particular component in eye movement behavior (e.g., blinks, microsaccades) that drives the observed effects.

      Reviewer #2 (Public review):

      Summary

      Schubert et al. recorded MEG and eye tracking activity while participants were listening to stories in single-speaker or multi-speaker speech. In a separate task, MEG was recorded while the same participants were listening to four types of pure tones in either structured (75% predictable) or random (25%) sequences. The MEG data from this task was used to quantify individual 'prediction tendency': the amount by which the neural signal is modulated by whether or not a repeated tone was (un)predictable, given the context. In a replication of earlier work, this prediction tendency was found to correlate with 'neural speech tracking' during the main task. Neural speech tracking is quantified as the multivariate relationship between MEG activity and speech amplitude envelope. Prediction tendency did not correlate with 'ocular speech tracking' during the main task. Neural speech tracking was further modulated by local semantic violations in the speech material and by whether or not a distracting speaker was present. The authors suggest that part of the neural speech tracking is mediated by ocular speech tracking. Story comprehension was negatively related with ocular speech tracking.

      Strengths

      This is an ambitious study, and the authors' attempt to integrate the many reported findings related to prediction and attention in one framework is laudable. The data acquisition and analyses appear to be done with great attention to methodological detail. Furthermore, the experimental paradigm used is more naturalistic than was previously done in similar setups (i.e.: stories instead of sentences).

      Weaknesses

      While the analysis pipeline is outlined in much detail, some analysis choices appear ad-hoc and could have been more uniform and/or better motivated (other than: this is what was done before).

      Reviewer #3 (Public review):

      I thank the authors for their extensive revision of this paper, and I found some elements greatly improved.

      In particular, the authors do embrace a somewhat more speculative tone in the current version, which I think is fitting for this work, as the data seem (to me) to be not fully conclusive. The data set collected here is clearly valuable and unique (and I would encourage the authors to make it publicly available!), however, my overall impression is that the specific analyses reported here might not fully

      Despite the revised description of methods, results and figures, I still have trouble understanding many of the results and the authors conclusive interpretation of them. These are my main reservations:

      (1) Regarding "individual prediction tendency" - thank you for adding clarifying methodological details and showing the data in a new Figure (#2). Honestly, however, I still can't say that I fully understand the result. For example, why is there also a significant response in the random condition as well? And how do you interpret the interesting time-course (with a peak ~200ms prior to the stimulus, and a reduction overtime from there? Also (I may have missed this, but..) what neural data was used to train the classifier and derive the "prediction tendency" index? Was it just the broadband neural response? Is there a way to know which sensors contributed to this metric (e.g., are they predominantly auditory? Frontal?)? And is there a way to establish the statistical significance of this metric (e.g., how good the decoder actually was in predicting behavioral sensitivity?). I don't see any statistics in the results section describing the individual prediction tendency.

      (2) Regarding the TRF analysis - Thanks for clarifying the approach used to obtain 2-second long "segments" of speech tracking. This is an interesting approach, however I think quite new(?), and for me it raises a whole new set of questions, as well as additional controls and data that I would have liked to see, to be convinced that results are significant. I will elaborate:

      - Do I understand correctly that you segment the real and predicted neural response into 2-second long segments and then calculate the Pearsons' correlation between them to assess the goodness of the model? This is very unclear, since in the methods section you state only that "the same" analysis was performed as for the full data - but what exactly? Clearly, values will be very different when using such short segments. I feel that additional details are still required (and perhaps data shown) to fully understand the "semantic violation" analysis of TRFs.

      - I would like to reiterate my previous comment regarding the use of permutation tests to verify the validity of TRF-based measures derived. This would be especially important when using new approaches (such as the segmentation used here). The authors argue that this is not needed since this was not done in their previously published study. However, this sounds a bit like "two wrongs make a right" argument... why not just do it, and let us know that this 2-second segmentation approach allows estimating reliable speech tracking?

      - Following up on my previous comment that defining "clusters" as at least two neighboring channels (Figure 3) - the fact that this is a default in Fieldtrip is by no means sufficient justification! This seems quite liberal to me, especially given the many comparisons performed. Here too, permutations can help to determine the necessary data-driven threshold for corrections. This is of course critical for interpreting the result shown in Figures 3E&G that are critical "take home messages" of the paper - i.e., that the prediction-index from the first part of the experiment is related to speech tracking in the second part of the experiment. To my eyes, this does not look extremely convincing, but perhaps the authors can show more conclusive data to support this (e.g., scatter plots of the betas across participant?). - A similar point can be made for the effect of semantic violations (though here the scalp-level result is somewhat more clustered). The authors point out that the semantic effect is a "replication" of their result reported in Schubert et al. 2023, but if I am not mistaken the results there were somewhat different (as was the manipulation). It would be nice to explicitly discuss the similarity/difference between these effects.

      (3) Regarding the ocular-TRFs -

      - Maybe this is just me, but I believe that effects that are robust should be clearly visible in the data, without the need for fancy "black-box" statistical models. In the case of the ocular TRFs, it is hard for me to see how these time-courses are not just noise (and, again, a permutation test would have helped to convince me.). The inconsistent results for horizontal and vertical eye-movements vis a vis the experimental conditions (single vs. multi-speaker conditions) don't help either, despite the authors argument that these are "independent" - but why should this be the case, especially if there is nothing really to look at in this task? - I remain with this scepticism for the mediation-portion of the analysis as well... But perhaps replications from other groups or making the data public will help shed further light on this in the future.

      Minor

      - Thanks for adding information about the creation of semantic-violation stimuli. Since the violations and lexical-controls were taken from different audio recordings, it would have been nice to verify that differences between neural responses cannot be attributed to differences in articulations (e.g., by comparing their spectro-temporal properties)

      We are grateful to the Reviewing Editor and the three reviewers for the care and time they have devoted across the review rounds; their input has strengthened the paper. We are happy to confirm that we have carried out the two additional analyses the Assessment identifies as the path to a "solid" rating for methodological rigor. In brief:

      (1) Multiple-comparison correction for the prediction-tendency effect (Figure 3). We would like to be transparent about why our implementation differs from the specific procedure that was suggested. Testing the relationship between prediction tendency and neural speech tracking properly requires a mixed-effects model: the nested, repeated-measures structure of the data demands a per-subject random intercept (1|subject) to absorb the between-subject variability that the fixed effects do not capture. This is a standard and widely recommended approach for nested data, not an exotic one. Crucially, to our knowledge no implementation of a cluster-based permutation test exists for mixed-effects models — so the permutation/cluster procedure that was recommended cannot be applied to the very model the data structure requires. This constraint, rather than preference, is what led us to Bayesian estimation in the first place. We would add that the effect replicates Schubert et al. (2023), with a consistent left-frontal localisation across studies, which further speaks to its robustness.

      We fully share the concern underlying the recommendation, however — that a Bayesian model must equally guard against multiple comparisons across channels — and we have addressed it directly within the framework the data require. We tightened the HDI inclusion criterion to be analogous to a frequentist multiple-comparison correction (with boundaries comparable to a Bonferroni-corrected p-value) and re-estimated the model (encoding ~ prediction tendency * condition + (1|subject)) with 200,000 draws per chain (previously 2,000) to obtain stable posterior tails. The spatial pattern reproduces Figure 3E. At a Bonferroni-comparable per-channel boundary — itself widely regarded as overly conservative, just as the conventional 5% threshold is often criticised as arbitrary — no single channel survives in isolation; but considering the three-channel cluster jointly, the accumulated posterior places only ~0.7% of its mass at or below zero, i.e. a ~99.3% posterior probability of a positive effect (see Author response image 1). Rather than commit to a single arbitrary boundary, we report this accumulated posterior and invite the reviewers and readers to apply whatever criterion they consider appropriate.

      Author response image 1.

      (2) Circular-shift procedure for the mediation analysis (eye movements). Following Reviewer 1's suggestion, we replaced the previous shuffling control with a circularly shifted eye-movement predictor (shifted by half its total duration) as the control model. This isolates the genuine envelope contribution from any reduction in envelope weights that arises simply from adding a second predictor, and provides a stricter test of the specific temporal relationship between eye movements and the speech envelope. The mediation results are robust under this procedure (now Figure 5): vertical eye movements significantly mediate neural tracking of clear speech across all three principal components, and horizontal eye movements mediate tracking of target speech in the multi-speaker condition, with an early, sustained component peaking at ~0.18 s. Importantly, this stricter control did not materially change the interpretations we draw from these results: the pattern of mediation, and the conclusions about ocular speech tracking, remain consistent with those obtained under the previous approach.

      We have chosen to focus this revision specifically on these two points, which the Assessment identifies as what is needed to move the work from "incomplete" to "solid." We are sincerely grateful for the many further suggestions the reviewers raised, several of which we agree are thoughtful and point to worthwhile directions for future work. Since the previous round, however, the circumstances of the two authors who led this work have changed substantially: the corresponding author has since left academia altogether, and the shared first author's professional circumstances have likewise changed considerably. We are therefore not in a position to take on further analyses, or to prepare a point-by-point reply to the remaining comments, beyond the two above; we hope the editors and reviewers will understand that this necessarily defines the final scope of our revision. We remain sincerely grateful for the time and care they have invested in the manuscript.

    1. eLife Assessment

      This study presents a useful alternate labeling-based approach to identify or profile bioenergetic changes between cells. Overall, the approaches and data are solid, but have limitations in interpretation and so are as yet incomplete. The method itself can be of interest to a variety of biochemists, cell and systems-biologists interested in quantitative cellular metabolism.

    2. Reviewer #1 (Public review):

      The idea behind this paper is to have an alternate, reliable and quantitative approach to assess cell-cell metabolic heterogeneity. This study tries to achieve that using translationally-coupled energetic responses to metabolic stress. This is interesting because, in general, most quantitative measurements of metabolic outputs are 'bulk' and average for many cells. To overcome this, many recent studies use some read-outs of translation (presuming that translation is the single major energy sink in cells - however, this is objectively correct only in rapidly proliferating cells). That said, the authors take an interesting approach - to use two clickable methionine analogs, and assess baseline vs metabolically coupled translation within the same cell.

      The highlight is the methodology development where two distinct, clickable CMAs are used (to replace methionine in proteins). The labelling and approach are clever, and can be useful if carefully used. But there is going to be a challenge in using this, since the depletion of methionine itself (required for labelling), and a bias towards incorporation in some proteins (because these reagents are not highly permeable) will make it challenging to obtain precise metabolic state information - which otherwise can be obtained directly and far more precisely using a combination of other methods (ATP/flux measurement, respiratory capacity, translation rates etc). I therefore only make broader comments in my review below - to help structure this study better, clearly identify key limitations (and there are several that are clearly seen), and better clarify what MARBL may be useful for.

      (1) These reagents used for MARBL are not highly cell permeable/transported into cells, and largely work on the surface proteome. Which means the measurements related to changes in translation are indirect - quantified based on changes on the surface proteome (seen with labelling), and not the entire proteome.

      (2) A major limitation - which will confound any interpretation- is the need to use methionine-free media; this can be a problem beyond protein synthesis/met incorporation since this will almost instantly deplete SAM pools in cells. There is little data provided on the impact of using this approach on SAM pools (time kinetics, how quickly SAM pools are affected, how much of the impact on metabolism comes from purely that, etc).

      This is important to establish because (i) of the continuous, very high flux of SAM -> SAH (eg. in the folate pathway, other methylations), and a constant need for SAM synthesis from methionine. This information will set the limits of capabilities of this method, as well as help delineate how much you can interpret results related to metabolic states between compared cells/states, etc. The labelling process is ~2 hours, while effects on SAM can be seen within minutes of methionine starvation in media, in metabolically active cells.

      (3) What is the effect on overall adenylate charge/ATP due to shifting to methionine-free media + label addition? How does it vary between the cells tested (suspension vs adherent)? This should be established before data related to Fig. 2. Does it correlate with extent of AHA incorporation?

      The primary conclusion that this method suitably reflects overall changes in energetics comes from the titration of 2DG/glycolytic inhibition.

      (4) Relatedly, if this label incorporation experiment is carried out (for ~2 hrs), and subsequently there is a washout/replacement with fresh, methionine-supplemented medium, (how quickly) do the cells recover and restore their energetic allocations?

      (5) One possible advantage of a system like this can be to address questions in single cells/study cell metabolic heterogeneity. However, these are best done if the attaching moiety has a (selective) fluorescence increase and/or other read-out that can be quantitatively obtained at a single cell level. Largely, using AHA or HPG effectively only leads to bulk estimates (which can be sub-sorted towards single-cell estimates indirectly). This means that this method cannot really be used to study cell-cell metabolic heterogeneity effectively - compared to far simpler approaches, for example using a mitochondrial potentiometric dye with high fluorescence, or reporters for glycolytic activity, etc. This would also be related to Figure 5 - at best, this approach may complement existing approaches towards identifying heterogeneous sub-populations of cells.

      However, I do agree that MARBL is flexible, stable, and can be internally normalised and used through flow-based platforms. It can supplement existing approaches to perturb bioenergetics, and also supplement existing approaches to understand metabolic state in live cells, particularly in suspension cells.

    3. Reviewer #2 (Public review):

      Summary:

      Delacruz et al. describe a new method, called "MARBL" (Methionine Analogues for Ratiometric Bioenergetics in Live cells) to measure metabolic activity in single cells. The concept is similar to the SCENITH (anti-puromycin flow cytometry) assay to measure energy metabolism by measuring protein translation activity, yet offers, in theory, two advantages: 1) it keeps cells alive for downstream biological assays and 2) it is a ratiometric measurement, measuring both baseline translation and translation in the presence of metabolic inhibitors to correct for inherent cell-to-cell translation differences.

      Specifically, this method takes advantage of two click-chemistry-active methionine analogs, and then clicks fluorophores onto newly-synthesized surface proteins that have incorporated these analogs to measure translational activity. One methionine analog is given to cells for 2-4 hours to measure baseline translational activity, then metabolism is blocked using 2-deoxyglucose and oligomycin and the second methionine analog given to measure "metabolically-linked" translation activity. The authors establish this technique and show that mouse T cells polarized as pathogenic Th17 cells are more translationally active ("resilient") compared to non-pathogenic Th17 cells, and when sorted, the resilient cells produce more interferon-gamma. This latter finding requires live cells after the metabolic measurement assay, showcasing findings that are inaccessible to the SCENITH assay.

      Strengths:

      The approach used is conceptually clever. It is appealing to measure metabolic/translational activity and to then be able to carry out further assays on sorted cell populations with different degrees of metabolic activity. This would indeed represent a useful advance.

      Weaknesses:

      In principle, one key benefit of this technique is that cells can be used for biological assays after the metabolic measurement. Indeed, this would represent a valuable tool in the field.

      However, in this technique, cells are subjected to methionine deprivation, addition of non-natural methionine analogs, click chemistry, and high doses of toxic metabolic inhibitors 2-deoxyglucose and oligomycin. Indeed, the authors show in Figure S5F that 1/3 more of the post-MARBL cells die relative to cells not subject to this technique (60% viability in unclicked control, 40% in MARBL-measured cells). This data suggests that cells after this technique may be stressed and not reflective of the biological function of unmanipulated cells. More controls on viability and cell function (e.g. cytokine production) at more time points after the MARBL assay would have been valuable to address this issue.

      Another weakness of the paper is limited benchmarking against established metabolic assays in the field. The main assays used currently in the field are SCENITH and Seahorse. The authors do not compare their findings to SCENITH. They do compare their results to Seahorse, but the data shown don't address the key question: how does energy production measured by Seahorse, say in unmanipulated vs 2dg+oligomycin-treated cells, compare to the MARBL measurement? (Instead, they show a calculated "glucose dependence" metric in cells subjected to low vs high inhibitor dose, not showing the underlying data or cells that didn't receive an inhibitor).

    4. Author response:

      We are grateful to the reviewers and eLife editors for providing thoughtful and constructive commentary on our manuscript. The major concerns raised during review center on (A) the broader impact of methionine depletion on cellular metabolism beyond protein synthesis, including potential effects on cellular bioenergetics, and B) the extent to which assay conditions may induce cellular stress that influences downstream functional measurements. We will address both points during the revision period through experiments we have outlined below, most of which are already underway. 

      Related to (A), we will define the metabolic consequences of methionine deprivation on cellular metabolism with greater precision and are currently optimizing a third non-canonical alkyne amino acid, β-ethynylserine (β-ES), for labeling surface-exposed proteins in living cells. β-ES is a clickable threonine analogue that is efficiently incorporated into the proteome in the presence of physiological threonine and therefore does not require metabolic deprivation (1). We are in the process of optimizing β-ES for the MARBL workflow as a substitute for homopropargylglycine (HPG), which reduces the duration of methionine deprivation from six hours to two hours. Thus, β-ES eliminates the SAM-depletion concern for the baseline-translation half of the MARBL ratiometric measurement and constitutes a genuine methodological advance beyond the original submission. 

      Related to (B), we plan to generate additional data benchmarking MARBL against Seahorse and other established assays for measuring cellular bioenergetics, while also performing phenotypic and functional cellular characterization at key stages of the workflow. Experiments that are planned during the revision period are described in our point-by-point response below.

      Public Reviews:

      Reviewer #1 (Public review):

      The idea behind this paper is to have an alternate, reliable and quantitative approach to assess cell-cell metabolic heterogeneity. This study tries to achieve that using translationally-coupled energetic responses to metabolic stress. This is interesting because, in general, most quantitative measurements of metabolic outputs are 'bulk' and average for many cells. To overcome this, many recent studies use some read-outs of translation (presuming that translation is the single major energy sink in cells - however, this is objectively correct only in rapidly proliferating cells). That said, the authors take an interesting approach - to use two clickable methionine analogs, and assess baseline vs metabolically coupled translation within the same cell.

      The highlight is the methodology development where two distinct, clickable CMAs are used (to replace methionine in proteins). The labelling and approach are clever, and can be useful if carefully used. But there is going to be a challenge in using this, since the depletion of methionine itself (required for labelling), and a bias towards incorporation in some proteins (because these reagents are not highly permeable) will make it challenging to obtain precise metabolic state information - which otherwise can be obtained directly and far more precisely using a combination of other methods (ATP/flux measurement, respiratory capacity, translation rates etc). I therefore only make broader comments in my review below - to help structure this study better, clearly identify key limitations (and there are several that are clearly seen), and better clarify what MARBL may be useful for.

      (1) These reagents used for MARBL are not highly cell permeable/transported into cells, and largely work on the surface proteome. Which means the measurements related to changes in translation are indirect - quantified based on changes on the surface proteome (seen with labelling), and not the entire proteome.

      We appreciate the reviewer’s careful consideration of the MARBL labeling strategy and the opportunity to clarify an important feature of the method. We interpret this comment to mean that the detection reagents used in MARBL are not cell-permeable, rather than that the methionine analogues AHA and HPG are inefficiently transported into cells. HPG and AHA are transported through the ubiquitously expressed sodium-dependent neutral amino acid transporter SLC1A5 (2), and are incorporated into the proteome by intracellular translational machinery. The bias identified by the reviewer arises from the click detection step, not from metabolic labeling. We agree with the reviewer that MARBL measures the appearance of a subset of newly synthesized proteins where the amino acid analogue is accessible at the cell surface rather than the nascent production of the entire proteome. However, this is an intentional feature of MARBL and represents a key technical advance. It enables MARBL to generate a translation-dependent signal in live cells without requiring destructive interventions like fixation and permeabilization to access the entire proteome, as is required for other approaches such as BONCAT, THRONCAT, CENCAT, and SCENITH (1,3–5). Importantly, we empirically demonstrate that the surfaceaccessible fraction of the nascent proteome provides sufficient signal for quantitative measurements of translation that respond predictably to inhibition of protein synthesis as well as to metabolic perturbations that alter cellular energetics (Main Figure 1D-G, 2C). MARBL is also readily compatible with fixation and permeabilization, allowing the same labeling strategy to quantify analogue incorporation when a measure of total protein synthesis is desired and there is no need for the recovery of live cells. In the revised manuscript, we will include data directly comparing MARBL surface labeling with total nascent protein synthesis measured following fixation and permeabilization to show that these signals track with one another.

      (2) A major limitation - which will confound any interpretation- is the need to use methionine-free media; this can be a problem beyond protein synthesis/met incorporation since this will almost instantly deplete SAM pools in cells. There is little data provided on the impact of using this approach on SAM pools (time kinetics, how quickly SAM pools are affected, how much of the impact on metabolism comes from purely that, etc).

      This is important to establish because (i) of the continuous, very high flux of SAM -> SAH (eg. in the folate pathway, other methylations), and a constant need for SAM synthesis from methionine. This information will set the limits of capabilities of this method, as well as help delineate how much you can interpret results related to metabolic states between compared cells/states, etc. The labelling process is ~2 hours, while effects on SAM can be seen within minutes of methionine starvation in media, in metabolically active cells.

      The reviewer is correct in pointing out that cellular SAM pools are dynamic and adapt rapidly to methionine availability within hours of its deprivation or add-back (6,7). In the current version of the manuscript, we observe an ~100-fold decrease in intracellular SAM with 2 hours of methionine-free incubation in Jurkat cells, corresponding to the same time window over which we label with AHA in MARBL (currently presented as a min-max scaled heat map in Supplementary Figure 1A-B). We agree this needs to be presented more clearly as a potential limitation and quantified more explicitly. To address this, we plan to (i) reformat the existing SAM/ SAH/ methionine kinetics as absolute peak areas (Author response image 1A-C), (ii) add a novel dual-isotope simultaneous measurement of ATP turnover and SAM turnover across a panel of human cell lines to define the extent to which changes in SAM metabolism during the MARBL labeling workflow are likely to influence cellular energy demand (using <sup>13</sup>C<sub>5</sub>-methionine and H<sub>2</sub><sup>18</sup>O), and (iii) add β-ES as a threonine-based alternative to HPG for measuring baseline translation that substantially reduces the duration of methionine deprivation required for MARBL. β-ES is a clickable, alkyne-modified threonine analogue that is incorporated into newly synthesized proteins by mammalian cells in complete medium (1,4). Compared to methionine, which is a major part of the methionine cycle and 1-carbon metabolism that are required for cell proliferation and survival, threonine plays a less extensive role in mammalian intermediary metabolism. We plan to optimize and validate β-ESàAHA as well as AHAàβ-ES dual labeling and determine whether this modified MARBL workflow preserves the dynamic range and metabolic responsiveness of the HPGàAHA approach as in the original submitted manuscript. These experiments are currently underway, and we have already found that β-ES produces robust surface signal above background within 2-4 hours of labeling using our established MARBL protocol in the presence of normal threonine levels (Author response image 2A-D). We are currently optimizing the surface labeling protocol to further reduce the duration of baseline β-ES incubation in the revised manuscript. We are unable to eliminate methionine depletion entirely from the MARBL workflow. This is because endogenous methionyl-tRNA synthetases possess drastically higher affinities for canonical methionine (8–10), which prevents the incorporation of methionine analogues into proteins. However, these planned experiments will better define the impact of methionine limitation while also providing an alternative MARBL implementation that restricts methionine withdrawal to the shorter AHA-labeling window.

      Author response image 1.

      Dynamics of intracellular methionine and its derived metabolites. (A-C) LC-MS/MS raw peak areas in Jurkat cells of methionine (A), SAM (B), and SAH (C) after 2, 4, 6, or 8 hours of incubation in Met-free RPMI media supplemented with Met, AHA, or HPG. Abbreviations: AHA = Azidohomoalanine, HPG = Homopropargylglycine, Met = Methionine, SAM = Sadenosylmethionine, SAH = S-adenosyl-L-homocysteine. Statistics: Graphs display mean ± SD (A-C).

      Author response image 2.

      Extension of live surface labeling protocol using clickable threonine analogue to monitor surface translation. (A) Chemical structures for threonine and its clickable analogue β-ES. (B-C) Representative flow cytometric histograms (B) and corresponding gMFIs of extracellular Azd-647 signal after 1, 2, or 4 hours of β-ES incorporation. (D) gMFIs of Alk-647 (corresponding to AHA) or Azd-647 (corresponding to HPG or β-ES) after 1, 2, or 4 hours of incorporation. Abbreviations: β-ES = β-Ethynylserine, AHA = Azidohomoalanine, HPG = Homopropargylglycine, Met = Methionine, CHX = Cycloheximide, gMFI = Geometric mean fluorescence intensity. Statistics: Graphs display mean ± SD (C-D).

      (3) What is the effect on overall adenylate charge/ATP due to shifting to methionine-free media + label addition? How does it vary between the cells tested (suspension vs adherent)? This should be established before data related to Fig. 2. Does it correlate with extent of AHA incorporation?

      The primary conclusion that this method suitably reflects overall changes in energetics comes from the titration of 2DG/glycolytic inhibition.

      We thank the reviewer for raising these important points, which are related to point #2 above, as both ask how methionine-free labeling conditions alter cellular metabolism. We agree that a more comprehensive set of experiments exploring methionine deprivation will better delimit interpretations that can be drawn from a MARBL assay related to cellular energetics. As described in our response to point #2, we will measure baseline adenylate energy charge across a panel of adherent and suspension cell lines under methionine-replete, methionine-free, and methionine-free +AHA/HPG labeling conditions and determine its relationship to AHA incorporation. We will also perform analogous experiments during β-ES supplementation to determine how this alternative labeling condition affects cellular energetic state and β-ES incorporation. 

      We also wish to clarify that the adenylate energy charge measurements in Main Figure 2D and the measurements of newly synthesized ATP by H<sub>2</sub><sup>18</sup>O turnover in Main Figure 2E were performed in methionine-free media supplemented with either AHA or methionine. We apologize that this was not made clear in the figure key and legend. In addition, we inadvertently covered the methionine-replete control in the original figure with the key. This has been corrected in Author response image 3A-B and will be updated in the revised manuscript. In Jurkat cells, both adenylate energy charge and ATP synthesis are similar in the presence and absence of methionine.

      We note, however, that cells with intact energy-generating systems maintain adenylate energy charge within a relatively narrow range (11). Given this buffering capacity, we do not expect methionine deprivation to produce a large change in adenylate energy charge in most cell types. 

      Author response image 3.

      Validation of surface translation as a readout of cellular energetics in methionine-free and AHA-supplemented media. (A) Correlation between changes in adenylate energy charge versus normalized Alk-647 Alkyne flow cytometric signal in methionine-free RPMI supplemented with either AHA or Met (Pearson’s R = 0.9044, Pearson’s R<sup>2</sup> = 0.8179, p = 0.0052). (B) Correlation between changes in newly synthesized ATP versus normalized Alk-647 Alkyne flow cytometric signal in methionine-free RPMI supplemented with either AHA or Met (Pearson’s R = 0.8854, Pearson’s R<sup>2</sup> = 0.7840, p = 0.008). Abbreviations: AHA = Azidohomoalanine, HPG = Homopropargylglycine, Met = Methionine, CHX = Cycloheximide, gMFI = geometric mean fluorescence intensity, 2DG = 2-Deoxy-D-Glucose, Omy = Oligomycin A. Statistics: Pearson’s correlation coefficient was used for correlation and significance (A-B). The Met + DMSO condition was excluded from the Pearson correlation analysis. Graphs display mean ± SD (A-B).

      (4) Relatedly, if this label incorporation experiment is carried out (for ~2 hrs), and subsequently there is a washout/replacement with fresh, methionine-supplemented medium, (how quickly) do the cells recover and restore their energetic allocations?

      If we observe substantial changes in adenylate energy charge associated with methionine restriction, as assessed in the experiments planned in Points #2 and #3 above, we will perform methionine add-back experiments to define the kinetics of recovery. In addition, our plan to provide an alternative workflow that replaces HPG with β-ES will reduce the total duration of methionine depletion and further address this concern.

      (5) One possible advantage of a system like this can be to address questions in single cells/study cell metabolic heterogeneity. However, these are best done if the attaching moiety has a (selective) fluorescence increase and/or other read-out that can be quantitatively obtained at a single cell level. Largely, using AHA or HPG effectively only leads to bulk estimates (which can be sub-sorted towards single-cell estimates indirectly). This means that this method cannot really be used to study cell-cell metabolic heterogeneity effectively - compared to far simpler approaches, for example using a mitochondrial potentiometric dye with high fluorescence, or reporters for glycolytic activity, etc. This would also be related to Figure 5 - at best, this approach may complement existing approaches towards identifying heterogeneous sub-populations of cells. 

      However, I do agree that MARBL is flexible, stable, and can be internally normalised and used through flowbased platforms. It can supplement existing approaches to perturb bioenergetics, and also supplement existing approaches to understand metabolic state in live cells, particularly in suspension cells.

      We thank the reviewer for raising this thoughtful point, which we will address with textual changes in the discussion. We note that MARBL provides a quantitative single-cell readout when flow cytometry is used as the analytical endpoint, because the dual-color labeling scheme provides an internal control for baseline translation that accounts for variation in protein synthesis rates independent from cellular energetics. This permits a single metabolic resilience index to be calculated per cell and interpreted either at single-cell resolution or after grouping cells into sub-populations as a bulk estimate. Both approaches have utility depending on the biological question and intended downstream application. We agree with the reviewer that MARBL can be paired with other reporters for metabolism to obtain a more comprehensive understanding of bioenergetic heterogeneity and appreciate the reviewer’s recognition of its value as a complementary approach for studying metabolic state in live cells.

      Reviewer #2 (Public review):

      Summary: 

      Delacruz et al. describe a new method, called “MARBL” (Methionine Analogues for Ratiometric Bioenergetics in Live cells) to measure metabolic activity in single cells. The concept is similar to the SCENITH (anti-puromycin flow cytometry) assay to measure energy metabolism by measuring protein translation activity, yet offers, in theory, two advantages: 1) it keeps cells alive for downstream biological assays and 2) it is a ratiometric measurement, measuring both baseline translation and translation in the presence of metabolic inhibitors to correct for inherent cell-to-cell translation differences.

      Specifically, this method takes advantage of two click-chemistry-active methionine analogs, and then clicks fluorophores onto newly-synthesized surface proteins that have incorporated these analogs to measure translational activity. One methionine analog is given to cells for 2-4 hours to measure baseline translational activity, then metabolism is blocked using 2-deoxyglucose and oligomycin and the second methionine analog given to measure “metabolically-linked” translation activity. The authors establish this technique and show that mouse T cells polarized as pathogenic Th17 cells are more translationally active (“resilient”) compared to nonpathogenic Th17 cells, and when sorted, the resilient cells produce more interferon-gamma. This latter finding requires live cells after the metabolic measurement assay, showcasing findings that are inaccessible to the SCENITH assay.

      Strengths:

      The approach used is conceptually clever. It is appealing to measure metabolic/translational activity and to then be able to carry out further assays on sorted cell populations with different degrees of metabolic activity. This would indeed represent a useful advance.

      We thank the reviewer for recognizing the conceptual strengths of MARBL and its potential to enable downstream analysis of live cell populations with distinct metabolic states.

      Weaknesses:

      In principle, one key benefit of this technique is that cells can be used for biological assays after the metabolic measurement. Indeed, this would represent a valuable tool in the field.

      However, in this technique, cells are subjected to methionine deprivation, addition of non-natural methionine analogs, click chemistry, and high doses of toxic metabolic inhibitors 2-deoxyglucose and oligomycin. Indeed, the authors show in Figure S5F that 1/3 more of the post-MARBL cells die relative to cells not subject to this technique (60% viability in unclicked control, 40% in MARBL-measured cells). This data suggests that cells after this technique may be stressed and not reflective of the biological function of unmanipulated cells. More controls on viability and cell function (e.g. cytokine production) at more time points after the MARBL assay would have been valuable to address this issue.

      The reviewer makes important points regarding the conditions needed for the workflow of a MARBL assay that may impact cellular fitness. Perturbational methods, particularly techniques that interrogate bioenergetics, inherently require media-based or pharmacologically induced stress to evaluate cellular responses. Still, we agree that comprehensively profiling the fitness of cells following a MARBL assay and sorting is important since our technique aims to link cellular bioenergetics to functional outcomes. First, we would like to highlight that the MARBL-processed pathogenic and non-pathogenic TH17 cells depicted in Main Figure 5 were rested overnight after Fluorescence-Activated Cell Sorting (FACS) in complete RPMI medium prior to the restimulation assay. This is in line with standard practice to allow cells to recover from the shear stress of sorting prior to subsequent experiments (12,13). In addition, both pathogenic and non-pathogenic TH17 cells maintained the expression of lineage-defining transcription factors throughout the MARBL workflow, as shown by analyzing rested cells stained with antibodies for T-bet (expressed by pathogenic TH17) as well as RORγt (expressed in both cell types) by flow cytometry (Author response image 4A-C). To address this in the revised manuscript, we plan to repeat our non-pathogenic and pathogenic TH17 dual-MARBL and sorting experiment (Main Figure 5A-B) and rest sorted cells in RPMI with IL-2 for longer periods of time (24 or 48 hours) before restimulation for viability and cytokine analysis. These controls will provide more information on cellular fitness throughout the MARBL workflow, and we appreciate the reviewer’s suggestion.

      Author response image 4.

      Expression of lineage-defining transcription factors is maintained post-MARBL processing and fluorescence-activated cell sorting (FACS). (A) Experimental schematic. Ex vivo differentiated pTH17 and npTH17 cells were stained with CD45.2 antibodies conjugated to different color fluorophores, processed via MARBL, mixed at a 1:1 ratio, sorted, and then re-cultured in IL-2-supplemented RPMI media. After resting overnight, the expression of RORγt and T-bet were evaluated by intracellular staining and flow cytometry, distinguishing pTH17 from npTH17 cells based on prior CD45.2 staining. (B-C) Frequency of RORγt (B) and T-bet (C) positivity in murine Th17 cells post-MARBL processing, sorting, and overnight rest in IL-2-supplemented media. Abbreviations: npTH17 = non-pathogenic TH17, pTH17 = pathogenic TH17, HPG = Homopropargylglycine, AHA = Azidohomoalanine, Met = Methionine, 2DG = 2-Deoxy-D-Glucose, Omy = Oligomycin A. Statistics: Graphs display mean ± SD (B-C).

      Another weakness of the paper is limited benchmarking against established metabolic assays in the field. The main assays used currently in the field are SCENITH and Seahorse. The authors do not compare their findings to SCENITH. They do compare their results to Seahorse, but the data shown don't address the key question: how does energy production measured by Seahorse, say in unmanipulated vs 2dg+oligomycin-treated cells, compare to the MARBL measurement? (Instead, they show a calculated "glucose dependence" metric in cells subjected to low vs high inhibitor dose, not showing the underlying data or cells that didn't receive an inhibitor).

      We appreciate the opportunity to perform additional benchmarking to define how the MARBL signal compares to existing methods for measuring cellular energetics. For our initial validation, we chose to benchmark the single-color MARBL signal against direct LC-MS/MS quantification of adenylate energy charge as well as newly synthesized ATP (Main Figure 2A-E). Although Seahorse and SCENITH are widely used standards in the field, these assays still provide indirect measures of cellular energetics, whereas LC-MS/MS quantifies the high-energy nucleotide pools that dictate cellular energy status. Both the adenylate energy charge and newly synthesized ATP decreased in response to increasing concentrations of 2-deoxy-D-glucose (2DG) and oligomycin A (Omy) treatments that impair ATP regeneration, and this energetic response displayed a linear relationship with the AHA click signal (Main Figure 2D-E). Nevertheless, we agree that additional benchmarking suggested by the reviewer will strengthen the methodological foundation of MARBL and help users understand how MARBL measurements relate to those obtained using more established metabolic assays. For this purpose, it is important to account for differences in what each assay measures. Seahorse resolves oxidative and glycolytic activity through simultaneous measurements of OCR and ECAR, whereas MARBL (as well as SCENITH and CENCAT) integrate the energetic contributions of these pathways into a single translation-dependent readout. This is the reason why we compared MARBL with Seahorse using the glucose-dependence calculation employed by SCENITH in the current version of the manuscript (Main Figure 2FH) (5). As suggested by the reviewer, we will assess how OCR and ECAR measured by Seahorse vary relative to the MARBL signal across different oligomycin and 2-deoxyglucose treatment conditions, providing a more comprehensive view of the bioenergetic responses captured by MARBL. 

      References

      (1) Ignacio BJ, Dijkstra J, Mora N, Slot EFJ, van Weijsten MJ, Storkebaum E, et al. THRONCAT: metabolic labeling of newly synthesized proteins using a bioorthogonal threonine analog. Nat Commun. 2023 Jun 8;14(1):3367. doi:10.1038/s41467-023-39063-7 PubMed PMID: 37291115; PubMed Central PMCID: PMC10250548.

      (2) Pelgrom LR, Davis GM, O’Shaughnessy S, Wezenberg EJM, Van Kasteren SI, Finlay DK, et al. QUAS-R: An SLC1A5-mediated glutamine uptake assay with single-cell resolution reveals metabolic heterogeneity with immune populations. Cell Reports. 2023 Aug 29;42(8):112828. doi:10.1016/j.celrep.2023.112828

      (3) Dieterich DC, Link AJ, Graumann J, Tirrell DA, Schuman EM. Selective identification of newly synthesized proteins in mammalian cells using bioorthogonal noncanonical amino acid tagging (BONCAT). Proceedings of the National Academy of Sciences. 2006 Jun 20;103(25):9482–7. doi:10.1073/pnas.0601637103 PubMed PMID: 16769897.

      (4) Vrieling F, van der Zande HJP, Naus B, Smeehuijzen L, van Heck JIP, Ignacio BJ, et al. CENCAT enables immunometabolic profiling by measuring protein synthesis via bioorthogonal noncanonical amino acid tagging. Cell Rep Methods. 2024 Oct 21;4(10):100883. doi:10.1016/j.crmeth.2024.100883 PubMed PMID: 39437716; PubMed Central PMCID: PMC11573747.

      (5) Argüello RJ, Combes AJ, Char R, Gigan JP, Baaziz AI, Bousiquot E, et al. SCENITH: A flow cytometry based method to functionally profile energy metabolism with single cell resolution. Cell Metab. 2020 Dec 1;32(6):1063-1075.e7. doi:10.1016/j.cmet.2020.11.007 PubMed PMID: 33264598; PubMed Central PMCID: PMC8407169.

      (6) Mentch SJ, Mehrmohamadi M, Huang L, Liu X, Gupta D, Mattocks D, et al. Histone Methylation Dynamics and Gene Regulation Occur through the Sensing of One-Carbon Metabolism. Cell Metab. 2015 Nov 3;22(5):861–73. doi:10.1016/j.cmet.2015.08.024 PubMed PMID: 26411344; PubMed Central PMCID: PMC4635069.

      (7) Chen Z, Chen W, Reheman Z, Jiang H, Wu J, Li X. Genetically encoded RNA-based sensors with Pepper fluorogenic aptamer. Nucleic Acids Res. 2023 Sep 8;51(16):8322–36. doi:10.1093/nar/gkad620 PubMed PMID: 37486780; PubMed Central PMCID: PMC10484673.

      (8) Kiick KL, Saxon E, Tirrell DA, Bertozzi CR. Incorporation of azides into recombinant proteins for chemoselective modification by the Staudinger ligation. Proc Natl Acad Sci U S A. 2002 Jan 8;99(1):19–24. doi:10.1073/pnas.012583299 PubMed PMID: 11752401; PubMed Central PMCID: PMC117506.

      (9) Beatty KE, Liu JC, Xie F, Dieterich DC, Schuman EM, Wang Q, et al. Fluorescence visualization of newly synthesized proteins in mammalian cells. Angew Chem Int Ed Engl. 2006 Nov 13;45(44):7364–7. doi:10.1002/anie.200602114 PubMed PMID: 17036290.

      (10) Kiick KL, Weberskirch R, Tirrell DA. Identification of an expanded set of translationally active methionine analogues in Escherichia coli. FEBS Lett. 2001 Jul 27;502(1–2):25–30. doi:10.1016/s0014-5793(01)02657-6 PubMed PMID: 11478942.

      (11) De la Fuente IM, Cortés JM, Valero E, Desroches M, Rodrigues S, Malaina I, et al. On the dynamics of the adenylate energy system: homeorhesis vs homeostasis. PLoS One. 2014;9(10):e108676. doi:10.1371/journal.pone.0108676 PubMed PMID: 25303477; PubMed Central PMCID: PMC4193753.

      (12) Pollizzi KN, Patel CH, Sun IH, Oh MH, Waickman AT, Wen J, et al. mTORC1 and mTORC2 selectively regulate CD8<sup>+</sup> T cell differentiation. J Clin Invest. 2015 May 1;125(5):2090–108. doi:10.1172/JCI77746 PubMed PMID: 0.

      (13) Roth TL, Puig-Saus C, Yu R, Shifrut E, Carnevale J, Li PJ, et al. Reprogramming human T cell function and specificity with non-viral genome targeting. Nature. 2018 Jul;559(7714):405–9. doi:10.1038/s41586-018-03265

    1. eLife Assessment

      This valuable pharmacological MEG study implicates GABA-A receptors in the shaping of intrinsic neuronal timescales in the human brain. While the authors provide solid evidence for the claims that (i) GABAergic neuronal inhibition prolongs timescales in specific cortical networks and (ii) NMDA receptor manipulations have no detectable effect, the manuscript could be further strengthened by resolving some apparent discrepancies in the results as well as additional validation of the network analysis used. This work will be of interest for experimental and theoretical neuroscientists studying brain-wide network dynamics and the large-scale organization of neurotransmitter systems.

    2. Reviewer #1 (Public review):

      Summary:

      The authors investigate whether systematic manipulation of GABA-A and NMDA receptors influences large-scale cortical neuronal timescales and transient network dynamics. 57 healthy male participants completed placebo, lorazepam and D-cycloserine sessions in a double-blind, within-participant, cross-over design. The authors estimated timescales from the knee frequency of the aperiodic component of MEG power spectra, and identified transient large-scale cortical networks using a time-delay embedded hidden Markov model (TDE-HMM) analysis. The authors report that lorazepam increases estimated timescale across multiple cortical areas, with particularly pronounced changes in states interpreted as frontal default-mode network (DMN) and dorsal attention network (DAN). These changes include increased fractional occupancy of the DMN and decreased for DAN. D-cycloserine did not significantly affect neuronal timescales.

      Strengths:

      The study has several notable strengths:

      (1) The pharma-MEG study is well designed and executed (e.g., a crossover design, collecting subjective, cardiovascular and additional control measurements).

      (2) Neuronal-timescales maps are validated against previously published cortical timescale and hierarchy maps.

      (3) TDE-HMM findings are examined using alternative numbers of hidden states and different preprocessing pipelines.

      (4) The manuscript is generally clear and well written.

      Weaknesses:

      Nevertheless, several aspects of the primary neuronal timescale measure and the TDE-HMM analysis require further validation, and several points should be clarified:

      (1) A central concern is whether the reported measure of neuronal timescale, on its own, is sufficient to support the interpretation assigned to it. Lorazepam has been shown to alter several spectral parameters, including oscillatory peaks and the aperiodic component of power spectra, in ways that could influence knee-frequency fitting. The manuscript, however, does not provide sufficient information about fit quality (across participants, parcels, and conditions), parcels within each participant/condition with identifiable knees, or the sensitivity of the results to the fitting range and peak settings (i.e., FOOOF parameters). Importantly, Figure S3 indicates that neither lorazepam nor D-cycloserine significantly affects knee frequency, although the neuronal timescale measure is mathematically derived from that knee frequency (i.e., tau = 1/(2*pi*f_knee)). Although such a result is mathematically possible, the pattern is unusual and requires explanation because timescales are not measured independently, but rather derived from the knee frequency. Furthermore, the neuronal timescale maps show a relatively homogenous increase across the cortex under lorazepam (Figure 2A), whereas the knee-frequency maps show a heterogeneous spatial pattern, with increased knee frequency under lorazepam in frontal regions, and decreased in occipital and temporal regions (Figure S3). A widespread significant effect appearing only after inversion could reflect a genuine effect, particularly in parcels with low knees to begin with. However, it could also arise if a small number of extremely low fitted knee frequencies (below 1Hz?) produce very long timescale estimates and disproportionately affect the statistical analysis.<br /> The authors should address this discrepancy by presenting the distributions of knee-frequencies and neuronal timescales, and their ranges. Additional information about outliers and the robustness of the findings to extreme fitted values would also be valuable.

      (2) A second concern relates to the TDE-HMM analysis. The model identifies states from the covariance matrix of time-delayed time-courses. Consequently, states are differentiated partly on the basis of their spectral and temporal characteristics. Knee frequencies for each state are then estimated from the power spectra of the same states used to define them. Differences in neuronal timescales across states may therefore be expected, at least in part, simply by the way the states were inferred. This point weakens the claim that cortical states are independent, and "operate on distinct timescales". A clearer separation between state definition and neuronal timescale estimation would be needed to establish that the reported state-specific differences are not partly an expected consequence of the fitted model. This could be addressed using, e.g., cross-validation or simulations. This concern is less substantial for the drug-related changes in neuronal time scale within individual states.

      (3) The result and methods sections provide insufficient detail and statistical reporting for the TDE-HMM analysis. In the result section (page 10), the authors report only the *range* of fractional occupancies across all states. A range of 1% to 48% is substantial. It is therefore important to determine whether some states occupied only 1% (corresponding to ~3sec) while others occupied a much larger proportion. The authors should report the mean fractional occupancy of each state, together with its SD and range across participants.

      (4) The authors should explain the apparent discrepancy between the relatively homogeneous increase in neuronal timescales under lorazepam across the complete recording (Figure 2) and the heterogeneous state-specific changes (both increases and decreases; Figure 4B). Could this pattern be explained, at least in part, by the differences in fractional occupancy of individual states? This provides an additional reason to report state-specific occupancy values in greater detail.

      (5) On page 13, 2nd paragraph, the authors state that the findings indicate that the frontal DMN is an important driver of the global prolongation of neuronal timescales, based on the strong similarities to the time-averaged timescales. However, Figure S5 appears to show that the spatial correlation between state-specific and time-averaged timescales is as high for State 8 and is similar for State 2. The current statement therefore appears to be an overinterpretation. Demonstrating that the correlation for State 3 is significantly larger than the correlations for the other states would provide stronger support for this claim.

      (6) Spatial maps are presented inconsistently throughout the main manuscript and in the supplementary. Specifically, some figures only show thresholded maps (at p<0.05 or p<0.001; e.g., Figure 4b), whereas others show unthresholded maps. Reporting only the number of parcels exhibiting a significant effect, without presenting the corresponding thresholded maps, makes it difficult to evaluate the spatial distribution of the reported changes. At least for the main findings (e.g., lorazepam effects on neuronal timescales), I recommend presenting both thresholded and unthresholded maps within the same figure.

      (7) Figure 3: The rationale for presenting the mean power between 3-30 Hz is unclear. The authors should explain why this metric was selected to represent each state. Visual inspection suggests that the state-specific power spectra differ across several dimensions, including the aperiodic exponent, offset, alpha power, etc. Characterizing states using these parameters may be more informative. Each of these features could potentially also influence the estimated knee frequency.

      (8) The conclusion on page 17 (and similar statement in the intro): "...these findings provide causal evidence that microscale synaptic inhibition directly shapes both local neuronal timescale organization... ") appears to be overstated based on the evidence presented. Although the lorazepam manipulation supports a causal effect of the drug on MEG-derived cortical timescales and network dynamics, the study does not directly measure microscale synaptic inhibition or establish a direct mechanistic link across scales. The phrasing should therefore distinguish the observed pharmacological effects from the inferred role of GABAergic inhibition and be phrased more cautiously.

      (9) The method section lacks several critical details, including: criteria for excluding MEG channels and noisy segments (see comment below), details on source reconstruction (e.g., number of vertices used to project the sensor-level data), procedures used to generate the null distributions for the spatial autocorrelation preserving permutation test, transition probability analysis and more. Relatedly, the preprocessing pipeline of the MEG data appears to be based on manual inspection for the removal of channels, ICA components and noisy segments. This procedure is inherently subjective. The authors should provide details on the exact criteria used to make these decisions. Critically, the authors should also report the duration of usable resting-state data remaining after segment rejection. The Methods section should provide sufficient detail to allow readers to evaluate the validity of the study and reproduce the analysis as closely as possible. In its current form, it does not do so.

    3. Reviewer #2 (Public review):

      This work provides empirical data on how GABA and NMDA agonists globally affect timescales as measured through MEG. The authors reproduce the previously observed gradient of intrinsic timescales in the placebo condition, as well as its relationship to cortical hierarchy in T1/T2w maps. Timescales were not fixed, but dynamic, as revealed by large-scale network analysis separating into discrete network states. Pharmacologically, GABA agonist Lorazepam produced a brain-wide increase in timescales while NMDA agonist D-cycloserine did not. Furthermore, the GABA-mediated increase in timescale was area- and state-dependent, and more detailed analyses show changes in state occupancy mainly for DMN and DAN.

      Overall, the paper contributes valuable data on a relevant topic of research in understanding the timescales of network dynamics at the local and global level. The question is well-motivated, and the analyses are technically sound and described in a straightforward manner. The network-level analysis in TDE-HMM is interesting and provides a complementary and more fine-grained perspective to the global timescale gradient, both in terms of space and time. The hypotheses were straightforward since both GABA and NMDA have relatively long timescales (of the dominant synaptic currents), though the lack of effect from NMDA-agonist is quite surprising but reasonably explained by the voltage-dependence of NMDA receptors in such a task-free setting. The paper overall is clearly written, and the figures are of high-quality, though some things could be presented in slightly more informative ways (see below). I have some questions and minor suggestions, but don't have too much to criticize as a whole.

      My biggest question is the following: the mixed effects model shows that lorazepam additionally mediates timescale over and above the hierarchy (myelination map). This leaves a very clear gap. What the authors also probably want to show is that greater GABA_A receptor expression (of any or all the subunits) results in greater change under lorazepam, not against the myelination map only, or that the residue can be explained by the GABA_A maps. This, of course, would not explain the non-stationary nature of the dynamics (and state-dependent timescale maps), but would give a more direct explanation of the spatial effect. Since you already compared to the prior MEG map in Shafiei et al., 2023, I guess it's not a huge technical effort to grab the gene maps from neuromaps (https://github.com/netneurolab/neuromaps). I think this could strengthen the current manuscript.

    1. eLife Assessment

      This study presents a valuable investigation of how inhibition of the WNK-SPAK/OSR1 pathway influences neuronal chloride homeostasis and seizure-like activity in organotypic hippocampal slice cultures. The evidence supporting the principal mechanistic conclusions is incomplete because several proposed mechanisms are inferred from changes in chloride dynamics rather than directly demonstrated. The work will be of interest to researchers studying chloride homeostasis, inhibitory neurotransmission, and epilepsy.

    2. Reviewer #1 (Public review):

      Summary:

      The authors used extracellular field potential recordings and two-photon imaging to monitor neuronal network activity and intracellular chloride concentration ([Cl-]i) in organotypic hippocampal slices from mice expressing the genetically encoded chloride fluorophore Clomeleon. These slices were used as a model of acute traumatic brain injury and epileptogenesis in vitro. The study provides evidence that blocking the WNK-SPAK/OSR1 pathway with WNK463 alleviates epileptic activity, and that this anticonvulsant effect involves suppression of the chloride loader NKCC1 and enhancement of chloride extrusion via KCC2. Overall, this is a solid study with a comprehensive pharmacological analysis.

      Strengths:

      The conclusions are well supported by the detailed pharmacological analysis.

      Weaknesses:

      The only weakness I see is the absence of cellular-level electrophysiology, which precludes interpretation of the imaging data in the context of GABA action polarity.

    3. Reviewer #2 (Public review):

      Summary:

      The authors investigate whether inhibition of the WNK-SPAK/OSR1 pathway using the allosteric inhibitor WNK463 improves neuronal chloride homeostasis and suppresses epileptiform activity in organotypic hippocampal slice cultures. Using Super Clomeleon imaging combined with extracellular field recordings, they demonstrate that WNK463 accelerates recovery of intracellular chloride following chloride loading, reduces interictal chloride accumulation, and progressively suppresses recurrent ictal-like discharges. Pharmacological inhibition and siRNA-mediated knockdown of NKCC1 and KCC2 are then used to investigate the contribution of these transporters to the anti-ictal effects of WNK463.

      Strengths:

      The study addresses an important question in the field of chloride homeostasis and epilepsy and combines complementary experimental approaches. In particular, the distinction between baseline chloride measured in the presence of TTX and activity-dependent interictal chloride accumulation provides a useful conceptual framework for interpreting previous studies of WNK-SPAK inhibition. The imaging, electrophysiological, and pharmacological data are internally consistent and support the conclusion that WNK463 alters chloride dynamics and substantially suppresses ictal-like activity in this model.

      Weaknesses:

      The principal limitation of the manuscript is that several mechanistic conclusions extend beyond the experimental observations. Throughout the results and discussion, the authors interpret the observed changes in intracellular chloride dynamics as evidence of enhanced CCC-mediated chloride extrusion, while the pharmacological and siRNA-mediated experiments are interpreted as supporting coordinated NKCC1 inhibition and KCC2 activation, ultimately leading to restoration of GABAergic inhibition and negative shifts in EGABA. While these interpretations are plausible and consistent with the data, they remain inferential because transporter phosphorylation or activity, EGABA, and inhibitory synaptic function were not directly assessed in the current study. Moreover, although the pharmacological and knockdown experiments support a contribution of NKCC1 and KCC2 to the actions of WNK463, they do not definitively establish coordinated modulation of both transporters as the primary mechanism underlying seizure suppression. These mechanistic conclusions should therefore be presented more cautiously.

      The manuscript would also benefit from broader contextualization within the current literature. The introduction largely focuses on previous work from the authors' group and provides a relatively narrow overview of chloride homeostasis in epilepsy. In particular, the discussion would benefit from broader consideration of studies examining KCC2 dysfunction in human epilepsy and experimental models, alternative mechanisms regulating KCC2 activity following seizures, and recent therapeutic strategies targeting KCC2.

      Finally, although the authors appropriately acknowledge that the experiments were performed exclusively in vitro, the discussion could more explicitly address the limitations of the organotypic hippocampal slice model, including how culture-induced network reorganization and spontaneous epileptiform activity may influence chloride homeostasis and the extent to which these findings generalize to traumatic brain injury and chronic epilepsy in vivo. In addition, the statistical analysis would benefit from clarification regarding the experimental unit and the treatment of repeated measurements.

    4. Reviewer #3 (Public review):

      Summary:

      Dzhala and colleagues present findings from organotypic slice cultures suggesting that simultaneous modulation of the complementary cation-chloride cotransporters NKCC1 and KCC2 through inhibition of the WNK-SPAK/OSR1 pathway reduces seizure-like activity. The manuscript is generally well written, and the data support the conclusion that WNK463 exerts robust anti-ictal effects in this model. However, several issues should be addressed to strengthen the mechanistic interpretation and statistical rigor of the study, and improve confidence in the conclusions.

      Strengths:

      (1) The study addresses an important mechanistic question by investigating how inhibition of the WNK-SPAK/OSR1 pathway with WNK463 influences seizure activity and neuronal chloride homeostasis.

      (2) The experimental design is logical and comprehensive, progressing from characterization of chloride dynamics to pharmacological and genetic interrogation of the underlying mechanism using multiple complementary approaches, including pharmacological inhibition, siRNA-mediated knockdown, electrophysiology, and chloride imaging.

      (3) The combination of simultaneous extracellular electrophysiology and two-photon chloride imaging provides complementary functional and mechanistic information and represents a major technical strength of the study.

      (4) The TTX experiments elegantly distinguish activity-dependent chloride accumulation from resting intracellular chloride concentration, substantially strengthening the central mechanistic conclusions.

      Weaknesses:

      (1) The mechanistic conclusions regarding KCC2 activation and NKCC1 inhibition are stronger than the data directly support. Throughout the manuscript, the authors conclude that WNK463 activates KCC2 and inhibits NKCC1. Although this interpretation is consistent with the established biology of the WNK-SPAK/OSR1 pathway, the evidence presented here is indirect. Specifically, the authors infer KCC2 activation and NKCC1 inhibition from the observation that pharmacological inhibition or siRNA-mediated knockdown of these transporters alters the effects of WNK463, together with measurements of chloride dynamics. While these findings are compatible with a KCC2- and NKCC1-dependent mechanism, they do not directly establish that WNK463 regulates either transporter. Direct evidence would require measurements of transporter activity, phosphorylation state, membrane expression, or other biochemical indices of transporter regulation. I therefore recommend that the authors both temper the mechanistic language throughout the manuscript and explicitly acknowledge in the Discussion that the proposed regulation of KCC2 and NKCC1 is inferred from indirect evidence rather than directly demonstrated in the present study.

      (2) The conclusions drawn from the siRNA-mediated knockdown experiments should be interpreted more cautiously. First, it is unclear whether silencing NKCC1 or KCC2 induced compensatory changes in the expression or function of the complementary cotransporter. Given the well-established interplay between NKCC1 and KCC2 in regulating intracellular chloride homeostasis, compensatory adaptations could influence the interpretation of these experiments and should be addressed or acknowledged as a limitation. Second, the sample size for the siRNA experiments appears relatively small. It is unclear how many independent animals contributed slices to each experimental group, making it difficult to assess the degree of biological replication. In addition, effect sizes are not reported. Clarifying the number of biological replicates and reporting effect sizes would improve the rigor of the statistical analysis and increase confidence in these findings.

      (3) The final pharmacological experiments require clarification, as the conclusions appear internally inconsistent. Earlier experiments suggest that the anticonvulsant effects of WNK463 depend on coordinated regulation of both NKCC1 and KCC2. However, in the final experiment, the authors state that simultaneous pharmacological inhibition of NKCC1 and KCC2 does not prevent the anticonvulsant effects of WNK463. In contrast, the accompanying statistical analysis indicates that combined transporter inhibition significantly reduces the effect of WNK463 relative to control conditions. These interpretations appear inconsistent and make it difficult to determine the extent to which the anticonvulsant action of WNK463 depends on NKCC1 and KCC2. The authors should clarify whether simultaneous inhibition of both transporters completely abolishes, partially attenuates, or merely reduces the magnitude of the WNK463 response, and revise the text accordingly. If the effect is only partially attenuated, alternative mechanisms contributing to the anticonvulsant actions of WNK463 should also be considered and discussed.

      (4) The statistical analysis and reporting require further attention. First, median values should not be reported with standard deviations, as standard deviation describes variability around the mean rather than the median. For non-normally distributed data, the authors should report median values together with an appropriate measure of variability, such as the interquartile range (25th-75th percentile) or another suitable summary. Second, in several instances, ANOVA results are reported using only a single degree of freedom value (e.g., page 6, DF = 53). This is incomplete, as an F statistic is defined by two degrees of freedom: the numerator degrees of freedom (between-group variability) and the denominator degrees of freedom (within-group variability). Reporting statistical results using standard notation (F(df_between, df_within) = F statistic, p = value) would improve clarity and allow proper interpretation of the analyses.

    1. eLife Assessment

      This valuable study combines free-flight kinematics, aerodynamic modeling, optogenetics, and connectomics to ask how Drosophila stabilize roll perturbations during flight. The kinematic and modeling methods are solid, but the evidence for redundant motor control of roll stabilization is incomplete: optogenetic silencing experiments lack a positive control, and no experiment directly tests the central claim that redundant motor pathways drive robust roll control. This work will interest researchers studying insect flight and robust motor control.

    2. Reviewer #1 (Public review):

      Summary:

      The authors developed a free flight perturbation assay that induces roll instability. They set out to understand how steering muscles control and stabilize roll instability. In addition to the previously reported changes in wing amplitude, they find that a pure roll perturbation induces changes in stroke deviation. They silence motor neurons of individual steering muscles to show that silencing any one muscle alone is not sufficient to disrupt the recovery from roll perturbations. Through aerodynamic modelling, they argue that this occurs despite the fact that each muscle alone can induce changes in wing amplitude and/or stroke deviation. They suggest that this robustness to roll perturbation may be due to redundancy in the motor control program. Finally, they confirm redundant projections from the published haltere connectome study and show that indeed muscles which produce similar effects on wing kinematics receive redundant connections from the haltere afferents.

      Strengths:

      This study's strength lies in integrating findings from several adjacent areas in insect flight control research. The authors combine steering muscles physiology, wing kinematic quantification, aerodynamic modelling, and sensory (haltere) control of wing kinematics. This integrated approach provides a comprehensive discussion of the mechanisms that may underlie recovery from roll perturbations.

      Weaknesses:

      The biggest weakness is that, although the authors generate a very plausible and interesting hypothesis, much of the supporting evidence already occurs in existing datasets. In fact, a distributed or redundant many-to-one muscle control for flight kinematics is not a new idea. It has been suggested wherever researchers have examined muscle control for wing kinematics, across studies in flies, moths and other insects (Heide and Gotz, 1996; Balint and Dickinson, 2001; Lindsay et al, 2017; Melis et al, 2024, etc; Wood et al., 2024, etc). Therefore, it is not particularly surprising that Drosophila employs a multi-muscle strategy for roll stabilization. Although it is interesting to see that inhibition of even the phasic muscles (b3/ I2) alone did not disrupt the recovery, further experiments are required to address how these muscles contribute to wing kinematics responsible for roll control. Unfortunately, although the new evidence provided in this manuscript strengthens the idea of redundant muscle control, it does not test it directly.

    3. Reviewer #2 (Public review):

      Summary:

      This manuscript investigates the kinematics, aerodynamics, and neural control of free-flight roll perturbation in fruit flies.

      Strengths:

      The paper employs a variety of appropriate methods, including magnetically sourced in-flight perturbations, free-flight wing kinematic measurement, and optogenetic silencing of specific motor units (and thus steering muscles). The results are generally consistent with prior work, showing that 5 different bilateral pairs of steering muscles contribute to the roll response, affecting the wing stroke amplitude and wing pitch. Furthermore, the roll response - both the overall animal performance and the details of the wing motion - is not detectably altered by knocking out any one of the five muscle pairs.

      The reverse approach, optogenetic activation of specific phasic muscle pairs or silencing of tonic pairs, confirms that the selected muscles produce changes to wing kinematics appropriate for a roll response (confirmed by quasi-steady aerodynamic modeling). This set of results corroborates the main conclusions.

      Weaknesses:

      The authors refer to this as robust control of roll, though exactly what is meant by this is not clearly defined, and the word "detectably" may be important to understanding the limitations of the findings, since the large amount of variability in many of the experimental measurements would make it challenging to detect differences among treatments. As with the silencing experiments, the wing kinematics after optogenetic activation were highly varied, making it challenging to identify differences between the effects of individual muscles.

    4. Reviewer #3 (Public review):

      Summary:

      Ludlow et al. investigate the control strategy that flies use to stabilize flight during small roll perturbations. Using 3D kinematic analysis of freely flying Drosophila, Ludlow et al. ask how manipulating steering motor neurons alters the fast stabilization reflex, which spans only a few wingbeats. Bilaterally activating i1 or i2 wing steering motor neurons during free flight pitches the fly down via decreases in the wing stroke amplitude, whereas inhibiting the b3 wing steering motor neurons pitches the fly up via increases in the wing stroke amplitude. These results suggest that asymmetric recruitment of these steering muscles may be used to rotate the fly around the roll axis. Then, using quasi-steady aerodynamic modeling, they linearly interpolate changes in six kinematic features of wing strokes to describe which wing parameters produce the greatest corrective roll torque. They repeated this analysis with data from optogenetic activation of wing steering motor neurons. They argue that modeling of these optogenetic perturbations supports the hypothesis that i1, i2, and b3 muscles contribute to rotation around the roll axis by calculating roll torque changes from changing kinematics of a single wing, despite bilateral optogenetic activation. Finally, they support their claims about the redundancy of steering motor neurons by presenting connectomic analyses of the direct pathways from haltere sensory neurons to wing steering motor neurons. This analysis reveals two independent pathways from halteres to two wing muscle groups (b1 and b2 vs. i1, i2, and b3). Taken together, these data provide some evidence for redundant roll stabilization control strategies in Drosophila.

      Strengths:

      The central strength of this work is the high-quality free-flight kinematic dataset, which provides a detailed picture of how wing-stroke parameters change during natural roll perturbations and correction. The integration of quasi-steady aerodynamic modeling with these kinematic data offers a principled framework for linking muscle activity to torque generation. The use of split-Gal4 lines to target individual wing steering motor neurons with cell-type specificity provides a potentially precise approach to probe the contribution of specific muscles to roll torque. Together, these tools position this study to make a meaningful contribution to understanding the sensorimotor control strategies underlying flight stabilization in Drosophila.

      Weaknesses:

      The GtACR1 silencing experiments lack validation that the optogenetic manipulation actually suppresses motor neuron activity. Without a positive control to calibrate light intensity and duration, the absence of kinematic effects cannot be interpreted with confidence. The confocal images of the driver lines provided are insufficient to uniquely identify the targeted motor neurons and do not clearly show expression in the brain and nerve cord. The kinematic modeling relies on flight profiles derived from a single fly and a single trial, raising questions about whether they capture the full range of natural variation. Finally, the connectomics analysis largely recapitulates prior work and provides little new insight.

      There is no evidence that optogenetic silencing of wing motor neurons with GtACR1 is actually suppressing their activity. There are many factors that could impact the efficacy of this manipulation: transgene expression, light intensity and duration, etc. While electrophysiology experiments would be ideal, this would be challenging. Another option would be to use a positive control with an obvious phenotype to calibrate the light intensity and duration. For example, silencing all motor neurons (e.g., using OK371-Gal4 or another driver labeling glutamatergic neurons) should essentially paralyze the fly. Without additional evidence, the lack of a kinematic effect in the GtACR experiments is not convincing.

      The confocal images in Figure 1C are of poor quality and do not help identify the motor neurons. They only point to cell body locations, and the morphologies of the neurons are not clear. The glial sheath also appears to be labeled. The images as they currently exist are not sufficient to uniquely identify the wing motor neurons and should be updated, including brains, since the experiments manipulate activity across the nervous system. It should also be clarified that these split Gal4 lines were not generated in this work.

      The connectomics analysis does not add much on top of what was already known. It doesn't necessarily need to be removed, and the authors acknowledge that it is basically repeating prior analyses in prior publications from another connectome dataset. But it could be reduced to a schematic summarizing prior work. The one thing that should be clarified is what the "haltere afferents" actually are. Campaniform sensilla only or other sensory neurons (e.g., haltere chordotonal neurons)?

      In Figures 3 and 4, 50 kinematic profiles are created using data from a single fly and a single trial. Are these kinematic profiles representative of the range that flies use? Would the results generalize across different bouts? Figure 1i-m shows a much narrower distribution of kinematic features compared to Figure 3, which might reduce the resulting torque calculations. Why not sample from the distribution of data in 1i-m in Figure 3? Could there be some sort of bootstrapping for the modeling in Figures 3-4?

    1. eLife Assessment

      This important study provides solid evidence that, in an α-synucleinoathy model, brain atrophy and associated functional decline generalize across host genotype and fibril species, while being strongly shaped by the site of seeding and regional vulnerability. With an unusually large longitudinal in vivo MRI dataset combined with behavioural and computational modelling, the study offers a comprehensive characterization of neurodegeneration across multiple biological and experimental conditions. The openly shared imaging dataset and analysis pipeline are an excellent resource, and the scale, longitudinal design, methodological breadth, and openness of the study make it a substantial contribution to the field. This work will be of interest to scientists in synucleinopathy- and network- neurodegeneration communities.

    2. Reviewer #1 (Public review):

      Summary:

      Tullo et al. address an important and currently unresolved mechanistic question: does the prion-like spreading of alpha-synuclein (aSyn) generalize across three biological factors: host genotype, preformed-fibril (PFF) species, and disease epicentre (brain region); and can the resulting neurodegeneration be predicted computationally? Using a longitudinal design, adult M83 A53T-hemizygous mice and wild-type littermates received intrastriatal human- or mouse-PFF or PBS, with in vivo brain MRI at 7T, motor testing, survival and weight followed to 120 days post-injection. A parallel experiment seeded human-PFF or PBS into the hippocampal dentate gyrus. Atrophy was quantified by deformation-based morphometry, brain-behaviour coupling by partial least squares, and spread was simulated with a Susceptible-Infected-Removed agent-based model constrained by the Allen mouse connectome and SNCA expression. The authors conclude that aSyn-associated atrophy generalizes across genotype and fibril species but is anatomically distinct for the two epicentres, emphasizing regional vulnerability. This is a technically strong and ambitious study, reflecting a substantial and well-executed research effort.

      Strengths:

      This is a technically accomplished and ambitious study from a group with clear expertise in mouse neuroimaging and network modelling. The study addresses a real knowledge gap, relevant for mechanistic explorations of alpha-synucleinopathies: genotype and fibril inoculum species have rarely been compared head-to-head, and the relationship between aSyn propagation and downstream atrophy outside the striatum has been under-examined so far. The central hypothesis, that regional vulnerability constrains aSyn-associated neurodegeneration, together with the first attempt to model aSyn-induced atrophy computationally in rodents, is conceptually and methodologically valuable and translationally relevant.

      The longitudinal dataset is unusually rich (687 in vivo scans), and both the data and the analysis pipeline are openly shared (OpenNeuro ds007671, Zenodo, GitHub), which is extremely important for reproducibility and of clear value to the community. The work uses a multi-modal methodology in which the same phenomenon is examined across anatomical MRI, multivariate brain-behaviour modelling, and a mechanistic simulation, and is the first to model aSyn-induced atrophy computationally in mice.

      Another strength is the consistent incorporation of sex as a biological variable throughout, including sex-stratified survival, behavioural, and voxelwise atrophy analyses. This allowed the authors to identify sex differences in disease progression, notably in survival and symptom onset, while indicating that the core PFF-induced atrophy pattern was largely preserved across sexes.

      The core descriptive findings, that PFF-induced atrophy and motor impairment are reproducible across genotype and fibril species, and that striatal and hippocampal seeding yield distinct anatomical signatures, are convincingly supported and represent a very valuable advance.

      Weaknesses:

      The strongest mechanistic and epicentre interpretations would benefit from some additional support, although the study's core findings are robust.

      First, the framing centres on aSyn propagation, but the sole in-cohort readout is MRI-derived atrophy assessment; no aSyn/phospho-Ser129 pathology is shown for these animals. The atrophy-propagation link remains inferential.

      Second, the computational model's fit is reported as the peak correlation across simulation time steps and is not yet benchmarked against null or baseline models, so it is somewhat difficult to determine how much the connectome and dynamics contribute beyond gene expression alone; parameter provenance is also not described in the text.

      Third, the epicentre difference is well supported empirically at matched inoculum, but the computational comparison (SIR) is so far inoculum-mismatched: the striatal model was evaluated against mouse-PFF atrophy while the hippocampal model used human-PFF. A matched striatal human-PFF map is already available, so this could be reconciled without new data. The reduced hippocampal vulnerability despite higher hippocampal SNCA expression also remains unexplained.

      Appraisal and impact:

      The authors largely achieve their aims, and the generalization of atrophy across genotype and fibril species, together with the epicentre-specific anatomy, is well supported by a strong and openly available dataset. The more mechanistic conclusions - aSyn propagation specifically, connectome-driven vulnerability, and epicentre-determined resistance - would be strengthened where feasible by pathology validation, model benchmarking, and completing the already-available matched computational comparison, and should be interpreted with corresponding caution. Even so, the combination of a large longitudinal imaging resource, a factorial in vivo design, and the first rodent computational model of aSyn-related atrophy makes this a valuable contribution that is likely to be a useful reference and methodological template for the synucleinopathy and network-neurodegeneration communities.

    3. Reviewer #2 (Public review):

      Summary:

      This study explores risk factors for neural atrophy following alpha-synuclein injection from two complementary perspectives. First, it evaluates the effect of biological and experimental factors (genotype, alpha-synuclein species, biological sex, seeded brain region and time since injection) on the extent of neural atrophy. Second, it assesses whether regional biological features (gene expression and structural connectivity) can predict the spatial distribution of that atrophy. Using longitudinal in vivo MRI, the authors map brain volume changes over time. They relate the brain changes from striatum seeding to behavioral outcome, identifying factors associated with more severe pathology. Finally, the authors validate a previously developed in silico model for predicting brain atrophy from alpha-synuclein seeding. The model is based on the alpha-synuclein prion-like spreading hypothesis and uses local gene expression and structural connectivity to predict atrophy following the injection. They conclude that the model accurately predicts atrophy following striatal seeding but performs poorly for hippocampal seeding. They further show that structural connectivity alone is insufficient to explain the observed atrophy after striatal seeding, and that incorporating regional gene expression substantially improves model performance.

      Strengths:

      The authors have expanded on their previous work by systematically evaluating how multiple biological and experimental variables influence the development of brain atrophy. The use of MRI to map structural changes and the subsequent analysis is well validated by this group and enables comprehensive whole-brain quantification across a large number of experimental conditions. The evaluation of the in silico model linking regional gene expression and structural connectivity to patterns of atrophy under different experimental conditions is important for expanding our understanding of how atrophy develops in synucleinopathies.

      Weaknesses:

      My principal concern is that the manuscript is framed as an investigation of alpha-synuclein propagation, whereas the primary outcome measured throughout the study is a change in regional brain volume. Although atrophy is likely related to the underlying spread of pathological alpha-synuclein, the spatial distribution of alpha-synuclein pathology is not directly quantified. Conclusions regarding propagation of alpha-synuclein and the relationship with tissue loss are inferred from the performance of the in silico model in predicting atrophy. I think the manuscript could be revised to make this distinction clearer.

      A second concern relates to the comparison between striatal and hippocampal seeding. A key conclusion of the manuscript is that the in silico model accurately predicts atrophy following striatal seeding but not hippocampal seeding. However, the two analyses use different experimental group comparisons (striatum: M83 Ms-PFF versus WT PBS; hippocampus: M83 Hu-PFF versus M83 PBS). It would be helpful to demonstrate that the observed difference in model performance is not attributable to these differing experimental/ control groups.

    4. Reviewer #3 (Public review):

      Summary:

      This work studied the prion-like α-synuclein spreading hypothesis from the view of different host genotypes (M83 transgenic vs wild-type), fibril species (mouse vs human PFFs), and disease epicenter (striatum vs hippocampus). Major results include tracking neurodegeneration longitudinally with in vivo MRI, behavior, and survival in the same mice. Furthermore, this work sought to link atrophy patterns to structural connectivity and regional SNCA expression. Finally, the authors tested whether a connectome-based SIR spreading model could predict the atrophy in silico and generalize across seed sites.

      Strengths:

      (1) Same mice imaged repeatedly across four timepoints (−7, 30, 90, 120 dpi), giving true within-subject volumetric trajectories rather than cross-sectional snapshots.

      (2) Investigate the atrophy pattern for striatal-vs-hippocampal seeding in PD.

      Weaknesses:

      (1) The hypothesis (regional vulnerability) is not novel, although the manuscript presents compelling and interesting results supporting it in Figures 2 and 3.

      (2) The findings primarily establish statistical associations rather than causal mechanisms. This limitation appears inherent to the cross-cohort dataset utilized, which the authors should explicitly address in the discussion.

      (3) The descriptions of the statistical analyses in Sections 2.5 and 2.6 lack sufficient detail. The authors should provide additional technical specifics to ensure reproducibility.

      (4) Given that VBM was used to determine atrophy patterns, it is necessary to address how the multiple comparisons problem was handled in the statistical analysis to control for false positives.

    1. eLife Assessment

      This is a valuable study that contributes to our understanding of transcriptomic responses in microglia to HIV infection in the human brain. The evidence provided remains incomplete, and further analyses are required (i.e., a bigger sample size) to eliminate confounding factors so that the study can unequivocally ascertain the persistent inflammatory state of microglia in HIV-suppressed brains.

    2. Reviewer #1 (Public review):

      Summary:

      HIV can persist in brain microglia despite ART and is linked to ongoing neuroinflammation and altered cellular function.

      Strengths:

      The authors demonstrate an innovative cell-type-specific analysis of human postmortem brain tissue from aviremic and viremic people with HIV. It uses FANS, bulk and single-nucleus RNA-seq to show that HIV DNA is concentrated mainly in microglia and that inflammatory and synaptic abnormalities persist despite ART.

      Weaknesses:

      The evidence is exploratory, based on a small, heterogeneous postmortem cohort. Therefore, the findings are suggestive rather than definitive.

      The study would be stronger if a larger cohort, particularly more HIV-negative controls, were included. Also, less heterogeneity would reduce confounding from co-infections and terminal illness. The findings would also be better supported by longitudinal or matched peripheral data, protein-level validation, and direct evidence of viral activity rather than proviral DNA alone. A larger sample size would improve statistical power and make the cell-type differences more reliable.

    3. Reviewer #2 (Public review):

      Summary:

      The authors use FANS of rapidly obtained postmortem brain tissue from DPWH, seven aviremic, four viremic and three HIV-negative controls to characterize the CNS HIV reservoir and cell-type-specific transcriptional changes.

      Strengths:

      The study addresses a genuinely important and understudied question: the effect of viral suppression specifically on the CNS reservoir and transcriptome using a rare and well-characterized specimen set.

      Weaknesses:

      I have some reservations about the conclusions, because of the confounders, mechanistic narrative, and the data itself.

      (1) With n = 3 negative, n = 4 viremic, and n = 6 aviremic (post-H5 exclusion), every DEG and enrichment result rests on very few individuals. Rather than HCA reporting effect-size distributions and per-gene sample support, the authors should consider sensitivity/leave-one-out analyses to show that results are not driven by single donors. To me, it is as in Figure 3: major changes in the DGE are between the viremic vs aviremic, interestingly not with the negative control.

      (2) HIV-negative controls were significantly older (74,76 & 83, inflammaging) and entirely male (sex-based immune differences). Both bias the immune comparisons that anchor the paper. PCA reassurance with n = 3 is weak. The authors should address this quantitatively, e.g., age/sex as model covariates, or explicit discussion of directionality of bias for each key pathway.

      (3) HIV DNA was detectable in only 5/11 DPWH, and the microglial reservoir signal comes from ~3 individuals. The 10³-10⁴ copies/million figure and "dominant reservoir" claim should be framed against this limited detection and the focal distribution of infection.

      (4) Only two participants had documented cognitive symptoms, and histopathology showed no neuropathology in anyone. The transcriptome-to-HAND link is currently asserted rather than demonstrated. The authors should state this limitation prominently and avoid implying an established relationship.

      (5) A large fraction of DPWH had TB (one TBM), and controls had SARS-CoV-2. TBM alone causes microglial activation. Excluding H5 does not remove the broader TB signal. The authors should analyze/discuss TB status as a potential driver of the microglial immune signature in the retained cohort.

      (6) Sorting on IRF5 cannot distinguish microglia from perivascular macrophages, as correctly stated by the authors in the discussion, so the "microglial" reservoir may include other myeloid populations. The authors should change the cell-type attribution accordingly.

  2. Aug 2026
    1. eLife Assessment

      This valuable study characterizes the population structure and ancestry of Helicobacter pylori in Cabo Verde. Using a carefully sampled population-based cohort and paired host-bacterial genomic data, the authors identify distinct local lineages but with limited concordance between human and bacterial ancestry. The evidence is solid for the principal findings concerning bacterial diversity, ancestry and recent lineage expansion, while the support for the claim that this reflects host adaptation, enhanced transmission, reduced virulence or a mix thereof is weaker. The work has particular relevance for researchers interested in microbial population genomics, host-pathogen coevolution and the genetic consequences of human migration and admixture.

    2. Reviewer #1 (Public review):

      Summary:

      The authors are trying to characterize the sources of H. pylori in an island system. They find that, like the humans, the bacteria are admixed, but there is no correlation within the island between human ancestry and bacterial ancestry.

      Strengths:

      The study has taken particular care to characterize the humans from which isolates were obtained. Thus, it is a particularly convincing demonstration of a bacterial "melting pot".

      Weaknesses:

      The GWAS is highly confounded by population structure. In particular, there is a large group of strains that lack the cag pathogenicity island, and also differ in frequencies of other genes. So it's not clear that these differences are other than in cag status.

      Figure 1 seems very inconclusive.

    3. Reviewer #2 (Public review):

      Summary:

      This study investigates the population structure and ancestry of Helicobacter pylori in Cabo Verde, where the human population has mixed West African and European ancestry. The authors combine a population survey, serum markers, bacterial genome analysis, and paired human-bacterial ancestry data. They report high H. pylori seropositivity, several distinct bacterial groups, and limited correlation between human ancestry and bacterial ancestry. They also identify one European-derived bacterial group that appears to have undergone a recent expansion and carries fewer well-known virulence-related genes.

      The study is interesting, and the dataset is valuable, especially because population-based H. pylori genomic data from Cabo Verde and West Africa are limited. The results provide useful information on bacterial diversity, historical migration, and host-bacterial ancestry. However, some of the main conclusions are stronger than the evidence currently supports, particularly the claims of host adaptation, increased transmission, and reduced virulence.

      Strengths:

      (1) A major strength is the study population. Participants were recruited from the general population and not only from patients with gastrointestinal disease. This gives a broader view of H. pylori diversity in Cabo Verde than studies based only on hospital patients.

      (2) The number of participants tested for H. pylori antibodies is substantial, and the authors also obtained a relatively large number of bacterial genomes. The combination of human and bacterial genomic information is another important strength. This allows the authors to directly examine whether human ancestry is related to the ancestry of the colonising bacteria.

      (3) The population genetic analyses are extensive. The authors use several different approaches, and these generally support the existence of African-derived and European-derived bacterial groups in Cabo Verde. The identification of two low-diversity European-derived groups is also interesting and suggests a relatively recent expansion.

      (4) The addition of new strains from Ghana and Portugal improves the reference dataset. The results may help future studies of H. pylori population structure in Africa, Europe, Cabo Verde, and populations affected by historical Atlantic migration.

      (5) The finding that human ancestry and bacterial ancestry are only weakly related in this population is potentially important. It suggests that the long-term relationship between human and bacterial ancestry may be less stable in recently admixed populations.

      Weaknesses:

      The main weakness is that several biological conclusions are based on indirect evidence. The genomic results support recent expansion of one bacterial group, but they do not directly show that this expansion was caused by adaptation to local hosts or by increased transmission. Founder effects, population history, geographic clustering, household transmission, or random expansion may also explain the pattern. The wording should therefore be more cautious.

      The conclusion of reduced virulence is also not fully supported. The expanded lineage often lacks the cag pathogenicity island and carries less virulent forms of vacA, which suggests lower virulence potential. However, this does not prove that the strains cause less gastric damage or lower disease risk. There are no endoscopic or histological data, and serum pepsinogen values are only indirect markers.

      The description of the study population as having limited gastric inflammation is too strong. Serum pepsinogen measurements are useful for estimating gastric atrophy, but they do not directly measure the degree of histological gastritis. In addition, participants were recruited independently of symptoms, but this does not mean that they were all asymptomatic.

      The epidemiological estimate is based on antibody testing. This measures seropositivity and cannot clearly distinguish current from previous infection. Therefore, terms such as active infection or colonisation should be used carefully.

      The proposed new West-Central African bacterial group is based on a small number of reference strains from Ghana and Nigeria. The result is interesting, but broader sampling from African countries is needed before this group can be considered firmly established.

      The interpretation related to the trans-Atlantic slave trade is plausible, but the data mainly show patterns consistent with known historical migration. They do not directly demonstrate when or how the bacterial lineages moved.

      The gastric cancer comparison may also be affected by bacterial population structure. Differences between the Cabo Verdean lineage and gastric cancer strains may reflect ancestry or lineage differences rather than disease association alone.

      Overall, the study achieves its main aim of describing H. pylori diversity and ancestry in Cabo Verde. The evidence is strong for the population structure and ancestry findings, but less strong for the proposed mechanisms of adaptation, transmission, and reduced disease-causing potential. The work will be useful to the field, but the main conclusions should be stated more carefully.

    4. Author response:

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The authors are trying to characterize the sources of H. pylori in an island system. They find that, like the humans, the bacteria are admixed, but there is no correlation within the island between human ancestry and bacterial ancestry.

      We thank the reviewer for their consideration of our work.

      Strengths:

      The study has taken particular care to characterize the humans from which isolates were obtained. Thus, it is a particularly convincing demonstration of a bacterial "melting pot".

      Weaknesses:

      The GWAS is highly confounded by population structure. In particular, there is a large group of strains that lack the cag pathogenicity island, and also differ in frequencies of other genes. So it's not clear that these differences are other than in cag status.

      We ran the mentioned GWAS using the Cabo Verdean European lineage as controls group and the European gastric cancer lineage as cases to investigate the genomic differentiation between the Cabo Verdean European lineages and gastric cancer-associated lineages. Our objective was not to replicate previously reported associations or to identify novel H. pylori loci associated with gastric cancer. Therefore, we will revise the manuscript to make our objective clearer and remove statements that suggest direct associations with gastric cancer. In addition, because we employed a linear mixed model that accounts for population structure through the estimation of a genome-wide kinship matrix, we will include the genomic inflation factor. These measures will provide readers with a clearer assessment of the extent of population structure influencing the association analysis.

      Figure 1 seems very inconclusive.

      We agree that Figure 1 shows overlap between H. pylori-seropositive and seronegative distributions; in addition, its visual interpretation might be also influenced by a number of outlying observations. Nevertheless, the figure illustrates differences in pepsinogen I and II concentrations and pepsinogen I/II ratios between H. pylori-seropositive and seronegative individuals. To improve clarity,  we will include statistics in the plot, and we will revise the text to provide a more complete description of the key findings shown in Figure 1. If the reviewers and editor consider that these descriptive data are not central to the main conclusions of the study, we would be happy to move Figure 1 to the Supplementary Material, where it can provide supporting context without detracting from the presentation of the principal results.

      Reviewer #2 (Public review):

      Summary:

      This study investigates the population structure and ancestry of Helicobacter pylori in Cabo Verde, where the human population has mixed West African and European ancestry. The authors combine a population survey, serum markers, bacterial genome analysis, and paired human-bacterial ancestry data. They report high H. pylori seropositivity, several distinct bacterial groups, and limited correlation between human ancestry and bacterial ancestry. They also identify one European-derived bacterial group that appears to have undergone a recent expansion and carries fewer well-known virulence-related genes.

      The study is interesting, and the dataset is valuable, especially because population-based H. pylori genomic data from Cabo Verde and West Africa are limited. The results provide useful information on bacterial diversity, historical migration, and host-bacterial ancestry. However, some of the main conclusions are stronger than the evidence currently supports, particularly the claims of host adaptation, increased transmission, and reduced virulence.

      We thank the reviewer for their comments and careful consideration of our work.

      Strengths:

      (1) A major strength is the study population. Participants were recruited from the general population and not only from patients with gastrointestinal disease. This gives a broader view of H. pylori diversity in Cabo Verde than studies based only on hospital patients.

      (2) The number of participants tested for H. pylori antibodies is substantial, and the authors also obtained a relatively large number of bacterial genomes. The combination of human and bacterial genomic information is another important strength. This allows the authors to directly examine whether human ancestry is related to the ancestry of the colonising bacteria.

      (3) The population genetic analyses are extensive. The authors use several different approaches, and these generally support the existence of African-derived and European-derived bacterial groups in Cabo Verde. The identification of two low-diversity European-derived groups is also interesting and suggests a relatively recent expansion.

      (4) The addition of new strains from Ghana and Portugal improves the reference dataset. The results may help future studies of H. pylori population structure in Africa, Europe, Cabo Verde, and populations affected by historical Atlantic migration.

      (5) The finding that human ancestry and bacterial ancestry are only weakly related in this population is potentially important. It suggests that the long-term relationship between human and bacterial ancestry may be less stable in recently admixed populations.

      Weaknesses:

      The main weakness is that several biological conclusions are based on indirect evidence. The genomic results support recent expansion of one bacterial group, but they do not directly show that this expansion was caused by adaptation to local hosts or by increased transmission. Founder effects, population history, geographic clustering, household transmission, or random expansion may also explain the pattern. The wording should therefore be more cautious.

      We thank the reviewer for underlining the limitations of inferring the causes of lineage expansion from indirect genomic patterns alone. We will review the manuscript to more carefully reflect the uncertainty surrounding the drivers of this lineage expansion. We will also expand the Discussion to outline analytical approaches that may help distinguish between demographic and selective scenarios in a highly recombining species (where high levels of recombination may have obscured local signatures of selection). Where feasible, we will explore additional analyses of the genomic data to assess the extent to which the observed patterns are consistent with these alternative hypotheses.

      The conclusion of reduced virulence is also not fully supported. The expanded lineage often lacks the cag pathogenicity island and carries less virulent forms of vacA, which suggests lower virulence potential. However, this does not prove that the strains cause less gastric damage or lower disease risk. There are no endoscopic or histological data, and serum pepsinogen values are only indirect markers.

      We will review our manuscript to adopt a more cautious description of these strains. However, we emphasize that, as show in  in Figure 4 – source data3, we could not find strains belonging to this expanded lineage carrying an active cagPAI,  or a virulent vacA allele. Our GWAs analyses show high divergence between these strains and gastric cancer strains both at a genome-wide level, and at previously identified virulence genes (as sabA and BabA, in addition to the ones already mentioned above). Finally, serum pepsinogen serum analysis, although indirect, have been shown to compare with  to the “gold standard” method, histopathological biopsy microscopy (ex: Telaranta-Keerie et al 2010, 10.3109/00365521.2010.487918; Kitamura et al 2015, 10.1111/jgh.12987; Miftahussurur et al, 2020, 10.1371/journal.pone.0230064; please see more about this in our following comment). Taken together these results provide minimal support for the hypothesis that the expansion of this lineage is driven by increased virulence. We will review the manuscript in order to reflect this.

      The description of the study population as having limited gastric inflammation is too strong. Serum pepsinogen measurements are useful for estimating gastric atrophy, but they do not directly measure the degree of histological gastritis. In addition, participants were recruited independently of symptoms, but this does not mean that they were all asymptomatic.

      The epidemiological estimate is based on antibody testing. This measures seropositivity and cannot clearly distinguish current from previous infection. Therefore, terms such as active infection or colonisation should be used carefully.

      As mentioned above, serum pepsinogen serum analysis has been widely shown to reliably detect both chronic and atrophic gastritis (e.g. Miftahussurur et al, 2020, 10.1371/journal.pone.0230064). We adopted conservative cut-offs for serum pepsinogen values after a careful review of the literature in our analysis. However, we acknowledge that these cut-offs may vary between populations. Therefore, we agree with the reviewer that a more precise assessment would have included a sensitivity and specificity analysis of  serum pepsinogen measurements against histopathological biopsy microscopy in a subset of the individuals. Taking this is into consideration, we will review the manuscript to include this discussion and to adopt a more cautious description of these results. Although we acknowledge that seropositive does not distinguish active from past infection in the Discussion, we will bring this discussion into the Results section as well.

      The proposed new West-Central African bacterial group is based on a small number of reference strains from Ghana and Nigeria. The result is interesting, but broader sampling from African countries is needed before this group can be considered firmly established.

      We thank the reviewer for pointing the need to assess higher African diversity, which is an issue that we also mention in the Discussion. Although the new West-Central African group comprised fewer isolates than the other comparison populations, the clustering pattern is unlikely to be explained solely by sample size because the analysis was based on a normalised chromosome-painting coancestry matrix, which reduces the influence of unequal donor numbers. In our review, we will provide additional analyses to confirm the observed clustering pattern: 1) we will provide fineSTRUCTURE MCMC tree assignments; 2) we will repeat the Chromopainter/ fineSTRUCTURE analyses under random downsampling of the larger groups; 3) as suggested in further reviewer recommendations, we will repeat the Chromopainter/ fineSTRUCTURE analyses after removing the new Ghanaian sequenced strains.

      The interpretation related to the trans-Atlantic slave trade is plausible, but the data mainly show patterns consistent with known historical migration. They do not directly demonstrate when or how the bacterial lineages moved.

      We thank the reviewer for highlighting this point. The chromosome-painting analysis identifies shared ancestry and gene flow between populations but does not directly estimate the timing of those events. However, several lines of evidence suggest that the majority of the strains have been present in Cabo Verde for an extended period rather than representing recent introductions. First, the isolates were obtained from individuals who, as well as both of their parents, were born in Cabo Verde. Second, Cabo Verdean African strains show an excess of European ancestry relative to their putative parental African populations. And vice-versa: Cabo Verdean European strains exhibit an excess of African ancestry compared with their putative parental European populations. To further investigate this question, as also suggested in reviewer recommendations, we will explore the feasibility of identifying clonal or near-clonal relationships between Cabo Verdean and assess whether dating approaches such as BactDating can provide additional insights into the timescale over which these lineages have diversified and admixed within Cabo Verde.

      The gastric cancer comparison may also be affected by bacterial population structure. Differences between the Cabo Verdean lineage and gastric cancer strains may reflect ancestry or lineage differences rather than disease association alone.

      We thank the reviewer for highlighting this point. As outlined in our response to Reviewer 1, the primary objective of this analysis was to investigate genomic differentiation between the Cabo Verdean and gastric cancer-associated lineages. Therefore, the strong contribution of ancestry and lineage effects is central to the interpretation of this comparison. We will make this clearer in our final manuscript.

      Overall, the study achieves its main aim of describing H. pylori diversity and ancestry in Cabo Verde. The evidence is strong for the population structure and ancestry findings, but less strong for the proposed mechanisms of adaptation, transmission, and reduced disease-causing potential. The work will be useful to the field, but the main conclusions should be stated more carefully.

      We thank  the reviewers and the editor for their time and effort in assessing the manuscript. We will submit a revised version that carefully addresses these points and incorporates the suggested changes.

    1. eLife Assessment

      In this fundamental study, the authors argue that the most widely used behavioral measure of the ability to stop an action is largely determined by the time it takes for visual signals to reach the brain and for movements to be executed, rather than by the control processes it is assumed to index. The evidence is convincing, combining reanalyses of seven existing datasets with a preregistered new experiment, and its force comes from the authors doing three things at once: identifying the problem, showing that it is present wherever it can be measured, and offering remedies applicable to both archived and future data. Various concerns were raised as to the interpretation of the data and their robustness to experimental manipulations, but if these were properly addressed, the implications of the work could extend across psychology, psychiatry and neuroscience.

    2. Reviewer #1 (Public review):

      This interesting manuscript challenges the current interpretation of the well-established stop signal reaction time task (SSRT), commonly used in many research areas. SSRT has been traditionally thought of as primarily a measure of inhibitory control. In this work, the authors argue that this is influenced significantly by sensory and motor transmission times, and that these low-level processes may systematically confound SSRT estimates both in individual groups and also in many clinical populations.

      Conceptually, this raises an important and significant question regarding the construct validity of one of the more widely used behavioral measures of response inhibition. The authors provide a clear theoretical framing and importantly address the overlooked assumptions in SSRT modelling, the underexplored source of variability.

      Further, direct evidence separating peripheral sensory and motor contributions from central inhibitory processes is certainly needed, as well as a more balanced interpretation of prior methodological refinements in the field. Is a correlation between T0 and SSRT sufficient to conclude that SSSRT may be predominantly driven by peripheral delays? What proportion of SSRT variance remains unexplained after one accounts for T0? If T0 and inhibitory processes share neural processing speed, does that mean they covary? If one corrects T0, do the group differences become smaller or perhaps disappear?

      Furthermore, how robust is the T0 estimate in noisy environments of all sorts and especially across modalities (visual and motor)? The authors argue that this relationship is clearer in "good quality data"; how do you define that, and what happens with noisy data? For instance, the authors refer to a range of clinical disorders such as ADHD and PD where noise is abundantly present, partly due to the disease itself or due to treatment.

      In conclusion, this is certainly a thought-provoking and potentially influential contribution in the literature that raises important questions about the interpretation of the stop signal reaction time task.

    3. Reviewer #2 (Public review):

      Continuing their work distinguishing sensory latencies of "cognitive" processes, the authors turn their attention to "response inhibition". The "square quotes" are being used to highlight how this manuscript aims to challenge previous descriptions of performance data and inferred computational processes. The authors assert that previous descriptions of the measure known as "stop signal reaction time" (SSRT) are flawed because they did not account for sensory latencies empirically or theoretically.

      Enthusiasm for the manuscript cannot be high in light of many weaknesses countering the possible strengths. Strengths include offering an opportunity to more carefully characterize the quantity SSRT and a specific empirical approach offered to the research community. However, these strengths are countered by the following structural, theoretical, and empirical weaknesses:

      As announced by the elephant in the title, the writing could be described as excessively polemical. However, the characterization and interpretation of previous empirical and theoretical work is disputable.

      The major theoretical claim regarding sensory delays inherent in SSRT is not novel. The authors assert, "...this corpus of work may have been misinterpreted because the SSRT is systematically influenced by low level sensory and motor transmission times, arguably more so than by inhibition or cognitive processes." This was certainly recognized by Logan and Cowan in their original work. They wrote, "An act of control, like any other act, must take time. The theory provides methods for measuring the latency of control even when the act of control is not directly observable." (page 298) Also, "... the estimate of stop-signal reaction time includes the latency of the internal response to the stop signal and the duration of the ballistic process." (page 316-317). Moreover, subsequent computational and empirical work, some noted by the authors, has distinguished the sensory encoding interval from the interval during which the STOP process interrupts the GO process.

      The theoretical suggestion that an accounting for sensory delays undermines the functional interpretation of SSRT mischaracterizes the original literature. For example, in the Abstract the authors write "Sensory and motor contributions must be ruled out before linking SSRT results to inhibition or cognition". The original Logan and Cowan theory was about what happens at the end of SSRT, and that was described only as an "act of control", in perfectly positivist fashion. For example, Logan and Cowan wrote, "Estimates of stop-signal reaction time provide a measure of the latency of control." (page 315). Thus, the authors are misstating what was meant originally by SSRT. In addition, the authors offer no specific or formal definition to specify what they mean by "inhibition or cognitive processes".

      Confidence in the new empirical conclusions of the manuscript must be low because the new performance data are of questionable quality. The first issue is that the stopping accuracy (or inhibition functions in original terminology) shown in Figure S3 is very problematic for the interpretation of the authors' empirical work in this manuscript. There are two problems. First, these plots should span from nearly 0% to nearly 100%. It is not possible to resolve the span of each individual in the figure, but it is clear that many, if not most, in both the Manual and Saccadic data span just 20-30%. Second, the plots should span the 50% success value. It is clear that the maximum or minimum values for many participants do not reach the 50% value. These two problems indicate that many (most?) participants were not really sensitive to the stop signal.

      The second issue concerns the pattern of response times (RTs) on "ignore" trials. The authors portray performance as exemplifying a "pause-then-go" strategy. This is not uncommon, but it is not the only way participants perform. Many participants across multiple studies of selective stimulus stopping produce RTs on "Ignore" trials essentially indistinguishable from RTs on no-stop trials. The authors must acknowledge and account for such individual variability. In fact, the "T_s" value is measured by the difference in distributions of RT on no-signal and ignore trials. If these distributions are not different, then the measurement and interpretation of this quantity is questionable.

      Related, the distributions of RT on stop trials, particularly for saccade responses, are portrayed with a second mode in the schematic illustrations and clearly peaking at SSRT in Figure S1. This second mode is not observed in other saccade stop signal studies. This indicates that the participants in this study were in a peculiar mode of performance.

      Finally, given the pivotal role of measures of differences of RT distributions and the pronounced variation of stopping accuracy (Figure S3), the authors must show the distributions for all of their new participants. The authors' claim to higher resolution obliges them to reveal every step of analysis.

      In its current form, this manuscript is unlikely to change the thinking of modelers or practitioners of the stop signal task.

    4. Reviewer #3 (Public review):

      Summary:

      Statham and colleagues test an assumption underpinning a very large literature: that the stop-signal reaction time (SSRT) indexes the speed or efficacy of top-down inhibitory control. They argue instead, and support their claims with a total of eight datasets, that SSRT is substantially occupied by visuomotor deadtime (i.e., incompressible sensory and motor delays common to all visually guided responses), which varies across individuals, conditions and populations in ways that mimic effects usually attributed to inhibitory control. They propose two remedies: subtracting an independent estimate of visuomotor deadtime (T₀) from SSRT, and a new index, the selective stopping delay (ΔT), from the stimulus-selective stopping task.

      Strengths:

      The paper's principal strength is the combination of these components. That SSRT must contain peripheral delays is not itself new, as the authors point out (Boucher et al., 2007; Salinas and Stanford, 2013; Bompas et al., 2020). What is new is the quantification of the problem at scale, across seven archival datasets and a preregistered replication, together with the demonstration that T₀ can be recovered from existing stop-task data. That is important, as it provides a diagnostic that can be applied to data already collected. The authors' offer to assist others in doing so is exemplary. The supplementary analyses of trial numbers and participant pooling are very useful, and the paper provides important sanity checks, notably confirming that stop and ignore signals produce indistinguishable initial interference before pooling them.

      Weaknesses:

      The evidence for the central claim is strong but presented in a way that overstates it. Figure 2 reports 85% and 80% shared variance between SSRT and T₀, but these pool across datasets and, more critically, across response modality: manual and saccadic estimates from the same participants are plotted together with a single regression line through both. Because manual and saccadic deadtimes differ by roughly 130 ms, the resulting correlation largely reflects a between-condition difference rather than covariation among individuals. The numbers that speak to individual differences are more modest (40% for manual responses; 7% for saccades). The manual result is convincing and consequential; the saccadic result is not, and the explanation in terms of restricted range, while plausible, is offered after the fact and is directly testable by reporting the reliability of saccadic T₀ or correcting the correlation for attenuation. This limitation is arguably good news for the paper's practical message, since it implies saccadic measures are relatively protected, but the manuscript should make clear (including in the abstract) that the strong individual-differences case rests on the manual data, where motor execution delay is the main driver.

      A related point concerns interpretation rather than analysis. Since SSRT is, on the authors' own account, approximately the sum of T₀ and a decision-related component, covariation between the two is expected on structural grounds; the preregistered correlation with reaction times from separate speeded blocks mitigates this, but the finding is less surprising than its current framing implies. What would determine whether past conclusions must be revised is not whether SSRT correlates with T₀ across individuals, but whether the decision-related component tracks the independent variable in any given study. The alcohol reanalysis could be a test case for this: the authors show that alcohol raises T₀ commensurately with SSRT and conclude the effects are "consistent with these effects being fully driven by visuomotor delays," yet (unless I missed something) they do not report the corrected measure for these data, while they do so for signal contrast and response modality. Running that analysis, and stating plainly what Campbell et al. (2017) would have concluded under the proposed treatment, would be an important demonstration.

      The case for ΔT is the least developed part of the paper. ΔT is a difference between two independently estimated, individually noisy quantities, extracted by a non-trivial procedure (see also below), and no reliability estimates are reported for T₀, TS or ΔT. This would be possible based on the two-session design (and the group has prior work on the reliability of cognitive control measures). This matters because the argument that ΔT is superior rests, to some extent, on null findings: ΔT does not correlate with SSRT, with stopping accuracy, or with the differential response to stop and ignore trials. These null correlations are interpreted as freedom from confounds, but an unreliable measure would produce the same pattern, and the seven participants with implausible negative ΔT values indicate that noise is not negligible.

      In addition, the subjective correction of dip onsets ("Departure points were visually inspected and adjusted if it was deemed that the algorithm had placed them in inappropriate places"), which is critical to the paper's central measurement, should be blinded to condition or show inter-rater agreement. Since T₀ and TS are compared across conditions and ΔT is their difference, this introduces researcher degrees of freedom.

      This is particularly critical when a dip is not easy to extract. Figure 3 depicts an idealized ignore-trial distribution with a clean, deep dip. Real distributions are unlikely to look like this, and dip depth should depend on the behavioral relevance and salience of the ignored event; published work on rapid manual inhibition indicates that dips to behaviorally irrelevant events can be very shallow. Since ΔT is extractable only where the dip is resolvable, the generality of the method can be questioned. Ideally, the empirical distributions underlying every dataset analyzed should be shown to alleviate this concern.

      One uncontrolled procedural difference also deserves comment. Corrective feedback about stopping too often or stopping too rarely was given after manual blocks only; saccadic blocks received none, and fixed rather than staircased delays were used throughout. Since the manual-saccadic contrast carries much of the argument, and saccadic blocks yielded both lower stopping accuracy (36% vs 52%) and far more exclusions (8 vs 1 of 37), this asymmetry offers an alternative to the interpretation that saccades are simply harder to inhibit.

      A final point concerns the comparison between response modalities. Raw SSRT suggests that saccadic inhibition is faster than manual (174 vs 266 ms), while both proposed corrections reverse this, with SSRT−T₀ and ΔT each indicating that saccadic inhibition is slower (the latter consistently across nearly every participant). This is one of the clearest illustrations of the paper's thesis, but it is not taken up in the discussion, which returns to modality only to note that saccadic T₀ varies little (the one reference to variation across action modalities appears in the modeling section, without stating its direction). It would also benefit from a caveat. Both corrected measures subtract the same T₀, and manual and saccadic T₀ differ by roughly 130 ms, so the two do not corroborate one another independently (TS is itself longer for manual responses, and yields a shorter ΔT only once the larger manual T₀ is removed). The accuracy of the subtraction therefore matters here: if the manual regression slope of 0.75 reflects sub-additivity rather than attenuation, subtracting the full T₀ would overcorrect manual responses more than saccadic ones, and could produce the reversal on its own.

      These concerns qualify rather than undermine the contribution. The core observation is robust, the diagnostic is practical and immediately applicable, and the case that a large body of work requires re-examination is well made. If the corrected analyses are carried through on the datasets already in hand, this will be an important paper for anyone who uses the stop-signal task.

    5. Author response:

      We are grateful to the editor and reviewers for providing their time and expertise in the assessment of this article. We are glad that the overall evaluation is supportive.

      Reviewer #1 (Public review):

      This interesting manuscript challenges the current interpretation of the well-established stop signal reaction time task (SSRT), commonly used in many research areas. SSRT has been traditionally thought of as primarily a measure of inhibitory control. In this work, the authors argue that this is influenced significantly by sensory and motor transmission times, and that these low-level processes may systematically confound SSRT estimates both in individual groups and also in many clinical populations.

      Conceptually, this raises an important and significant question regarding the construct validity of one of the more widely used behavioral measures of response inhibition. The authors provide a clear theoretical framing and importantly address the overlooked assumptions in SSRT modelling, the underexplored source of variability.

      Further, direct evidence separating peripheral sensory and motor contributions from central inhibitory processes is certainly needed, as well as a more balanced interpretation of prior methodological refinements in the field. Is a correlation between T0 and SSRT sufficient to conclude that SSSRT may be predominantly driven by peripheral delays? What proportion of SSRT variance remains unexplained after one accounts for T0? If T0 and inhibitory processes share neural processing speed, does that mean they covary? If one corrects T0, do the group differences become smaller or perhaps disappear?

      We thank the reviewer for their positive assessment of our work. We agree all these questions are important. Our revision will provide repeatability estimates for all our indices, and use the repeatability of SSRT and T0 to estimate the proportion of SSRT variance that remains unexplained. We will mention that it is theoretically possible that shared neural processing speed between T0 and inhibitory processes contributes to the reported effect but, since the neural pathways are anatomically largely distinct, the available cognitive neuroscience literature suggests such contribution is unlikely to be strong enough to explain the observed relationship. The impact of correcting group differences in SSRT for T0 will entirely depend on where differences in SSRT between these groups come from. If it comes exclusively from peripheral delays, then the effect may indeed disappear. If it is central, then stronger differences may be revealed after correction. We appreciate these are key questions we would like answers to already, but identifying and reanalysing suitable datasets to answer it will be the purpose of future papers.

      Furthermore, how robust is the T0 estimate in noisy environments of all sorts and especially across modalities (visual and motor)? The authors argue that this relationship is clearer in "good quality data"; how do you define that, and what happens with noisy data? For instance, the authors refer to a range of clinical disorders such as ADHD and PD where noise is abundantly present, partly due to the disease itself or due to treatment.

      By good quality data, we mean enough trials at the optimal RT and SOA combinations, and an adequate preprocessing pipeline that excludes non-standard trials (poor fixation, pre-emptive responses, large undershoot …). Participants with increased intra-individual variance or low compliance will need more trials overall to get enough trials around divergence times for these to be accurately estimated. Supplementary figure 5 illustrates how low trial numbers lead to an overestimation of T0, which can be mitigated by pooling across participants. We expect noisy data to have a similar effect, although its impact may differ across clinical conditions based on where the additional variability comes from. We shall have more clarity on the reliability of T0 in clinical populations once relevant datasets have been reanalysed, and use this to define constraints for future data collection. Again, this will need to wait for future papers.

      In conclusion, this is certainly a thought-provoking and potentially influential contribution in the literature that raises important questions about the interpretation of the stop signal reaction time task.

      We thank the reviewer for their in-depth and thoughtful comments and suggestions

      Reviewer #2 (Public review):

      Continuing their work distinguishing sensory latencies of "cognitive" processes, the authors turn their attention to "response inhibition". The "square quotes" are being used to highlight how this manuscript aims to challenge previous descriptions of performance data and inferred computational processes. The authors assert that previous descriptions of the measure known as "stop signal reaction time" (SSRT) are flawed because they did not account for sensory latencies empirically or theoretically.

      Enthusiasm for the manuscript cannot be high in light of many weaknesses countering the possible strengths. Strengths include offering an opportunity to more carefully characterize the quantity SSRT and a specific empirical approach offered to the research community. However, these strengths are countered by the following structural, theoretical, and empirical weaknesses:

      As announced by the elephant in the title, the writing could be described as excessively polemical. However, the characterization and interpretation of previous empirical and theoretical work is disputable.

      The major theoretical claim regarding sensory delays inherent in SSRT is not novel. The authors assert, "...this corpus of work may have been misinterpreted because the SSRT is systematically influenced by low level sensory and motor transmission times, arguably more so than by inhibition or cognitive processes." This was certainly recognized by Logan and Cowan in their original work. They wrote, "An act of control, like any other act, must take time. The theory provides methods for measuring the latency of control even when the act of control is not directly observable." (page 298) Also, "... the estimate of stop-signal reaction time includes the latency of the internal response to the stop signal and the duration of the ballistic process." (page 316-317). Moreover, subsequent computational and empirical work, some noted by the authors, has distinguished the sensory encoding interval from the interval during which the STOP process interrupts the GO process.

      We agree that our manuscript should acknowledge that Logan and Cowan (1984) explicitly stated that internal and ballistic delays contribute to SSRT and thank the reviewer for highlighting the need to clarify the relationship between our work and the original Logan and Cowan framework. We will clarify that the novelty of our claim is not that peripheral delays contribute to SSRT, which is a logical necessity, but that differences in peripheral delays (across conditions or people) contribute to differences in SSRT, sometimes to a large extent. Such differences in SSRT are very widely assumed to reflect inhibitory control in the large corpus of work that followed this initial literature. This corpus has essentially ignored the message about stimulus processing and ballistic delays and their implications for individual differences or changes across conditions. Therefore, we maintain that SSRT differences may have been widely misinterpreted. We reference the articles where these implications were clearly spelled out: these empirical and modelling studies were based on a few monkeys or human participants, and therefore unable to provide the large-scale demonstration we provide here.

      The theoretical suggestion that an accounting for sensory delays undermines the functional interpretation of SSRT mischaracterizes the original literature. For example, in the Abstract the authors write "Sensory and motor contributions must be ruled out before linking SSRT results to inhibition or cognition". The original Logan and Cowan theory was about what happens at the end of SSRT, and that was described only as an "act of control", in perfectly positivist fashion. For example, Logan and Cowan wrote, "Estimates of stop-signal reaction time provide a measure of the latency of control." (page 315). Thus, the authors are misstating what was meant originally by SSRT. In addition, the authors offer no specific or formal definition to specify what they mean by "inhibition or cognitive processes".

      We thank the reviewer for pointing out that their original approach was mechanistically agnostic, which we will explicitly clarify in revision. We will remove the words “top-down” from our 4th sentence and reword the quoted sentence into "Changes in peripheral delays must be ruled out before linking changes in SSRT to inhibition or cognition". It remains the case that many hundreds of studies have since interpreted “ability to inhibit” and “latency of control” as specific to inhibition and control, and therefore have assumed that changes in SSRT directly reflect an inhibitory cognitive process. We agree that we do not currently offer a definition of what we mean by cognitive processes, except that they do not include incompressible sensory and motor delays. Based on the reviewer’s clarification, this common shortcut in the literature appears inconsistent with the spirit of the initial work, which we seek to rectify.

      Confidence in the new empirical conclusions of the manuscript must be low because the new performance data are of questionable quality. The first issue is that the stopping accuracy (or inhibition functions in original terminology) shown in Figure S3 is very problematic for the interpretation of the authors' empirical work in this manuscript. There are two problems. First, these plots should span from nearly 0% to nearly 100%. It is not possible to resolve the span of each individual in the figure, but it is clear that many, if not most, in both the Manual and Saccadic data span just 20-30%. Second, the plots should span the 50% success value. It is clear that the maximum or minimum values for many participants do not reach the 50% value. These two problems indicate that many (most?) participants were not really sensitive to the stop signal.

      We acknowledge our inhibition functions are narrow, but we do not believe this undermines the main conclusions, for several reasons. On the question of sensitivity to the stop signal, average spans for inhibition functions after participant exclusions were 37% for manual and 29% for saccades. This limited span is mainly attributable to our fixed SOA design, which was a necessary feature of a direct comparison between manual and saccadic behaviours. Figure S3 covers only 80 ms spread of SOA. The slopes are commensurate with most other studies, which cover much wider differences in SOA. If one selects the central 100 ms (where the slopes are steepest) from the figures in most previous papers, one will find stopping accuracy changes of around 30%. Therefore, sensitivity is similar.

      On the question of some functions not crossing 50%, the correlations between SSRT and T0 remain the same if we only keep those participants who crossed the 50% point (R(26)=0.64 for manual, R(11)=0.4 for saccades, same statistical significance levels). We will additionally rerun our SSRT and T0 correlation with stopping accuracy as a covariate. As stopping accuracy affects SSRT but not T0, it is unlikely to drive our results.

      The second issue concerns the pattern of response times (RTs) on "ignore" trials. The authors portray performance as exemplifying a "pause-then-go" strategy. This is not uncommon, but it is not the only way participants perform. Many participants across multiple studies of selective stimulus stopping produce RTs on "Ignore" trials essentially indistinguishable from RTs on no-stop trials. The authors must acknowledge and account for such individual variability. In fact, the "T_s" value is measured by the difference in distributions of RT on no-signal and ignore trials. If these distributions are not different, then the measurement and interpretation of this quantity is questionable.

      We agree that the issue raised by the reviewer would be important if no difference were present between the distributions. In our dataset, however, all participants showed a measurable distributional difference. We believe the difference in perspective that ‘many participants…produce RTs on ignore trials essentially indistinguishable from RTs on no-stop trials’ may be attributed to differences in the way we analyse data (RT distributions versus mean RT).

      Nearly all participants in our final sample had mean RTignore – RTgo > 10 ms (significant at the individual level), except for 2 in the manual condition (and none for saccades). Following Bisset & Logan (2014), a lack of clear mean RTignore - RTgo difference in these 2 participants might have been interpreted as reflecting a different strategy. However, all our participants showed clear dips between go and ignore RT distributions when locked on signal onset, and these two manual participants were no exception. Therefore, accounting for response probability and RT at each SOA in our distributional analysis revealed clear ignore versus go differences, masked when relying on mean RT. The lack of mean RT difference for these two participants in the manual modality had no impact on our main hypothesis testing because T0 was extracted by comparing signal-absent and signal-present trials (pooling ignore and stop), while TS was extracted by comparing ignore and stop (not ignore and go).

      In terms of interpretation and whether participants employ a pause-then-go strategy, we understand performance in the selective stopping task as reflecting a combination of automatic activation and interference, and endogenous activation and inhibition. Our interpretation is that the pause component primarily reflects automatic interference triggered by stimulus onset, although strategic factors may also contribute in some circumstances. Individual differences in mean RTignore - RTgo could reflect both automatic interference and endogenous pausing (both of which would increase the difference between ignore and go trials), as well as subsequent failing to go on ignore trials (omissions, which would decrease differences by removing longer latency responses from ignore distributions just as in stop signal distributions). Although some of this can be described as strategic, some won’t be, and we therefore refrain from inferring strategy based on mean RTignore – RTgo.

      Related, the distributions of RT on stop trials, particularly for saccade responses, are portrayed with a second mode in the schematic illustrations and clearly peaking at SSRT in Figure S1. This second mode is not observed in other saccade stop signal studies. This indicates that the participants in this study were in a peculiar mode of performance.

      The second mode indicates that participants occasionally ignore the stop signal, which is why it peaks at the same latency as the rebound for the ignore distribution. These are not unusual, in particular in selective stopping designs, but are not as easily seen on cumulative functions, which are the standard way of plotting the results in this field.

      Finally, given the pivotal role of measures of differences of RT distributions and the pronounced variation of stopping accuracy (Figure S3), the authors must show the distributions for all of their new participants. The authors' claim to higher resolution obliges them to reveal every step of analysis.

      We will save figures showing the individual distributions in the OSF folder. Note that these figures can be produced by running the code we shared, so each step is already fully transparent, but we will create tidy versions that also highlight manually corrected indices.

      In its current form, this manuscript is unlikely to change the thinking of modelers or practitioners of the stop signal task.

      We thank the reviewer for their in-depth and thoughtful comments and suggestions, so that the paper can be revised to address the concerns.

      Reviewer #3 (Public review):

      Summary:

      Statham and colleagues test an assumption underpinning a very large literature: that the stop-signal reaction time (SSRT) indexes the speed or efficacy of top-down inhibitory control. They argue instead, and support their claims with a total of eight datasets, that SSRT is substantially occupied by visuomotor deadtime (i.e., incompressible sensory and motor delays common to all visually guided responses), which varies across individuals, conditions and populations in ways that mimic effects usually attributed to inhibitory control. They propose two remedies: subtracting an independent estimate of visuomotor deadtime (T₀) from SSRT, and a new index, the selective stopping delay (ΔT), from the stimulus-selective stopping task.

      Strengths:

      The paper's principal strength is the combination of these components. That SSRT must contain peripheral delays is not itself new, as the authors point out (Boucher et al., 2007; Salinas and Stanford, 2013; Bompas et al., 2020). What is new is the quantification of the problem at scale, across seven archival datasets and a preregistered replication, together with the demonstration that T₀ can be recovered from existing stop-task data. That is important, as it provides a diagnostic that can be applied to data already collected. The authors' offer to assist others in doing so is exemplary. The supplementary analyses of trial numbers and participant pooling are very useful, and the paper provides important sanity checks, notably confirming that stop and ignore signals produce indistinguishable initial interference before pooling them.

      Weaknesses:

      The evidence for the central claim is strong but presented in a way that overstates it. Figure 2 reports 85% and 80% shared variance between SSRT and T₀, but these pool across datasets and, more critically, across response modality: manual and saccadic estimates from the same participants are plotted together with a single regression line through both. Because manual and saccadic deadtimes differ by roughly 130 ms, the resulting correlation largely reflects a between-condition difference rather than covariation among individuals. The numbers that speak to individual differences are more modest (40% for manual responses; 7% for saccades). The manual result is convincing and consequential; the saccadic result is not, and the explanation in terms of restricted range, while plausible, is offered after the fact and is directly testable by reporting the reliability of saccadic T₀ or correcting the correlation for attenuation. This limitation is arguably good news for the paper's practical message, since it implies saccadic measures are relatively protected, but the manuscript should make clear (including in the abstract) that the strong individual-differences case rests on the manual data, where motor execution delay is the main driver.

      We will add separate regression lines and R-values for manual and saccadic modalities on Fig.2A and an inset showing the variance only driven by individual differences across all archival data (i.e. z-scored per condition and datasets, R(215)=0.43, p<0.001). We will also state more explicitly that the overall correlation may not be the relevant one for researchers specifically interested in individual differences.

      While a large portion of SSRT literature is about individual differences, there are also many studies about differences between conditions, including comparing different response modalities. Therefore, it is a general question whether differences of any kind in SSRT reflect differences in inhibitory control or differences in sensory-motor delays. At a conceptual level, most users of the SSRT are intending to measure control, and have a conceptual model in which control is separate from the modality of response or the exact characteristics of stimulus delivery. Thus, it is important to point out that their measure of control is very much dependent on these things, and in fact to a much larger degree than the more subtle individual differences, group differences or conditions of interest.

      We agree that repeatability is critical for interpreting null results and will provide split-half repeatability for all our indices, including T0 and SSRT, and use these to correct their correlations. We thank the reviewer for this suggestion.

      A related point concerns interpretation rather than analysis. Since SSRT is, on the authors' own account, approximately the sum of T<sub>0</sub> and a decision-related component, covariation between the two is expected on structural grounds; the preregistered correlation with reaction times from separate speeded blocks mitigates this, but the finding is less surprising than its current framing implies. What would determine whether past conclusions must be revised is not whether SSRT correlates with T<sub>0</sub> across individuals, but whether the decision-related component tracks the independent variable in any given study. The alcohol reanalysis could be a test case for this: the authors show that alcohol raises T<sub>0</sub> commensurately with SSRT and conclude the effects are "consistent with these effects being fully driven by visuomotor delays," yet (unless I missed something) they do not report the corrected measure for these data, while they do so for signal contrast and response modality. Running that analysis, and stating plainly what Campbell et al. (2017) would have concluded under the proposed treatment, would be an important demonstration.

      We fully agree that the presence of a correlation is indeed entirely expected and obvious in our own account, as conveyed early on in the manuscript (“From Fig. 1D, it seems clear that SSRT and T<sub>0</sub> are inevitably connected”). We agree the main question is what is left for SSRT to explain. We will update our analysis of the Campbell et al. (2017) alcohol study as suggested. Future work can then focus on other “independent variables”.

      The case for ΔT is the least developed part of the paper. ΔT is a difference between two independently estimated, individually noisy quantities, extracted by a non-trivial procedure (see also below), and no reliability estimates are reported for T<sub>0</sub>, TS or ΔT. This would be possible based on the two-session design (and the group has prior work on the reliability of cognitive control measures). This matters because the argument that ΔT is superior rests, to some extent, on null findings: ΔT does not correlate with SSRT, with stopping accuracy, or with the differential response to stop and ignore trials. These null correlations are interpreted as freedom from confounds, but an unreliable measure would produce the same pattern, and the seven participants with implausible negative ΔT values indicate that noise is not negligible.

      We fully agree with all this and will update the wording surrounding the lack of correlation between ΔT and the other measures in light of its repeatability

      In addition, the subjective correction of dip onsets ("Departure points were visually inspected and adjusted if it was deemed that the algorithm had placed them in inappropriate places"), which is critical to the paper's central measurement, should be blinded to condition or show inter-rater agreement. Since T<sub>0</sub> and TS are compared across conditions and ΔT is their difference, this introduces researcher degrees of freedom.

      All divergence times (T0, T0,stop, T0,ignore and TS) were confirmed by two of the authors. Each index for each modality is plotted on a separate figure (showing all individuals). It is technically easy to compare, say, T0 and TS, for one individual, but we refrained from doing this (and indeed ended up with many T0 > TS). We agree that blinding and inter-rater reliability are important safeguards and that our current analysis fell short of this. We will explore whether this can be done retrospectively, and report on this exercise alongside guidance on criteria used for manual corrections.

      This is particularly critical when a dip is not easy to extract. Figure 3 depicts an idealized ignore-trial distribution with a clean, deep dip. Real distributions are unlikely to look like this, and dip depth should depend on the behavioral relevance and salience of the ignored event; published work on rapid manual inhibition indicates that dips to behaviorally irrelevant events can be very shallow. Since ΔT is extractable only where the dip is resolvable, the generality of the method can be questioned. Ideally, the empirical distributions underlying every dataset analyzed should be shown to alleviate this concern.

      We will make figures available in the OSF folder with individual distributions that supported the extraction of each index, flagging those that got manually corrected. This will make apparent that the vast majority of dips were very clear, for both T0 and TS. Unclear dips led to missing indices, and were therefore excluded from our hypothesis testing.

      One uncontrolled procedural difference also deserves comment. Corrective feedback about stopping too often or stopping too rarely was given after manual blocks only; saccadic blocks received none, and fixed rather than staircased delays were used throughout. Since the manual-saccadic contrast carries much of the argument, and saccadic blocks yielded both lower stopping accuracy (36% vs 52%) and far more exclusions (8 vs 1 of 37), this asymmetry offers an alternative to the interpretation that saccades are simply harder to inhibit.

      Indeed, blockwise feedback would have been hard to implement reliably for saccades. As manual and saccadic blocks were interleaved, our hope was that participants could use the feedback received for manual to adjust their strategy for both modalities. We will check how often the feedback was triggered for manual blocks and, if more than negligible, we will note the reviewer’s suggestion as a possible driver for modality differences.

      A final point concerns the comparison between response modalities. Raw SSRT suggests that saccadic inhibition is faster than manual (174 vs 266 ms), while both proposed corrections reverse this, with SSRT−T<sub>0</sub> and ΔT each indicating that saccadic inhibition is slower (the latter consistently across nearly every participant). This is one of the clearest illustrations of the paper's thesis, but it is not taken up in the discussion, which returns to modality only to note that saccadic T<sub>0</sub> varies little (the one reference to variation across action modalities appears in the modeling section, without stating its direction). It would also benefit from a caveat. Both corrected measures subtract the same T<sub>0</sub>, and manual and saccadic T<sub>0</sub> differ by roughly 130 ms, so the two do not corroborate one another independently (TS is itself longer for manual responses, and yields a shorter ΔT only once the larger manual T<sub>0</sub> is removed). The accuracy of the subtraction therefore matters here: if the manual regression slope of 0.75 reflects sub-additivity rather than attenuation, subtracting the full T<sub>0</sub> would overcorrect manual responses more than saccadic ones, and could produce the reversal on its own.

      We agree with all this. We will use the repeatability of SSRT and T0 to disattenuate the slopes and consider alternatives to subtraction to correct for visuo-motor deadtime. Before we can elaborate on the modality effect on the speed of inhibition, we need to simulate the effect of motor variability on T0, TS and SSRT. If motor noise affects TS or SSRT more or less than it affects T0, this will affect our conclusions. We will explore this issue in the revision and report any analyses that bear on the robustness of the modality effect.

      These concerns qualify rather than undermine the contribution. The core observation is robust, the diagnostic is practical and immediately applicable, and the case that a large body of work requires re-examination is well made. If the corrected analyses are carried through on the datasets already in hand, this will be an important paper for anyone who uses the stop-signal task.

      We thank the reviewer for their in-depth and thoughtful comments and suggestions

    1. eLife Assessment

      Winter months with short days are commonly associated with seasonal depression and hypersomnolence; the mechanisms behind this hypersomnolence however, remain unclear. Chen and colleagues identify a genetic basis for this phenomenon in the fly Drosophila - mutations in the circadian photoreceptor cryptochrome resulted in increased sleep under short photoperiods. These findings are valuable insights into the genetic mechanisms regulating sleep under short days. There is solid evidence that cryptochrome acts in GABAergic neurons, but only limited evidence for the proposed site of action. Further work with better techniques will be needed to identify the precise site of action of cryptochrome.

    2. Reviewer #1 (Public review):

      Summary:

      In this paper, Chen et al. identified a role for the circadian photoreceptor CRYPTOCHROME (CRY) in promoting wakefulness under short photoperiods. This research is potentially important as hypersomnolence is often seen in patients suffering from SAD during winter times. The mechanisms underlying these sleep effects are poorly known.

      Strengths:

      The authors clearly demonstrated that mutations in cry lead to elevated sleep under 4:20 Light-Dark (LD) cycles. Furthermore, using RNAi, they identified GABAergic neurons as a primary site of CRY action to promote wakefulness under short photoperiods. They then provide genetic and pharmacological evidence demonstrating that CRY acts on GABAergic transmission to modulate sleep under such conditions.

      Weaknesses:

      The authors then went on to identify the neuronal location of this CRY action on sleep. This is where this reviewer is much more circumspect about the data provided. The authors hypothesize that the l-LNvs which are known to be arousal promoting may be involved in the phenotypes they are observing. To investigate this, they undertook several imaging and genetic experiments.

      While the authors have made improvements in this resubmitted manuscript, there are still multiple concerns about the paper. I think the authors provide enough evidence suggesting that CRY plays a role in sleep under short photoperiod. The data also supports that CRY acts in GABAergic neurons. However, the identity of the GABAergic neurons involved in this CRY dependent sleep mechanism remains unclear. Similarly, whether l-LNvs are the target of this GABA mediated sleep regulation under short photoperiod is not fully demonstrated. The data presented suggests that but does not conclusively prove it.

    3. Reviewer #3 (Public review):

      Summary:

      In humans, short photoperiods are associated with hypersomnolence. The mechanisms underlying these effects is, however, unknown. Chen et al. use the fly Drosophila to determine the mechanisms regulating sleep under short photoperiods. They find that mutations in the circadian photoreceptor cryptochrome (cry) increase sleep specifically under short photoperiods (e.g. 4h light : 20 h dark). They go on to show that cry is required in GABAergic neurons and that the effects of the cry mutation on sleep are mediated by alterations in GABA signalling. Further, they suggest that the relevant subset of GABAergic neurons are the well-studied small ventral lateral neurons that they suggest inhibit the arousal promoting large ventral neurons via GABA signalling.

      Strengths:

      Genetic analysis to show that cryptochrome (but not other core clock genes) mediates the increase in sleep in short photoperiods, and circuit analysis to localise cry function to GABAergic neurons.

      Weaknesses:

      The authors' have substantially revised their manuscript, and the manuscript is much better for the revisions. However, the idea that the sLNvs are GABAergic is unfortunately still not well supported by the data. The authors have acknowledged the limitations of their methods though which is very welcome, and a substantial improvement.

    4. Reviewer #4 (Public review):

      Summary:

      Short photoperiod is an important experimental manipulation in neurobiology, endocrinology, and metabolism studies. However, the molecular mechanisms by which short photoperiod gives rise to behavioral phenotypes that are seen in seasonal affective disorders remain unknown. Using the classic circadian model organism Drosophila, this study examines short photoperiod-induced hypersomnolence and identifies the circadian photoreceptor cryptochrome as a regulator of GABAergic tone within the clock neural circuit to promote wakefulness under short photoperiod conditions. The discovery has broad implications for understanding how short photoperiod modulates neural inhibition in circadian circuits in regulating sleep.

      Strengths:

      The Drosophila model provided a powerful platform to dissect the molecular mechanisms underlying short photoperiod-induced hypersomnolence. A battery of behavioral, imaging, circuit-manipulation approaches were employed to test the novel hypothesis that the circadian photoreceptor cryptochrome modulates GABAergic tone within the clock neural circuit to promote wakefulness under short photoperiod conditions.

      Weaknesses:

      The current model proposed by the authors suggests that the small ventral lateral neurons of the Drosophila clock circuit are GABAergic; however, this remains unclear. At present, the field lacks sufficient data and validated reagents to definitively establish the GABAergic identity of these neuropeptidergic neurons.

    5. Author response:

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      In this paper, Chen et al. identified a role for the circadian photoreceptor CRYPTOCHROME (CRY) in promoting wakefulness under short photoperiods. This research is potentially important as hypersomnolence is often seen in patients suffering from SAD during winter times. The mechanisms underlying these sleep effects are poorly known.

      Strengths:

      The authors clearly demonstrated that mutations in cry lead to elevated sleep under 4:20 Light-Dark (LD) cycles. Furthermore, using RNAi, they identified GABAergic neurons as a primary site of CRY action to promote wakefulness under short photoperiods. They then provide genetic and pharmacological evidence demonstrating that CRY acts on GABAergic transmission to modulate sleep under such conditions.

      Weaknesses:

      The authors then went on to identify the neuronal location of this CRY action on sleep. This is where this reviewer is much more circumspect about the data provided. The authors hypothesize that the l-LNvs which are known to be arousal promoting may be involved in the phenotypes they are observing. To investigate this, they undertook several imaging and genetic experiments.

      While the authors have made improvements in this resubmitted manuscript, there are still multiple concerns about the paper. I think the authors provide enough evidence suggesting that CRY plays a role in sleep under short photoperiod. The data also supports that CRY acts in GABAergic neurons. However, there are still major issues with the quality of the confocal images presented throughout the paper. In many cases it appears that the images are oversaturated with poor resolution, making it hard to understand what is going on. In addition, none of the drivers used in this study are specific to the neurons the authors aim to manipulate. Therefore, the identity of the GABAergic neurons involved in this CRY dependent sleep mechanism remains unclear. Similarly, whether l-LNvs are the target of this GABA mediated sleep regulation under short photoperiod is not fully demonstrated. The data presented suggests that but does not prove it.

      Major concerns:

      (1) While the authors provided sleep parameters like consolidation or waking activity for some experiments. These measurements are still not shown for several experiments (for example Figures 2E, 3, 4, 5, and 6). These data are essential, these metrics must be reported for all sleep experiments.

      These metrics have now been added to Fig.2 S4 and 5, Fig.3 S1-3, Fig.4 S2 and 3, Fig.5 S2 and 4, as well as Fig.6 S1.

      (2) Line 144 "We fed flies with agonists of GABA-A (THIP) and GABA-B receptor (SKF-97541) (Ki and Lim, 2019; Matsuda et al., 1996; Mezler et al., 2001). Both drugs enhance sleep in WT," The proper citation is needed here, Dissel et al., 2015 PMID:25913403. Both THIP and SKF-97541 were used in that paper.

      Thank you for pointing this out. We have modified our manuscript accordingly.

      (3) Figure 2C and 2F: it appears that the control data is the same in both panels. That is not acceptable.

      Thank you for pointing this out. We are now using data from control flies that were monitored in the same experiments as the experimental groups.

      (4) Figure 4A: With the quality of the images, it is impossible to assess whether GABA levels are increased at the l-LNvs soma.

      We apologize for the poor quality. Unfortunately, the GABA immunostaining does not work very well in our hands and thus the background is high. We have now commented on this issue in the fourth paragraph of discussion and have toned down our conclusions regarding the GABAergic s-LNv—l-LNv circuitry in this revised version of the manuscript.

      (5) Fig 4 S1A shows colabeling of l-LNvs and Gad1-Gal4 expressing neurons. They are almost 100% overlapping signals. This would indicate that the l-LNvs are GABAergic themselves, or that there is a problem with this experiment.

      Fig 4 S1A demonstrates the expression pattern of SYT-GFP driven by Gad1GAL4, which should label the synaptic terminals of GABAergic neurons. Therefore, the labeling observed at l-LNvs suggest that GABAergic neurons project to l-LNvs. This is further validated by the GRASP and trans-Tango experiments.

      (6) Fig 4 S1B: Again, I can see colabelling of the GFP and PDF staining, suggesting that Gad1-Gal4 expresses in l-LNvs.

      Fig 4 S1B demonstrates anatomical sites where GABAergic neurons project to and form synaptic connections with PDF neurons. Therefore, GFP signals at the l-LNvs suggest that these cells receive synaptic inputs from GABAergic neurons, echoing the results shown in Fig 4 S1A.

      (7) Line 184: "Consistently, knocking down Rdl in the l-LNvs rescues the long sleep phenotype of cry mutants (Figure 4-figure supplement 1D)." This statement is incorrect as the driver used for this experiment, 78G01-GAL4 is not specific to the l-LNvs, so it is possible that the phenotypes observed are not coming from these neurons.

      Thank you for pointing this out. We have modified our manuscript to note this.

      (8) Figure 4G-K: None of these manipulations are specific to the l-LNvs. The authors describe 10H10-GAL4 and 78G01-GAL4 as l-LNvs specific tools, but this is not the case. Why not use the SS00681 Split-GAL4 line described in Liang et al., 2017 PMID: 28552314? It is possible that some of the effects reported in this manuscript are not caused by manipulating the l-LNvs.

      Thank you for pointing this out. We have now modified our manuscript to avoid misleading remarks. We have used SS00681 Split-GAL4 to express TrpA1 but did not observe any substantial effect on sleep duration under short photoperiod. Therefore, we did not use this line for further experiments.

      (9) Similarly for the manipulation of s-LNvs, the authors cannot rule out effect that are coming from other cells as R6-GAL4 is not specific to s-LNvs.

      We have now modified our manuscript to avoid misleading remarks.

      (10) The staining presented in Fig 5 S1 is not very convincing. Difficult to see whether Gad1-GAL4 only expresses in the s-LNvs.

      We have now quantified the GFP signal in the l-LNvs and s-LNVs in Fig.5 S1B and D. As can be seen, the s-LNvs show prominent signal above the background while the l-LNvs do not.

      Reviewer #3 (Public review):

      Summary:

      In humans, short photoperiods are associated with hypersomnolence. The mechanisms underlying these effects is however, unknown. Chen et al. use the fly Drosophila to determine the mechanisms regulating sleep under short photoperiods. They find that mutations in the circadian photoreceptor cryptochrome (cry) increase sleep specifically under short photoperiods (e.g. 4h light: 20 h dark). They go on to show that cry is required in GABAergic neurons and that the effects of the cry mutation on sleep are mediated by alterations in GABA signalling. Further, they suggest that the relevant subset of GABAergic neurons are the well-studied small ventral lateral neurons that they suggest inhibit the arousal promoting large ventral neurons via GABA signaling

      Strengths:

      Genetic analysis to show that cryptochrome (but not other core clock genes) mediates the increase in sleep in short photoperiods, and circuit analysis to localise cry function to GABAergic neurons.

      Weaknesses:

      The authors' have substantially revised their manuscript, and the manuscript is better for the revisions. However, the conclusion that the sLNvs are GABAergic is unfortunately still not well supported by the data. A key sticking point remains the anti GABA immunostaining, and specific driver lines for sLNvs and lLNvs.

      The authors should tone down their conclusions to reflect the fact that their data, as presented, does not support the model that cry acts in sLNvs to modulate GABA signalling onto lLNvs and thus modulate sleep.

      Thank you for the comments. We have now toned down our conclusions regarding the GABAergic s-LNv—l-LNv circuitry in this revised version of the manuscript in the Introduction, Results and Discussion.

      Reviewer #4 (Public review):

      Summary:

      Short photoperiod is an important experimental manipulation in neurobiology, endocrinology, and metabolism studies. However, the molecular mechanisms by which short photoperiod gives rise to behavioral phenotypes that are seen in seasonal affective disorders remain unknown. Using the classic circadian model organism Drosophila, this study examines short photoperiod-induced hypersomnolence and identifies the circadian photoreceptor cryptochrome as a regulator of GABAergic tone within the clock neural circuit to promote wakefulness under short photoperiod conditions. The discovery has broad implications for understanding how short photoperiod modulates neural inhibition in circadian circuits in regulating sleep.

      Strengths:

      The Drosophila model provided a powerful platform to dissect the molecular mechanisms underlying short photoperiod-induced hypersomnolence. A battery of behavioral, imaging, circuit-manipulation approaches was employed to test the novel hypothesis that the circadian photoreceptor cryptochrome modulates GABAergic tone within the clock neural circuit to promote wakefulness under short photoperiod conditions.

      Weaknesses:

      The current model proposed by the authors suggests that the small ventral lateral neurons of the Drosophila clock circuit are GABAergic; however, this remains unclear. At present, the field lacks sufficient data and validated reagents to definitively establish the GABAergic identity of these neuropeptidergic neurons.

      Thank you for the comments. We have now toned down our conclusions regarding the GABAergic s-LNv—l-LNv circuitry in this revised version of the manuscript.

      Recommendations for the authors:

      The manuscript has improved after revisions. However, the evidence in support of the claim that the sLNVs secrete GABA onto the lLNvs remains unconvincing. The evidence that loss of cry in GABAergic neurons modulates sleep is solid. However, the authors' claim that the sLNVs are the relevant GABAergic neurons is not sufficiently backed up by the evidence presented. We suggest that the authors tone down their conclusions to reflect this.

      Thank you for the comments. We have now toned down our conclusions regarding the GABAergic s-LNv—l-LNv circuitry in this revised version of the manuscript in the Introduction, Results and Discussion.

      Reviewer #3 (Recommendations for the authors):

      Minor points:

      (1) The authors suggest that the effects of cry on sleep and mediated by the Rdl receptor, and use the GABA agonist THIP as support of this argument. However THIP acts on the Lcch3 and Grd receptors, not Rdl

      Thank you for pointing this out. We have modified relevant discussion accordingly.

      (2) In several instances (e.g. line 66, line 148), the authors use 'consistently' in the sense of 'consistent with previous data'. It would be better if they rephrase this.

      This has been fixed.

      Reviewer #4 (Recommendations for the authors):

      It is my pleasure to serve as a reviewer for this revised manuscript. The authors have carefully revised the manuscript in response to the critiques raised by all previous reviewers and have used all the available reagents to conduct additional experiments to assess the GABAergic properties of the small ventral lateral neurons (sLNv). Although it remains unclear in the field whether sLNvs co-transmit GABA, this study raises this possibility within an interesting biological relevant context. I recommend this manuscript for final publication.

      Thank you for your comments.

    1. eLife Assessment

      This important study reports dynamic reprogramming of global H3K4me2 during the oocyte-to-embryo transition in mice. The data presented support the main conclusion and are generally convincing. The work will be of interest to researchers working on epigenetic reprogramming during embryonic development.

    2. Reviewer #2 (Public review):

      Chong Wang et al. investigated the role of H3K4me2 during the reprogramming processes in mouse preimplantation embryos. The authors show that H3K4me2 is erased from GV to MII oocytes and re-established in the late 2-cell stage by performing Cut & Run H3K4me2 and immunofluorescence staining. Erasure and re-establishment of H3K4me2 have not been studied well, and profiling of H3K4me2 in germ cells and preimplantation embryos is valuable to understanding the reprogramming process and epigenetic inheritance.

      (1) "The authors' assertion that H3K27me3 did not change from GV to MII stage (Author response) is noted; however, this data was not provided. To validate the technical success of the CUT&RUN protocol in MII oocytes-a stage characterized by chromatin condensation and low DNA input-it is essential that the authors provide their internal H3K27me3 data as a positive control.

      Without showing that a stable mark (like H3K27me3) can be successfully mapped in their MII samples, the 'disappearance' of H3K4me2 peaks cannot be distinguished from a technical failure of the assay at this developmental stage. Furthermore, I remain skeptical of the inclusion of the first polar body as a DNA quantity, as polar body chromatin is often undergoing degradation and may not reflect the epigenetic state of the oocyte itself.

      (2) I remain concerned by the inconsistencies regarding KDM1A (LSD1) expression. The authors claim in the text and Figure 4A that KDM1A is 'rarely expressed' during mouse embryonic development. However, their own Extended Data Figure 3A shows high expression of KDM1A in oocytes, 4-cell, and 8-cell stages. Both figures reportedly use published RNA-seq data. The authors must explain how the same gene in the same developmental stages can appear 'rarely expressed' in one figure and 'expressed' in another.

      (3) Page 6 (Line 161-165) The H3K4me2 demethylases KDM1B (LSD2) were also highly expressed in growing oocytes but showed decreased expression in MII oocytes (Figure S3A, see also revised heatmap). We have re-analyzed the expression data and corrected the heatmap normalization; the revised Figure S3A now accurately shows that Kdm1b is highly expressed in growing oocytes, with lower levels in 8- week and MII oocytes, consistent with its role in maternal imprint establishment.

      The expression data in growing oocytes are missing in Fig S3A.

      (4) The authors' proposal to use transcriptome data to confirm TCP specificity is insufficient (Author reply). Since TCP is a small-molecule inhibitor of protein activity, its primary effect is not the reduction of mRNA transcripts, but the inhibition of the enzymes themselves.

      To support their claim that the observed H3K4me2 resetting and developmental arrest are specifically due to KDM1B inhibition, I recommend:

      a. Perform Western Blot for KDM1A/B

      b. Use a Selective Inhibitor: Use a more selective KDM1A inhibitor (e.g. GSK-LSD1)

      c. Address the Discrepancy: Provide a clear biological explanation for why chemical inhibition leads to 4-cell arrest while a maternal KO of KDM1B survives to E10."

      (5) "The authors' response that IF and CUT&RUN cannot be compared because of 'different analysis models' is not scientifically sound in this context.

      Quantitative Contradiction: A 1,000-fold increase in global IF signal must be reflected in the genomic landscape. If only 251 genes show a gain in H3K4me2 in the CUT&RUN data, this represents a negligible fraction of the genome, directly contradicting the 'dramatic increase' claimed in the IF data.

      Lack of Spike-in Normalization: Did the authors use a spike-in for their CUT&RUN? Without global normalization, CUT&RUN only reports relative changes. If H3K4me2 increased everywhere, a non-normalized CUT&RUN library would fail to show it, making the data misleading. Currently, the two datasets provide two different versions of biological reality. A third validation (e.g., Spike-in Normalized CUT&RUN) would be recommended to determine which dataset is accurate."

    3. Reviewer #3 (Public review):

      Summary:

      The study "Resetting of H3K4me2 during mammalian parental-to-zygote transition" provides valuable insights into the dynamic changes in H3K4me2 during early embryonic development.

      Strengths:

      The findings provide valuable insights into the temporal and spatial dynamics of H3K4me2 and its potential role in zygotic genome activation (ZGA).

      Weaknesses:

      Key areas for improvement include enhancing the innovation and novelty of the study, providing robust functional validation, establishing a clear model for H3K4me2's role, and addressing technical and presentation issues. While the findings are significant, the current manuscript falls short in several critical areas. Addressing these major and minor issues will significantly strengthen the study's contribution to the field of epigenetic reprogramming and embryonic development.

      Comment on revised version:

      It would be better for the author to directly provide some experimental or analytical data rather than discussing and defending.

    4. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public Review):

      (1) The study emphasizes H3K4me2, which often serves as a precursor to H3K4me3, a well-studied modification during early development. Analyzing the new H3K4me2 dataset alongside published H3K4me3 data is crucial for comprehensively understanding epigenetic reprogramming post-fertilization and the interplay between histone modifications. However, the current analysis is preliminary and lacks depth.

      We fully agree with this valuable suggestion. Our research group has previously systematically profiled H3K4me3 dynamics in human and mouse early embryos, and the relevant results have been published in Science (2019). The core objective of the current study is to explore the erasure, re-establishment and biological functions of H3K4me2 during mammalian parental-to-zygote transition. To enrich our analysis, we have now integrated our H3K4me2 data with publicly available H3K4me3 datasets for joint analysis. The results clearly demonstrate that H3K4me2 is not merely a precursor of H3K4me3. These two histone marks present distinct genome-wide distribution patterns and perform independent regulatory roles in embryonic epigenetic reprogramming. We have added the joint analysis results and relevant discussions in the revised manuscript to elaborate the crosstalk between H3K4me2 and H3K4me3.

      Manuscript Revisions

      All supplementary analyses and discussions are located in the Results and Discussion sections, Results section (Page 6, Lines 156–165) (Page 7, Lines 197–201) (Page 10, Lines 274–278) and Discussion section (Page 13- 14, Lines 374–382) (Page 15, Lines 424–429.

      (2) Tranylcypromine (TCP) is known as an irreversible inhibitor of monoamine oxidase and LSD1. While the authors suggest TCP inhibits the expression of LSD2, this assertion is questionable. Given TCP's potential non-specific effects in cells, conclusions related to the experiments using TCP should be made with caution.

      We highly appreciate this important reminder about the off-target effects of TCP. We have supplemented two classic literatures (J Am Chem Soc, 2010; Mol Cell, 2010) which have proven that TCP acts as an irreversible inhibitor targeting both LSD1 (KDM1A) and LSD2 (KDM1B). According to our transcriptome and protein detection data, the endogenous expression level of LSD1 is extremely low in mouse early embryos. Therefore, LSD2 is the primary functional target of TCP in our embryonic experimental system. All conclusions derived from TCP treatment experiments are described prudently in the full text to avoid over-interpretation.

      Manuscript Revisions

      Relevant supplements are made in the Results section (Page 7, Lines 197–201).

      (3) Some batches of H3K4me2 antibody are known to cross-react with H3K4me3. Has the H3K4me2 antibody used in CUT&RUN been tested for such cross-reactivity? Heatmaps in the figures indeed show similar distribution for H3K4me2 and H3K4me3, further raising concerns about antibody specificity.

      Thank you for raising this critical question regarding antibody specificity, which is essential for the reliability of CUT&RUN experiments. The H3K4me2 antibody used in this study was purchased from Millipore (Cat. No. 07030). Based on the manufacturer’s product specification and our internal verification, this antibody has very low cross-reactivity with H3K4me3.The similar distribution shown in heatmaps is not caused by antibody contamination. Instead, it reflects the inherent spatial correlation between H3K4me2 and H3K4me3 on chromatin in early embryos.

      (4) Certain statements lack supporting references or figures (examples on page 9 can be found on line 245, line 254, and line 258).

      We apologize for the inadequate citation in the original manuscript. We have comprehensively checked the full text and added standard peer-reviewed references to all statements without literature support. Specifically, we have supplemented corresponding references for the content on Page 9, Line 259 and Line 266 as suggested. We also completed a full-text inspection to fix similar problems in other positions.

      Manuscript Revisions

      References are supplemented on Page 9, Line 259 and Line 266.

      (5) Extensive language editing is recommended to clarify ambiguous sentences. Additionally, caution should be taken to avoid overstatement - most analyses in this study only suggest correlation rather than causality.

      We fully accept this suggestion. We have thoroughly revised all ambiguous, redundant and grammatically problematic sentences throughout the manuscript to improve readability and academic rigour. Furthermore, we have carefully modified all overstated expressions. For all experimental results and bioinformatics analyses, we only use words such as correlate with, suggest, indicate to describe correlative relationships. All inappropriate causal inferences have been completely removed to ensure objective presentation of our data.

      Manuscript Revisions

      Full manuscript is polished and revised.

      Reviewer #2 (Public Review):

      (1) The authors claim that the Cut & Run worked for MII oocytes, zygotes, and the 2-cell embryos. However, it is unclear if H3K4me2 is erased during the stage or if the Cut & Run did not work for these samples. To support the hypothesis of the erasure of H3K4me2, the authors conducted immunofluorescence staining, and H3k4me2 was undetected in the MII oocyte, PN5, and 2-cell stage. However, the published papers showed strong staining of H3K4me2 at the zygote stage and 2-cell stage ((Ancelin et al., 2016; Shao et al., 2014)). The authors need to cite these papers and discuss the contradictory findings.

      The authors used 165 MII oocytes and 190 GV oocytes for the Cut & Run. The amount of DNA in MII oocytes is halved because of the emission of the first polar body. Would it be a reason that H3K4me2 has fewer H3K4me2 peaks in MII oocytes?

      Thank you for putting forward these thoughtful questions. Firstly, we have cited two published literatures (Ancelin et al., 2016; Shao et al., 2014) in the revised manuscript and discussed the inconsistent immunofluorescence results. The main reason for the discrepancy lies in different confocal microscope parameters including laser power, gain and exposure time adopted by different laboratories. In our study, we used unified imaging parameters to continuously observe samples from GV oocytes to blastocysts, so weak H3K4me2 signals at zygote and two-cell stages could not be detected. When we adjust parameters specifically for these stages, weak fluorescence signals can be observed. We have elaborated this point in the Discussion section.

      Secondly, we clarify that the reduction of H3K4me2 peaks in MII oocytes is not caused by decreased DNA content. Although MII oocytes extrude the first polar body during maturation, we collected the polar body together with oocytes in all CUT&RUN experiments, so the total DNA content of MII samples is not reduced. Combined with previous studies on human oocytes, we confirm that the loss of H3K4me2 peaks from GV to MII stage is a real physiological epigenetic change accompanying oocyte meiotic maturation and chromatin remodeling.

      (2) The authors claim that Kdm1a is rarely expressed during mouse embryonic development (Figure 4A). However, the published paper showed that KDM1a is present in the zygote and 2-cell stage using immunostaining and western blotting ((Ancelin et al., 2016)). Additionally, this paper showed that depletion of maternal KDM1A protein results in developmental arrest at the two-cell stage, and therefore, KDM1a is functionally important in early development. The authors should have cited the paper and described the role of KDM1a in early embryos.

      We apologize for the ambiguous expression in the original manuscript. What we described is a relative expression level: in mouse early embryos, the expression of KDM1A is lower than KDM1B, rather than the absolute absence of KDM1A.

      (3) The authors used the published RNA data set and interpreted that KDM1B (LSD2) was highly expressed at the MII stage (Figure S3A). However, the heat map shows that KDM1B expression is high in growing oocytes but not at 8w_oocytes and MII oocytes. The authors need to interpret the data accurately.

      We sincerely apologize for the data misinterpretation caused by improper data normalization in the original heatmap. We have completely re-normalized the RNA-seq data and redrawn Supplementary Figure S3A.

      The updated heatmap clearly shows that KDM1B is highly expressed in growing oocytes, while its expression decreases in 8-week oocytes and MII oocytes. Combined with Figure 4A, we have rewritten the description of KDM1B expression trends across different oocyte stages, and all textual descriptions are now consistent with the corrected data.

      Manuscript Revisions

      Supplementary Figure S3A is remade; data interpretation is revised in the Results section. Supplementary Figure S3A (remade); Results section (Page 42).

      (4) All embryos in the TCP group were arrested at the four-cell stage. Embryos generated from KDM1b KO females can survive until E10.5 (Ciccone et al., 2009); therefore, TCP-treated embryos show a more severe phenotype than oocyte-derived KDM1b deleted embryos. Depletion of maternal KDM1A protein results in developmental arrest at the two-cell stage ((Ancelin et al., 2016)). The authors need to examine whether TCP treatment affects KDM1a expression. Western blotting would be recommended to quantify the expression of KDM1A and KDM1B in the TCP-treated embryos.

      We dig the transcriptome data to confirm the specificity of TCP to KDM1b. In addition, the intervention of TCP on the whole fertilized egg in this study increased the H3K4me2 content, and the embryo development retarding effect was more significant than that obtained by crossing with normal paternal lines after knocking down KDM1B from the mother.

      (5) H3K4me2 is increased dramatically in the TCP-treated embryos in Figure 4 (the intensity is 1,000 times more than the control). However, the Cut & Run H3K4me2 shows that the H3K4me2 signal is increased in 251 genes and decreased in 194 genes in the TCP-treated embryos. The authors need to explain why the gain of H3K4me2 is less evident in the Cut & Run data set than in the immunofluorescence result.

      Thank you for this valuable question. The inconsistent data performance between immunofluorescence (IF) and CUT&RUN is determined by the essential differences between the two technical principles.

      Immunofluorescence is a global semi-quantitative method, which reflects the total content of H3K4me2 in the whole nucleus. The 1000-fold increase refers to the overall fluorescence intensity of the nucleus. In contrast, CUT&RUN combined with high-throughput sequencing is a locus-specific quantitative method, which detects H3K4me2 enrichment changes at individual gene loci. Different analytical models and threshold settings also lead to differences in final data presentation.

    1. eLife Assessment

      This important study provides new insights into how extracellular-domain disulfide configurations regulate ligand binding and endocytosis of the natural killer cell receptor KIR2DL4, identifying a conformational switch that provides a mechanistic framework for receptor trafficking. The evidence supporting the main conclusions is solid, with complementary biochemical, structural, and functional approaches providing strong support for alternative disulfide configurations and their role in regulating receptor behavior. The work will be of interest to researchers studying natural killer cell biology, immune receptor signaling, and protein disulfide regulation.

    2. Reviewer #1 (Public review):

      Summary:

      This paper asks how the NK cell receptor KIR2DL4 binds HLA-G and undergoes endocytosis. The authors propose that an allosteric disulfide-bond switch controls whether the receptor is in a ligand-binding or non-binding state, and they support this model using mutagenesis, imaging, mass spectrometry, and structural prediction.

      Strengths:

      A major strength is the use of diverse, complementary approaches to validate the central claim. The authors combined unbiased random mutagenesis to identify key residues, confocal microscopy to track cellular localization , and mass spectrometry to quantify the redox states of specific disulfide bonds. These methods consistently support a single model: an allosteric disulfide switch. The transition between a Cys10-Cys28 bond and a Cys28-Cys74 bond serves as a functional switch that controls whether the receptor resides at the plasma membrane to bind ligand or remains inactive in endosomes.

      Comments on revised version.

      The revision substantively addresses the core weaknesses I raised, particularly on direct binding evidence, which was the most important gap. The remaining points (oligomerization confound in the SPR comparison, C10L not independently validated by imaging, and the 293T-only functional readout) are real but secondary so I'd suggest they be addressed with a sentence or two of acknowledgment in the Discussion rather than additional experiments.

    3. Reviewer #2 (Public review):

      Summary:

      Rajagopalan et al. shows how extracellular domain features regulate KIR2DL4 internalization. The trafficking phenotypes of cysteine mutants are logically organized and well summarized in Table. The disulfide mapping and differential alkylation strategy is appropriate and provides strong support for alternative disulfide configurations in D0. The higher accessibility or more selective reduction of Cys10-Cys28 as compared to Cys28-Cys74 by PDI is a key mechanistic anchor.

      Strengths:

      The identification of a conformational switch in KIR2DL4 is conceptually novel. Experimental elegance, detailed and well written.

    4. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This paper asks how the NK cell receptor KIR2DL4 binds HLA-G and undergoes endocytosis. The authors propose that an allosteric disulfide-bond switch controls whether the receptor is in a ligand-binding or non-binding state, and they support this model using mutagenesis, imaging, mass spectrometry, and structural prediction.

      Strengths:

      A major strength is the use of diverse, complementary approaches to validate the central claim. The authors combined unbiased random mutagenesis to identify key residues, confocal microscopy to track cellular localization, and mass spectrometry to quantify the redox states of specific disulfide bonds. These methods consistently support a single model: an allosteric disulfide switch. The transition between a Cys10-Cys28 bond and a Cys28-Cys74 bond serves as a functional switch that controls whether the receptor resides at the plasma membrane to bind ligand or remains inactive in endosomes.

      Weaknesses:

      (1) The core model is interesting, but some of the strongest mechanistic claims still rely heavily on structure prediction rather than direct structural evidence, especially the proposed HLA-G contact surface in Figure 6 (now in Figure 7).

      The crystal structure of KIR2DL4 has a D0 domain in the C10-C28 disulfide configuration [1]. The AlphaFold prediction is different, having a C28-C74 disulfide bond in the D0 domain. It is understood that any prediction could be wrong. Nevertheless, the AlphaFold structure did point to the possibility that KIR2DL4 exists in two different disulfide-bonded forms. We went on to demonstrate experimentally that these two forms coexist in human cells. This conclusion is independent of the structure predicted by AlphaFold.

      The second difference predicted by AlphaFold is an allosteric change in a loop distant from the disulfide bond, suggesting the possibility that it could control binding of HLA-G. Again, this prediction could be wrong. Nevertheless, considering that the KIR2DL4 used to obtain a crystal structure was in a C10-C28 bond configuration and did not bind HLA-G [1], we wondered if HLA-G would bind to KIR2DL4 in a C28-C74 configuration.

      New experiments included in the revision have shown that a purified Cys10Leu KIR2DL4 mutant binds HLA-G (new Figure 6). Solving the structure of a KIR2DL4–HLA-G complex would be ideal, but this has not been possible thus far. The difficulty in crystallizing KIR2DL4 may be due, in part, to its propensity to form oligomers [1], as shown in the new Figure S6.

      The addition of both direct binding of HLA-G to KIR2DL4 and functional data showing that KIR2DL4 induces an ISG response when in the C28-C74 but not in the C10-C28 configuration strengthens the conclusion that disulfide switching controls ligand binding and downstream signaling relevant to NK cell interactions with HLA-G in early pregnancy.

      (2) The paper supports an effect of the disulfide state on trafficking and uptake, but the case for direct KIR2DL4-HLA-G binding still feels somewhat indirect. The manuscript itself notes that direct binding had not been previously shown, and the current explanation partly depends on inference about which disulfide state is present.

      Direct binding and affinity measurements of HLA-G bound to the Cys10Leu KIR2DL4 mutant (in a C28-C74 disulfide form) are in the new Figure 6. This crucial result is also consistent with functional data. New experiments (new Figure 5D) have shown that the ability of HLA-G to stimulate a transcriptional interferon-stimulated gene (ISG) response occurred with the C28-C74 form, but not the C10-C28 form of KIR2DL4.

      Surface plasmon resonance data showed for the first time direct binding between the C28-C74 form of KIR2DL4 and soluble HLA-G, with a K<sub>D</sub> of 1.6 mM (new Figure 6). Binding to WT KIR2DL4, which is in both configurations, C10-C28 and C28-C74, was also detected but with a lower affinity (K<sub>D</sub> = 19.4 mM). The purified WT KIR2DL4 formed oligomers (new Figure S6). In addition, binding of HLA-G to KIR2DL4 depended on the sequence of the peptide presented by HLA-G, as only one out of three peptides tested was compatible with KIR2DL4 binding.

      This new data was obtained in the laboratory of Jamie Rossjohn at Monash University, Victoria, Australia. He, along with Jan Peterson and Priyanka Chaurasia are new co-author on our revised manuscript.

      (3) Most of the main experiments are done in transfected 293T cells, so it is still not fully clear how strongly this mechanism carries over to the more relevant NK-cell setting discussed in the paper.

      Primary resting NK cells are not amenable to transfection. Despite this technical hurdle, we have included two key findings with primary NK cells in the revision.

      (1) As in the 293T transfected cell system, we have shown that inhibition of PDI caused reduced uptake of HLA-G in primary resting NK cells (New Figure 4E, F, G). This is consistent with uptake of HLA-G by the C28-C74 form of KIR2DL4 and with a switch from C10-C28 to C28-C74 catalyzed by PDI.

      (2) We have shown that cell-surface C28-C74 KIR2DL4 on primary NK cells, as detected by mAb 2388, decreased upon inhibition of PDI, again consistent with the role of PDI in maintaining a pool of C28-C74-bonded KIR2DL4 at the cell surface (New Figure 5E, F). As shown in the original Figure 3, PDI could reduce the C10-C28 bond in purified WT KIR2DL4 in vitro.

      (4) The cellular evidence for the PDI story is not specific, since it depends a lot on inhibitor and blocking experiments that could affect the broader extracellular redox environment.

      Using inhibitors that target PDIA1 selectively, namely Rutin (PDI-specific up to 30 microM), and a PDI-specific monoclonal antibody, we found that HLA-G uptake by primary NK cells was inhibited (Figure 4C, D). We admit that pCMPS and thiol blockade by DTNB (Figure 4A, B) affect the extracellular redox environment. Data that were obtained without PDI inhibitors include the reduction of the C10-C28 bond by PDI in WT KIR2DL4 in vitro (Figure 3F), direct binding of HLA-G to KIR2DL4 in a C28-C74 disulfide conformation (Figure 6), and a functional transcriptional response to HLA-G by C28-C74 KIR2DL4 and not with the C10-C28 KIR2DL4.

      Reviewer #2 (Public review):

      Summary:

      Rajagopalan et al show how extracellular domain features regulate KIR2DL4 internalization. The trafficking phenotypes of cysteine mutants are logically organized, and well-summarized in a Table. The disulfide mapping and differential alkylation strategy are appropriate and provide strong support for alternative disulfide configurations in D0. The higher accessibility or more selective reduction of Cys10-Cys28 as compared to Cys28-Cys74 by PDI is a key mechanistic anchor.

      Strengths:

      The identification of a conformational switch in KIR2DL4 is conceptually novel. Experimental elegance, detailed and well-written.

      Weaknesses:

      Most of the mechanistic work was shown in HEK293. The authors should exhibit relevance using primary NK cells (using primary NK)

      As primary NK cells are not amenable to transfection, it is difficult to dissect the role of each disulfide form of the receptor KIR2DL4.

      Instead, we have now included PDI inhibition experiments using primary NK cells and shown that PDI inhibition reduces HLA-G uptake by primary NK cells (New Figure 4E, F, G). This is consistent with uptake of HLA-G by the C28-C74 form of KIR2DL4 and with a switch from C10-C28 to C28-C74 catalyzed by PDI.

      Furthermore, inhibition of PDI caused a decrease of C28-C74 KIR2DL4 at the cell surface of primary NK cells (New Figure 5E, F). This data is consistent with a requirement for a switch from C10-C28 to C28-C74 catalyzed by PDI, which maintains a pool of C28-C74 KIR2DL4 at the cell surface for HLA-G binding and internalization. As shown in the original Figure 3, PDI can reduce the C10-C28 bond in purified WT KIR2DL4 in vitro.

      Recommendations for the authors:

      Reviewing Editor Comments:

      To improve the strength of the evidence and the overall impact of the paper, please address the following major points:

      (1) Validation in Primary Cells:

      The central biological framing of the paper involves decidual NK cell responses to soluble HLA-G. We strongly recommend performing a critical experiment using primary NK cells to test whether PDI inhibition or thiol blockade alters KIR2DL4 surface retention and HLA-G uptake in a manner consistent with your observations in 293T cells.

      We have added new experiments with primary, resting NK cells, as described in our response to the major point 3 of reviewer #1, and to the weakness raised by reviewer #2.

      Briefly, we have included experiments in the revised manuscript on the effect of PDI inhibition on HLA-G uptake in primary NK cells (New Figure 4E, F, G) and on transient accumulation of KIR2DL4 at the cell surface (in a C28-C74 bonded form) of primary NK cells (New Figure 5E, F). The data showed that HLA-G endocytosis by primary NK cells and the presence of KIR2DL4 at the plasma membrane of primary NK cells were reduced after inhibition of PDI.

      (2) Clarification of the "Switching" Mechanism:

      The current data points toward the coexistence of the Cys10-Cys28 and Cys28-Cys74 states. Please clarify or provide evidence regarding whether a dynamic conversion occurs (e.g., prior to binding, upon ligand engagement, or during trafficking) versus a model of stable coexistence of two distinct receptor pools.

      Stable coexistence of two distinct KIR2DL4 receptor pools was a plausible hypothesis but one that is not supported by some of our data. In such a scenario, the C10-C28 form would not bind HLA-G and would reside in endosomes. It could have a role that is not related to HLA-G nor to the transcriptional response induced by HLA-G. However, our recent paper [2] showed that the transcriptional response of primary NK cells to soluble mAb #33 (bound to C10-C28) is very similar (R<sup>2</sup>=0.89) to that of resting NK cells incubated with soluble HLA-G (bound to C28-C74). These two ligands were tested at the same time, at the same molarity, and with the same primary NK cells [2].

      We don’t have answers yet to some obvious questions: is there switching after internalization of KIR2DL4 bound to mAb #33? What is the fate of C28-C74 that internalizes with HLA-G? We are not aware of technology that would answer these questions.

      A C28-C74 form, as a separate pool with residency at the cell surface, could be functional and respond to HLA-G by internalization and signaling from endosomes. However, there is no stable pool of C28-C74 KIR2DL4 at the cell surface and C28-C74 is depleted from the cell surface in the presence of PDI inhibitor (new Figure 5E, F), suggesting that C28-C74 KIR2DL4 is generated by the activity of PDI (new Figure S5). The sum of our experiments points to a tightly regulated control of KIR2DL4 biology, rather than the coexistence of two separate pools. A separate pool of C10-C28 KIR2DL4 would remain in an inactive state as far as the response to HLA-G is concerned. We favor the model whereby functional C28-C74 is generated from C10-C28 by the activity of PDI.

      Why could the response to HLA-G not be simpler? We address this point in the Discussion. One reason is that C28-C74 KIR2DL4 signaling at the plasma membrane of NK cells could be subject to inhibition by LILRB1 and NKG2A-CD94, co-expressed on NK cells, which bind to HLA-G and HLA-E, respectively, on fetal trophoblasts that encounter maternal NK cells in the decidua. These inhibitory receptors are known to be dominant against activation signals [3]. Trophoblast cells that invade the maternal decidua and encounter decidual NK cells selectively express HLA-C, HLA-E, and HLA-G. Strong inhibition signals by LILRB1 and NKG2A-CD94 could prevent activation through KIR2DL4. However, KIR2DL4 signaling, which occurs in endosomes [4] where signaling is sustained [5], can bypass these inhibitory signals at the plasma membrane.

      (3) Specificity of the PDI Model:

      Please elaborate on the relevance of extracellular PDI. Specifically, how does PDI perturbation affect the relative abundance of the two disulfide forms in a cellular context?

      We show in Figure 5E that two mAb for KIR2DL4 recognize different forms of the receptor. While mAb #33 recognizes only the C10-C28 form of the receptor, which is not at the cell surface, mAb 2238 recognizes both forms of the receptor. This allowed us to examine the effect of PDI on surface expression of the C28-C74 form of KIR2DL4 as detected by mAb 2238. We show that PDI inhibition reduces surface staining of C28-C74 (new Figure 5F), consistent with a model whereby a switch from C10-C28 to C28-C74 is catalyzed by PDI.

      A quantitative assessment of the relative abundance of the two forms of KIR2DL4 upon inhibition by PDI in a cellular context would have to be carried out by mass spec analysis of the two forms before and after treatment. That would be a very challenging experiment to perform with intact cells rather than purified proteins.

      (4) Agonist Antibody Mechanism:

      The manuscript mentions mAb #33 as a KIR2DL4 agonist. It would be highly informative for the reader if you could elaborate on whether this antibody activates the receptor by stabilizing a specific disulfide state or by driving internalization independently of HLA-G.

      We have shown that the agonist mAb #33 recognizes only the C10-C28 form (Figure 5E). We do not yet understand how it activates KIR2DL4. We do know that mAb #33 is not driving internalization considering that the receptor internalizes constitutively and is predominantly located in endosomes in the absence of HLA-G. Instead, it is the C10-C28 form of the receptor that carries mAb #33 into endosomes. Understanding how mAb #33 may function as a receptor agonist will require crystallization of the antibody bound to the receptor and is beyond the scope of this study. Structural studies of KIR2DL4 have been very difficult, due in part to its isoforms and tendency to form oligomers. It is not possible to answer your interesting question at this time.

      Minor Revisions:

      (1) Imaging Quantification:

      Ensure all figure legends include the number of independent experiments (n), specific statistical tests used, and precise alignment with the Methods section.

      This information is now included in the Methods section.

      (2) Textual Flow:

      To enhance engagement, please integrate the logic of Table 1 more explicitly into the main text of the Results section.

      This has been done.

      (3) Structural Discussion:

      Acknowledge the limitations of using structure prediction for the binding interface and discuss how these models align with existing literature on KIR-ligand interactions.

      We have described the use of AlphaFold solely as a tool to make predictions. Predictions can be wrong. Even so, they can generate new and useful hypotheses, as they did here. Existing, traditional KIR-ligand interactions are not informative in the context of the D0 domain in KIR2DL4 for the following reasons:

      The KIR2DL1/2/3 receptors with 2 Ig domains (hence 2D) have a D1 and a D2 domain. A comparison with KIR2DL4, which has a D0 and a D2 domain, may not be informative.

      The KIR3D receptors have the three domains, D0, D1 and D2. A structure of KIR3DL1 bound to HLA-B has been solved [6] by our collaborator for the revision, Dr. Jamie Rossjohn. As shown and mentioned in our manuscript (Fig. S7C and Legend), “predicted” contacts of the KIR2DL4 D2 domain with HLA-G involve residues conserved in the heavy chains of HLA-B and HLA-G and residues conserved in the KIR3DL1 and KIR2DL4 D2 domains. It is therefore likely that the KIR2DL4 D2 domain contacts HLA-G in a similar way.

      As for the KIR2DL4 D0 domain, it is very different. Due to the similarity between D2 domains of KIR3DL1 and KIR2DL4, and to the lack of a D1 domain in KIR2DL4, the KIR2DL4 D0 domain is in a completely different space than the D0 domain of KIR3DL1. “Predictions” by AlphaFold show that there could be interactions between the KIR2DL4 D0 domain and HLA-G (Figures 7 and S7). These predictions could be wrong. Nevertheless, the disulfide switch in the KIR2DL4 D0 domain correlates with a predicted change elsewhere on D0 at a position compatible with proximity to HLA-G. Furthermore, the KIR2DL4 isoform with a Cys28-Cys74 bond is “predicted” to be more aligned with a potential binding site than the Cys10-Cys28 isoform. Having no structural guide as a reference on how KIR2DL4 D0 domain may interact with HLA-G, such predictions may generate testable hypotheses.

      As we clearly state in the manuscript: “Structures of KIR2DL4–HLA-G complexes obtained experimentally are required to determine how HLA-G distinguishes the D0 domain in the alternative disulfide-bonded configurations.” (Results), and “Rules that dictate HLA-G binding to KIR2DL4 await further studies and structures of KIR2DL4–HLA-G complexes.” (Discussion).

      In the revised manuscript, we have now included SPR binding data for KIR2DL4 with HLA-G. We also show a higher affinity of HLA-G for the C28-C74 form of KIR2DL4. This has strengthened the study as it validates our model whereby switching to the functional form of the receptor allows binding of HLA-G. In this regard, we also include data showing that only the C28-C74 form of KIR2DL4 can respond to HLA-G to induce transcription of an ISG response. This provides a functional correlate to the role of the different disulfide forms of the receptor.

      Reviewer #2 (Recommendations for the authors):

      Major points to address:

      (1) Exhibit relevance using primary NK cells (using primary NK). The central biological framing is decidual NK responses to soluble HLA-G during early pregnancy, yet most mechanistic work is in 293T transfectants. The authors can perform one of the critical experiments using primary NK cells with soluble HLA-G stimulation. They should test whether PDI inhibition/thiol blockade similarly alters KIR2DL4 surface retention and HLA-G uptake in primary NK cells

      These experiments have been performed in primary NK cells and are described in the new Figure 4E, F, G and Figure 5F.

      (2) The authors should detail more about the relevance of extracellular PDI and the effect of PDI perturbation on the abundance of the two disulfide forms in cells. They should also provide evidence or discuss whether switching occurs prior to ligand binding, upon ligand engagement, or during trafficking.

      Such experiments would be very challenging. The predicted structural change is minor and may not be detectable by changes in proximity of labeled reporters. Ligand is not required for switching. We do know that PDI can convert C10-C28 into C28-C74, presumably by accessibility to the KIR2DL4 Cys28 when bonded in a C10-C28 configuration (Figure 2). How ligands (mAb #33 or HLA-G) impact KIR2DL4 structure is unknown. Data are compatible with the possibility of a stabilization of C10-C28 by mAb #33 and of C28-C74 by HLA-G.

      (3) The authors should elaborate on whether mAb #33 activates by stabilizing or by driving internalization independent of HLA-G. This is very interesting to the reader, given mAb #33 as a KIR2DL4 agonist.

      The question is undeniably interesting. mAb #33 is not required for internalization but is required for signaling. The C10-C28 KIR2DL4 configuration to which it binds internalizes constitutively and resides mainly in endosomes. How mAb #33 internalization by KIR2DL4 (not the reverse) results in signaling is not known. Nor is it known for the alternative form, C28-C74, which binds HLA-G, internalizes it, and signals for a transcriptional response very similar to that of C10-C28 bound to mAb #33 [2]. The C28-C74 KIR2DL4 configuration is retained, probably transiently, at the cell surface, to be available for HLA-G binding and internalization.

      Minor points to address:

      (1) The authors should ensure that all imaging quantifications include n, the number of experiments, and statistical treatment. Some are described in the Methods section. Please align figure legends with the method in detail.

      Details of the imaging experiments are provided in the Methods section.

      (2) Please summarize the Table 1 logic in the main text for enhanced reader engagement.

      This has been done.

      (3) The authors identify both Cys10-Cys28 and Cys28-Cys74 states in human cells. The data points towards coexistence rather than towards dynamic conversion. Please provide clarity on switching versus stable coexistence of two forms.

      Stable coexistence of two distinct KIR2DL4 receptor pools was a plausible hypothesis but one that is not supported by some of our data. In such a scenario, the C10-C28 form would not bind HLA-G and would reside in endosomes. It could have a role that is not related to HLA-G nor to the transcriptional response induced by HLA-G. However, our recent paper [2] showed that the transcriptional response of primary NK cells to soluble mAb #33 (bound to C10-C28) is very similar (R<sup>2</sup>=0.89) to that of resting NK cells incubated with soluble HLA-G (bound to C28-C74). These two ligands were tested at the same time, at the same molarity, and with the same primary NK cells [2].

      We don’t have answers yet to some obvious questions: is there switching after internalization of KIR2DL4 bound to mAb #33? What is the fate of C28-C74 that internalizes with HLA-G? We are not aware of technology that would answer these questions.

      A C28-C74 form, as a separate pool with residency at the cell surface, could be functional and respond to HLA-G by internalization and signaling from endosomes. However, there is no stable pool of C28-C74 KIR2DL4 at the cell surface and C28-C74 is depleted from the cell surface in the presence of PDI inhibitor (new Figure 5E, F), suggesting that C28-C74 KIR2DL4 is generated by the activity of PDI (new Figure S5). The sum of our experiments points to a tightly regulated control of KIR2DL4 biology, rather than the coexistence of two separate pools. A separate pool of C10-C28 KIR2DL4 would remain in an inactive state as far as the response to HLA-G is concerned. We favor the model whereby functional C28-C74 is generated from C10-C28 by the activity of PDI.

      Why could the response to HLA-G not be simpler? We address this point in the Discussion. One reason is that C28-C74 KIR2DL4 signaling at the plasma membrane of NK cells could be subject to inhibition by LILRB1 and NKG2A-CD94, co-expressed on NK cells, which bind to HLA-G and HLA-E, respectively. These inhibitory receptors are known to be dominant against activation signals [3]. Trophoblast cells that invade the maternal decidua express HLA-E and HLA-G and encounter decidual NK cells that express LILRB1 and NKG2A-CD94. Strong inhibition signals induced by these two receptors could prevent activation through KIR2DL4. KIR2DL4 signaling in endosomes protects it from these inhibitory signals and benefits from the sustained signaling property of endosomal signaling platforms [5].

      (1) S. Moradi et al., The structure of the atypical killer cell immunoglobulin-like receptor, KIR2DL4. J Biol Chem 290, 10460-10471 (2015).

      (2) S. Rajagopalan et al., The fetal trophoblast cell marker HLA-G activates a type I interferon response in primary NK cells through the receptor KIR2DL4. Sci Signal 19, eadv2400 (2026).

      (3) E. O. Long, H. S. Kim, D. Liu, M. E. Peterson, S. Rajagopalan, Controlling natural killer cell responses: integration of signals for activation and inhibition. Annu Rev Immunol 31, 227-258 (2013).

      (4) S. Rajagopalan et al., Activation of NK cells by an endocytosed receptor for soluble HLA-G. PLoS Biol 4, e9 (2006).

      (5) M. Miaczynska, L. Pelkmans, M. Zerial, Not just a sink: endosomes in control of signal transduction. Curr Opin Cell Biol 16, 400-406 (2004).

      (6) J. P. Vivian et al., Killer cell immunoglobulin-like receptor 3DL1-mediated recognition of human leukocyte antigen B. Nature 479, 401-405 (2011).

    1. eLife Assessment

      This study attempts to predict spatial patterns of genetic diversity among coral reef dwellers, estimated with a computationally efficient, standardized k-mer pipeline, from satellite-derived seascape variables. The curated dataset of 18 coral reef taxa across 173 reefs is valuable, but the evidence for the central ecological claims is in its current form incomplete: pairwise genetic distances are modelled as independent observations, the reef-level random effects used as the response in the environmental analysis likely conflate species composition with environmental signal, and , and the coverage is likely too sparse to robustly support conclusions about loss of inter-and intra-specific within-reef diversity over time. The work will be of interest to conservation geneticists, marine biologists, and researchers developing Earth-observation approaches to biodiversity monitoring.