10,000 Matching Annotations
  1. Last 7 days
    1. eLife Assessment

      The results of this study are important and the approach to dissect the developmental contribution of RIF1 function is convincing. The finding that replication timing and gene expression may be independently controlled is intriguing and provides a strong foundation for future research. The work will be of interest for researchers both in the transcription and the replication field, especially for scientists investigating the interplay between the two processes.

    2. Reviewer #1 (Public review):

      The authors sought to determine how Rif1 contributes to DNA replication timing (RT), transcriptional regulation, and embryonic development using zebrafish. They generated a maternal-zygotic rif1 knockout line and examined developmental phenotypes, genome-wide replication timing profiles, RNA-seq, and nascent transcription (SLAM-seq) during early embryogenesis.

      Their major findings in this manuscript are

      (1) Rif1 is not essential for zebrafish viability, unlike its partially essential role in mice.

      (2) Rif1 deficiency causes defects in female sex determination, delayed epiboly, and reduced primitive erythropoiesis.

      (3) Genome-wide RT is altered by Rif1, but developmental stage has a much larger influence than Rif1 itself.

      (4) Rif1 is required for the proper maturation ("sharpening") of the RT program during development rather than for specific developmental RT switches.

      (5) Rif1 has a much stronger effect on transcription during zygotic genome activation (ZGA) than on replication timing at these early stages.

      (6) Loss of Rif1 leads to increased expression of early zygotic genes, indicating that Rif1 normally suppresses widespread transcription during ZGA.

      Overall, the work proposes that Rif1 independently regulates replication timing and transcription, with these two functions becoming most prominent at different developmental stages.

      The major strengths of the manuscript are as follows.

      (1) the study combines multiple genome-wide approaches including whole-genome RT profiling, RNA-seq, SLAM-seq in combination with gene KO and developmental analyses.

      (2) One of the strongest points is that the authors conducted the analyses at multiple developmental stages rather than a single point.

      (3) The most important conclusion is that the Rif1 regulates transcription during development in a manner largely independent of its RT function, which was further strengthened by the additional data provided in the revised manuscript.

      On the other hand, the weakness of the manuscript includes the followings.

      (1) Limited mechanistic insight. The questions such as where Rif1 binds on the chromatin (in relation to the transcriptional promoters/ enhancers and replication origins).

      (2) Which functional domains of RIf1 are involved in regulation of transcription and replication (Is PP1 recruitment required for transcription regulation?) are not addressed.

      (3) Since Rif1 is known to be involved in chromatin organization/ nuclear architecture regulation, the studies addressing this (Hi-C, compartment analyses, ATAC seq etc) would provide important mechanistic information.

      (4) Female sex determination phenotype is intriguing, but it remains largely descriptive, and its mechanisms are elusive at the moment.

      Overall, the results support the authors' conclusions and they have successfully provided answers to the authors' original questions on developmental roles of Rif1 in RT and transcription in vertebrate.

      Comments on revised version:

      The authors responded to my comments in a largely satisfactory manner. They have conducted additional analyses and concluded that Rif1 regulates transcription during ZGA largely independently of its classical RT function, which is an important finding.

      Although authors did not examine origin firing and replication fork rate in rif1 KO cells, which I suggested in my original review, this can be saved for their future studies.

      I think the revised manuscript has been improved and provides important basic information on the functions of the conserved Rif1 protein in RT and transcriptional regulation.

      I have no further recommendation for additional experiments or data analyses.

    3. Reviewer #2 (Public review):

      This study by Masser et al. analyzes global replication timing and gene expression in rif-1 null zebrafish. This work is an extension of their previous report of the normal replication timing pattern during wild-type zebrafish development. The major valuable finding here is that Rif1 is not essential for viability in zebrafish, and - counter to expectation from studies in cultured cells and other species - late replication does not strongly depend on Rif1. Instead, the data suggest that Rif1 subtly sharpens replication timing pattern during normal development rather than function generally to delay replication timing. In the absence of Rif1, the normal pattern establishment is somewhat delayed. The authors also document some changes in expression during development with more genes being repressed by Rif1 than activated at some early stages.

      The study and analysis are generally rigorous, and the conclusions are supported by convincing data. Given the strong link between replication timing and cell type/development, studying timing in a whole developing organism is important. The experimental approach is technically challenging, particularly the bioinformatic analysis. The scientific advance here is largely confined to documenting the timing of Rif1-affected transcription, the unanticipated effect of the rif1 deletion on replication timing and on sex determination, though the latter is not explored. The difference in timing of the transcription phenotypes and replication phenotypes suggests they may be very distinct Rif1 roles. The overall study a useful set of findings and detailed data for future work.

      Loss of Rif1 did not affect viability, but it did strongly influence sex determination, resulting in a lower population of females. This effect is the strongest organismal phenotype, but the study provides no mechanistic explanation for the loss of females from the data gathered here.

      Comments on revised version:

      We are generally satisfied with the revised version of this manuscript.

    4. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public Review):

      In this manuscript authors examined the effect of rif1 knockout on replication timing and transcription in early embryos of zebrafish. Contrary to the expectation, genome-wide replication timing domains did not significantly change upon Rif1 knockout, although the replication timing became less dynamic in the mutant, meaning the entire genomes are replicated toward the mid S. In contrast, transcriptional profiles change by rif1 mutation throughout the embryo stage. These effects were more predominantly observed after gastrulation at the early stages of zebrafish development.

      The results presented in this manuscript provide new information on the effects of rif1 mutation on early zebrafish development, although the underlying mechanism has not been explored. The information is useful for researchers in the field of early development, with specific focus on replication and transcription regulation.

      The genome wide analyses of replication timing has been conducted and analyzed properly. The transcriptional analyses are conducted by RNA-seq and SLAM-seq (determining the nascent mRNA), and the results convincingly show the overall transcriptional patterns at different developmental stages.

      This work shows that Rif1 regulates replication timing and transcription in zebrafish embryos, while the extents of the effects vary during the developmental process. Although the data convincingly illustrate the whole picture of Rif1 KO on replication and transcription during zebrafish development, the mechanistic insight is missing. Especially, how Rif1 may or may not coordinately regulate replication and transcription during the zebrafish development has not been addressed.

      We thank the reviewer for recognizing the value of combining genome-wide replication-timing, RNA-seq, and SLAM-seq analyses across zebrafish development. We agree that the original study did not establish a molecular mechanism linking Rif1-dependent transcriptional and replication-timing effects. To address whether these effects are locally coordinated, we added a gene-centred analysis comparing replication-timing values for genes with increased, decreased, or unchanged transcript abundance at Dome (Figure 5--figure supplement 2). Differentially expressed genes did not show a clear enrichment in early- or late-replicating regions, either at Dome or at pre-MBT. These results argue against replication timing state being the primary determinant of the Dome-stage transcriptional changes. We also expanded the Discussion to explain the limitations of the current study and the need for future measurements of origin use, fork progression, chromatin state, and cell-type-specific effects. The new discussion of Nakatani et al. (2025) further places our findings in the context of evidence that Rif1-dependent replication-timing changes can be uncoupled from transcriptional changes.

      Reviewer #2 (Public Review):

      This study by Masser et al. analyzes global replication timing and gene expression in rif-1 null zebrafish. This work is an extension of their previous report on the normal replication timing pattern during wild-type zebrafish development. The major valuable finding here is that Rif1 is not essential for viability in zebrafish, and - counter to expectation from studies in cultured cells and other species - late replication does not strongly depend on Rif1. Instead, the data suggest that Rif1 subtly sharpens replication timing pattern during normal development rather than function generally to delay replication timing. In the absence of Rif1, the normal pattern establishment is somewhat delayed. The authors also document some changes in expression during development with more genes being repressed by Rif1 than activated at some early stages.

      The study and analysis are generally rigorous, and the conclusions are supported by convincing data. The manuscript is well written, though there are aspects of the presentation that could be improved for a broader scientific audience. Given the strong link between replication timing and cell type/development, studying timing in a whole developing organism is important. The experimental approach is technically challenging, particularly the bioinformatic analysis. The scientific advance here is largely confined to documenting the timing of Rif1-affected transcription, the unanticipated effect of the rif1 deletion on replication timing and on sex determination, though the latter is not explored. The work is descriptive and feels like two relatively unconnected studies, transcription and replication plus a small bit of development, and the difference in timing of the transcription phenotypes and replication phenotypes suggests they may be very distinct Rif1 roles. There isn't a lot of new insight into the mechanism of how Rif1 affects either replication timing or gene expression. As such, the overall study is an useful set of findings and detailed data for future work, but it doesn't make a big step forward in understanding the role of Rif1 or the biological processes it affects.

      Weaknesses worth addressing include the following:

      (1) Loss of Rif1 did not affect viability, but it did strongly influence sex determination, resulting in a lower population of females. This effect is the strongest organismal phenotype, but the study provides no explanation for the loss of females from the data gathered here.

      (2) The approach to distinguish nascent zygotically expressed mRNAs from maternal mRNAs is a strength. Are the differentially expressed genes related at all to regions of the genome whose replication timing is most affected? Are any of them related to the sex determination or developmental phenotypes?

      We thank the reviewer for recognizing the rigor of the analyses and the value of studying replication timing in a developing vertebrate. We revised the manuscript extensively to make the experimental logic, zebrafish developmental context, replication-timing analyses, and figure legends more accessible to a broad audience. We also quantified the gastrulation phenotype, showing an approximately one-hour delay in completion of epiboly in maternal-zygotic rif1 mutants rather than a persistent developmental arrest.

      We agree that the mechanism underlying the sex-ratio phenotype remains unresolved. The transcriptomic experiments were performed in whole embryos at stages much earlier than zebrafish sex determination and therefore cannot resolve changes in primordial germ cells or supporting gonadal somatic cells. We have avoided making a mechanistic connection between the early embryonic transcriptional changes and the adult sex-ratio phenotype and identify this as an important area for future study. To address the relationship between transcription and replication timing, we added Figure 5--figure supplement 2. Genes with increased or decreased transcript abundance at Dome were not preferentially associated with early- or late-replicating regions. Together with the distinct developmental timing of the transcriptional and replication-timing phenotypes, this supports the interpretation that Rif1 has separable roles in the two processes rather than a single local mechanism that directly couples them.

      Reviewer #3 (Public Review):

      Using the zebrafish model system, this manuscript assessed the roles of Rif1 protein in replication timing control and transcription during early development, and successfully demonstrated the differential impact of Rif1 protein in replication timing control and transcription. Moreover, the comprehensive assessments of the impacts of mutating Rif1 on animal development (including animal survival and sexual development) were assessed. Although there are works that examined Rif1's implications in replication timing and transcription separately, this work is unique in assessing all these points at once.

      The strength of this manuscript is the genomic analyses of replication timing and transcription being combined in a single model system. Consequently, this manuscript clearly demonstrates the differential impact of Rif1 in these processes during zebrafish development.

      The weakness of this manuscript is, as the authors comment in the Discussion, analyses of replication timing and transcription were performed using bulk embryos. There is a possibility that tissue-specific changes could have been masked. Tissue-specific or single-cell analysis in the future will fill the gap in the knowledge.

      Some of the findings presented in this manuscript are consistent with previous findings using different models such as Drosophila and mice, whereas other findings do not necessarily agree. I hope further studies will reveal more clearly what is common in these systems, and what is different.

      Also, the suggestion that the Rif1 protein may be implicated in a function similar to Fanconi-Anemia genes/proteins is very intriguing.

      Overall, the data presented in this manuscript sufficiently justify the authors' claims. Moreover, this manuscript provides interesting insights into Rif1's function, as well as how development could be controlled.

      We thank the reviewer for highlighting the strength of analyzing replication timing, transcription, and developmental phenotypes in the same vertebrate model. We agree that bulk-embryo measurements may mask tissue- or cell-type-specific effects. We now emphasize this limitation and the need for future tissue-specific or single-cell studies, particularly in the cell populations relevant to sex determination. We also expanded the cross-species context by discussing the recent mouse-embryo study by Nakatani et al. (2025), which supports a conserved role for RIF1 in consolidation of the replication-timing program while also indicating that replication-timing and transcriptional effects can be uncoupled. We agree that defining which Rif1 functions are conserved across zebrafish, mouse, Drosophila, and other systems, including possible relationships to Fanconi-anaemia pathways, will be an important direction for future work.

      Reviewing Editor:

      While the paper was under revision, a relevant paper from the Torres-Padilla lab was published (Nakatani et al., Developmental Cell, 2025). It complements these studies and cites the previous version of this manuscript. I suggest adding a reference in the Discussion to support the conclusions.

      We thank the Reviewing Editor for bringing the recent study by Nakatani et al. to our attention. We have added a standalone paragraph near the end of the Discussion explaining how this work complements our findings, and we have added the complete reference to the bibliography. The new Discussion text reads:

      “A recent study in mouse embryos independently identified RIF1 as a regulator of the developmental consolidation of the RT program. RIF1 depletion produced a less-defined, developmentally immature RT program, while RIF1-dependent RT changes were not correlated with transcriptional changes (Nakatani et al., 2025). Together with our findings in zebrafish, these results support a conserved role for RIF1 in sharpening replication timing during vertebrate development and indicate that its effects on replication timing can be uncoupled from changes in gene expression.”

      Recommendations for the authors:

      Reviewer #1 (Recommendations For The Authors):

      The results presented in this manuscript provide new information on the effects of rif1 mutation on replication and transcription during early zebrafish development, although the underlying mechanism has not been explored. I suggest authors consider conducting the following experiments.

      (1) Does replication timing domains have any role in Rif1-mediated regulation of transcription? It is not clear from the data presented whether transcriptionally affected genes are in the early replicating domains or late replicating domains (that appear after the shield stage). This should be examined.

      We thank the reviewer for this helpful suggestion. To address whether transcriptional effects in rif1 mutants are associated with replication timing, we assigned each gene the nearest smoothed replication timing value and compared replication timing distributions for genes whose transcript levels increased at Dome, decreased at Dome, or were not significantly changed. This analysis is now shown in Figure 5—figure supplement 2. Genes with increased or decreased transcript abundance at Dome did not show a clear enrichment for either early- or late-replicating regions relative to genes with no significant transcript change. This was also true when replication timing was examined at pre-MBT, the stage preceding the major transcriptional changes detected at Dome. These results argue against replication timing state being the primary determinant of the Dome-stage transcriptional changes observed in rif1 mutant embryos. We have revised the Results to describe this analysis and added Figure 5—figure supplement 2.

      (2) It is of interest whether the Rif1-mediated regulation of transcription and replication are mediated by a common mechanism, e.g. through alteration of chromatin structures. Close look at the data in Figure 3D indicates that some genome segments convert replication timing or undergo significant changes of replication timing. It would be informative to know whether these segments (Rif1-regulated replication domains) are associated with the genes whose expression change upon rif1 knockout.

      We thank the reviewer for this insightful suggestion. We agree that an association between Rif1-dependent replication timing changes and Rif1-dependent transcriptional changes would be informative, and we considered this analysis. We attempted to identify Rif1-regulated replication timing domains using the same approach that we previously used to define developmentally regulated timing domains. However, the effect of Rif1 loss differed qualitatively from the developmental timing switches described in our prior work. Rather than producing a limited set of discrete timing-domain transitions, Rif1 loss caused a broad reduction in the dispersion of replication timing values across the genome, consistent with a general flattening of the timing profile. Under these conditions, an unbiased domain-calling approach preferentially identifies genomic regions with the most extreme early or late timing values in wild-type embryos, because these regions show the largest shift toward the mean in rif1 mutants. Thus, the resulting “Rif1-regulated replication domains” largely reflect the strongest wild-type timing domains rather than a discrete set of Rif1-specific regulatory intervals. For this reason, we do not think that assigning differentially expressed genes to such domains would provide a meaningful test of whether Rif1 regulates transcription and replication timing through a common local mechanism. Instead, we have now added a gene-centred analysis comparing replication timing values for genes with increased, decreased, or unchanged transcript abundance at Dome (Figure 5—figure supplement 2), which directly addresses whether transcriptionally affected genes are associated with early- or late-replicating regions.

      (3) Replication is analyzed only by timing analysis. Authors need to analyze frequency of origin firing and replication fork rate by DNA fiber analyses to see whether they are affected by rif1 knockout at various stages of development.

      We agree that measuring origin firing frequency and replication fork rate would provide valuable additional information about how Rif1 loss affects the replication program. However, performing DNA fibre analyses across multiple zebrafish developmental stages and genotypes would require substantial optimization and experimental expansion beyond the scope of the current revision. The current study was designed to measure genome-wide replication timing and transcript abundance across developmental stages, rather than single-molecule replication dynamics. We therefore have not added DNA fibre experiments. Instead, we have revised the Discussion to acknowledge this limitation and to clarify that replication timing reflects the combined effects of origin usage, fork progression, fork directionality, and fork stability. We added the following text to the Discussion:

      “A further limitation of this study is that we concentrated on replication timing without directly measuring other features of the replication program that contribute to this timing. These features include origin usage, replication fork spacing, fork directionality, fork progression, and fork stability. A more comprehensive understanding of how Rif1 loss affects these parameters will be important for defining the relationship between Rif1-dependent changes in replication timing and transcription.”

      Figure 4B, D and F: I did not see the blue lines which represent preMBT in the panels shown.

      We thank the reviewer for identifying this error. The pre-MBT data were not intended to be shown in Figures 4B, 4D, and 4F. We have corrected the figure legend by removing the reference to the blue pre-MBT line.

      Line 270: Figure 4G should be Figure 6G.

      We thank the reviewer for identifying this error. We have corrected the figure reference from Figure 4G to Figure 6G.

      No description of Figure 6E and 6F in the main text.

      We thank the reviewer for noting this omission. We have added text to the Results describing Figures 6E and 6F. The revised text explains that Dome Up-DEGs are normally upregulated from pre-MBT to Shield stages but show earlier upregulation in rif1 mutant embryos, whereas Dome Down-DEGs normally decrease between Dome and Shield stages but show earlier reduction in mutant embryos.

      Reviewer #2 (Recommendations For The Authors):

      (1) This study is an extension of the lab’s previous work which established the wild-type genome-wide replication timing pattern during zebrafish development. The experimental details and analysis are described in the methods, but the general strategy is sometimes treated very cursorily. A non-expert can only understand parts of it by going back to the Seifert study.

      We thank the reviewer for pointing this out. We agree that the replication-timing strategy should be understandable without requiring readers to consult our previous study. We have revised the manuscript to explain the general logic of the assay more clearly. Specifically, we now state that replication timing was inferred from copy-number differences between S-phase and G1-phase genomic DNA: genomic regions that replicate early in S phase are enriched in S-phase DNA relative to G1 DNA, whereas later-replicating regions are less enriched. We also clarified that pre-MBT, dome, and shield embryos were treated as S-phase samples because most cells are in S phase at these stages, whereas nuclei from bud and 24 hpf embryos were sorted by DNA content to isolate G1 and S-phase fractions. These additions make the experimental design and interpretation of the replication-timing profiles clearer in the main text and Methods.

      Figure 2 is meant to document developmental delay in early embryos, but the differences between the single wt and mutant examples in 2D are poorly described and labeled. Most readers will be unfamiliar with the specifics of zebrafish development. There is also no quantification of this developmental phenotype, and that quantification should be included along with better labeling and description of 2D.

      We thank the reviewer for pointing this out. We agree that the developmental delay shown in Figure 2D required clearer explanation and quantification for readers who are less familiar with zebrafish gastrulation. We have revised the Results to explain that epiboly is the process by which the blastoderm and yolk syncytial layer move toward the vegetal pole to envelop the yolk cell, and that zebrafish gastrulation stages are commonly described by the percentage of yolk coverage. We also added quantification of this phenotype. At 10 hpf, most wild-type embryos had completed epiboly, whereas most rif1 mutant embryos had not: 18 of 24 wild-type embryos, but only 2 of 24 mutant embryos, had reached 100% yolk coverage. By 11 hpf, all wild-type and mutant embryos had completed epiboly. These revisions clarify that rif1 mutant embryos show an approximately 1-hour delay in epiboly completion rather than a persistent arrest in gastrulation.

      (3) The presentation could be greatly improved with additional information about the experimental approach and display. As written, the text and figure legends assume readers are intimately familiar with replication timing experiments, zebrafish development, and differential gene expression analysis. Most of the figure legends are not sufficient to understand the figures themselves, and the necessary information is also not always in the results. An example is Figure 3 which is not well described (other than the PCA plots); the term “lag” which is the x-axis in 3C is not defined.

      We thank the reviewer for this helpful comment. We agree that several aspects of the replication-timing analysis required clearer explanation for readers who are less familiar with replication-timing experiments. We have revised the Results to explain the logic of the replication-timing assay more clearly and have added a more detailed description of the autocorrelation analysis in Figure 3C. Specifically, we now explain that autocorrelation measures how similar replication-timing values are across increasing genomic distances along the same chromosome, providing a quantitative readout of the peak-and-valley structure of the timing profile. We also clarified that increasing autocorrelation across hundreds of kilobases reflects the progressive establishment of broader replication-timing domains during development. In addition, we changed the x-axis label in Figure 3C from “lag” to “Genomic distance (Mb).” Together, these changes should make the experimental approach and display easier to understand without requiring readers to consult our previous replication-timing study.

      Figure 4 is generally poorly described and labelled (4B, D, and F graph legends indicate preMBT in the data, but there are no blue lines on the graphs), and Figures 6 and 7 are quite busy.

      We thank the reviewer for pointing this out. We agree that the Figure 4 legend incorrectly described the data shown in panels B, D, and F. The pre-MBT data were not intended to be plotted in these panels, and we have removed the corresponding reference from the figure legend. We recognize that Figures 6 and 7 contain several analyses, but we have retained the current organization because the panels in each figure address a connected set of questions. Figure 6 summarizes how Rif1 loss affects abundance of developmentally regulated transcripts, whereas Figure 7 extends this analysis by directly measuring nascent transcription using SLAM-seq.

      Reviewer #3 (Recommendations For The Authors):

      I do not think any additional experiments are required to justify the authors’ claims. Well done! However, for readers’ benefit, I propose the following changes or adding more explanations:

      (1) Page 2, line 86: I guess “single copy” means “single copy per haploid”. Better to clarify this point.

      We thank the reviewer for this helpful clarification. The reviewer is correct that “single copy” refers to a single copy per haploid genome. We have revised the text to state that the zebrafish genome has a single copy of the rif1 gene per haploid genome.

      (2) Related to the data presented in Figure 2C, do you have an explanation for why sex determination is affected in the heterozygotes, despite the change in Rif1 expression being subtle (Figure 1C)?

      We thank the reviewer for raising this point. We agree that the reduction in whole-embryo rif1 mRNA levels in heterozygotes appears modest relative to the sex-ratio phenotype. At present, we can only speculate about the basis for this difference. One possibility is that whole-embryo mRNA measurements do not accurately reflect Rif1 abundance in the specific cell populations that influence zebrafish sex determination, such as primordial germ cells or their supporting somatic cells. We have therefore avoided making a strong mechanistic conclusion from the heterozygous phenotype.

      (3) Related to the data presented in Figure 2D, did you observe a delay in heterozygotes?

      We thank the reviewer for this question. We have not quantitatively analyzed epiboly progression in heterozygous embryos. However, we did not observe an obvious developmental delay in heterozygotes during early development. The delay shown in Figure 2D was observed in maternal-zygotic rif1 homozygous mutants.

      (4) Figure 3D: it is not easy to distinguish WT and mutant lines, particularly for the Bud stage. Please consider changing the colour schemes or other aspects. For example, making colour lines thinner may help.

      We thank the reviewer for this helpful suggestion. We agree that the wild-type and mutant profiles in Figure 3D, particularly at the bud stage, were difficult to distinguish in the original version. We have revised Figure 3D by reducing the line width of the colored profiles, which improves the contrast between the wild-type and mutant traces.

      (5) Figure 3E: Could you avoid overlapping of WT and mutant plots?

      We thank the reviewer for this suggestion. We considered separating the wild-type and mutant density plots in Figure 3E, but we have retained the overlaid format because the purpose of this panel is to directly compare the distributions of replication timing values between genotypes at each developmental stage. Overlaying the plots makes the reduced dispersion of timing values in the rif1 mutants easier to visualize relative to the corresponding wild-type distribution.

      (6) Figure 4C and 4E: the point legends (WT and mutant) do not match the points used in the graph.

      We thank the reviewer for noting this potential source of confusion. In Figures 4C and 4E, point shape indicates genotype, with open squares representing wild-type samples and open circles representing rif1 mutant samples. Point color indicates developmental stage. We used separate visual encodings for genotype and stage to avoid a large legend containing every genotype-stage combination. To make this clearer, we have revised the figure legend to state explicitly that point shape denotes genotype and point color denotes developmental stage.

      (7) Figure 4D: Very difficult to recognise 24 hr mutant line. Please improve the way there are shown.

      We thank the reviewer for this helpful suggestion. We agree that the 24 hpf mutant profile in Figure 4D was difficult to distinguish in the original version. We have revised the figure by changing the appearance of the mutant lines to make them more visible while preserving the stage color scheme.

      (8) Related to data presented in Figure 4B. Is it possible to show a statistical evaluation of all (or a reasonably large number of samples from) DARs?

      We thank the reviewer for this suggestion. Figure 4A already provides a genome-wide analysis of the DAR set shown by example in Figure 4B. Specifically, Figure 4A plots the change in replication timing from shield to 24 hpf for all 2,498 putative enhancer-associated DARs in both wild-type and rif1 mutant embryos. The strong correlation between wild-type and mutant values indicates that DAR-associated timing changes are largely preserved in rif1 mutants. Because all DARs used for this analysis are included in the scatterplot, we did not add a separate statistical analysis of selected examples from Figure 4B.

      (9) Page 8, line 220: It is unclear what “all” means. Is it all the available replication timing values genome-wide? Please clarify.

      We thank the reviewer for noting this ambiguity. In this sentence, “all” refers to all genome-wide replication timing values calculated from the genomic windows used in our replication timing analysis. We have revised the text to make this clearer.

      (10) Figures 6C and 6D: Colour labels are too dark and it is almost impossible to read texts inside. Please reconsider the colour scheme.

      We thank the reviewer for pointing this out. We agree that the labels in Figures 6C and 6D were difficult to read because of insufficient contrast. We have changed the text colour inside the colored boxes to white to improve legibility.

      (11) Related to overall transcription studies: Is there any sign that Rif1 mutation affects the transcription of genes involved in sex determination?

      We thank the reviewer for raising this interesting question. We have not specifically analyzed whether genes involved in sex determination are differentially expressed in the early embryonic transcriptome data. Because zebrafish sex determination occurs substantially later than the embryonic stages analyzed here, and likely depends on specific cell populations such as primordial germ cells and supporting gonadal somatic cells, we do not think the current whole-embryo RNA-seq data can directly resolve this question. We therefore avoid drawing a mechanistic connection between the early transcriptional changes and the adult sex-ratio phenotype. Determining whether Rif1 mutation affects transcription in the cell populations that regulate zebrafish sex determination will be an important direction for future work.

    1. eLife Assessment

      In this manuscript, the authors analyse the nanoscale localisation of α5β1 and αVβ3 integrins in integrin adhesion complexes (IAC) by dual-colour STORM and DNA-Paint and assess the spatial organisation at the nano and mesoscale of their main adaptors (paxillin, talin and vinculin). This is an important work that provides detailed analyses that reveal how elements of these complex structures are really organised at the nanoscale, an essential perspective for a better understanding of how IACs function and regulate mechanotransduction processes. The evidence presented is convincing, using complementary super-resolution imaging techniques and subsequent computational modelling that enabled a quantitative assessment of the resulting data.

    2. Reviewer #1 (Public review):

      Summary:

      In recent years, it becomes increasingly evident how beautifully intricate IAC are at the nanoscale. Studies like the one presented here that shed light on the precise inner organisation of IAC are thus quite important and relevant to obtain better in-depth understanding of IAC functioning and the contribution of different integrin subtypes to cell adhesive and mechanotransductive processes.

      Interestingly, the authors found a distinct localisation of α5β1 and αVβ3 integrin nanoclusters within focal adhesion of human fibroblasts, with α5β1 integrin nanoclusters being at the periphery of IAC and αVβ3 integrin nanoclusters randomly distributed. Furthermore, a surprisingly high percentage of inactive integrins within IAC and relatively low spatial integrin colocalisation with adaptor proteins has been shown.

      Strengths:

      This is a very thoroughly performed STORM-based assessment of the nanodistribution of α5β1 and αVβ3 nanoclusters within IAC (and outside). The image quality is outstanding, and the authors have meticulously executed the experiments and the image analyses.

      Weaknesses:

      The only weakness is maybe that the manuscript remains descriptive. However, the high quality of the "description" of the nano-organisation of IAC by this scrupulous study is really important to better understand the inner workings of IAC. It provides a very solid foundation to look deeper into the (patho)physiological implications of this organisation, see recommendations (which are rather suggestions in this case).

      Comments on revision:

      The authors meticulously addressed all my questions and suggestions. I want to thank the authors for an exemplary revision.

    3. Reviewer #2 (Public review):

      Summary:

      In this study, dual-color super-resolution microscopy analysis was performed to study the co-operation between integrins and focal adhesion proteins in human fibroblast cells. The study focused on two integrins which have been previously found to be mainly responsible for focal adhesions, namely α5β1 and αvβ3.

      Specifically, the study tried to shed light on the nanoclustering of integrins in focal adhesions.

      In the current study, more integrin nanoclusters were observed in focal adhesions compared to other cell-matrix adhesion structures. The study revealed that both α5β1 and αvβ3 form nanoclusters and those appear segregated from each other. While αvβ3 nanoclusters organize randomly inside focal adhesions regardless of their activation state, α5β1 nanoclusters, and particularly the nanoclusters containing β1-integrin in active conformation preferentially organized at the edges of focal adhesions. The nanoclusters formed by each integrin were similar in size.

      Cytoplasmic adapter proteins appeared less in nanocluster assemblies, suggesting that integrin nanoclusters are also forming without the studied cytoplasmic adapter proteins (talin, vinculin, paxillin). Active integrins were identified with help of conformation-specific antibodies, and those enabled to study the colocalization between integrins and their cytoplasmic adapter proteins. This analysis revealed that activated integrins are strongly engaged with adapter proteins

      Strengths:

      The study stems from the thorough computational modelling of the nanoclusters, which enables quantification of the behavior of the clusters, including their mesoscale distribution.

      The study strengthens the view that α5β1 and αvβ3 have specific functions in focal adhesions, α5β1 nanoclusters localizing preferentially on focal adhesion edges. The study also revealed that nanoclusters localized at the edges of focal adhesion were enriched for talin and paxillin but not for vinculin.

      Analysis of adaptor protein nanoclusters (paxillin, talin, and vinculin) revealed that all adapter protein nanoclusters studied here close to active β1 nanoclusters are enriched on the focal adhesion edge region, whereas integrin adaptor nanoclusters far from active β1 appear to be more uniformly distributed.

      Importantly, the current study suggests that integrin subtype-specific nanoclusters are not only present at early stage of adhesion formation, but integrin nanoclusters remain segregated from each other also in mature focal adhesions, maintaining their sizes and number of molecules.

      Interestingly, the study revealed that selected cytoplasmic adaptors (paxillin, talin and vinculin), also form nanoclusters of similar size and number of single molecule localizations as the integrins, regardless of whether they locate inside or outside focal adhesions. The adapter nanoclusters are enriched in the focal adhesion "belt", colocalizing with the active α5β1 integrin nanoclusters.

      Weaknesses:

      The current study is highly dependent on the antibodies. It is possible, that antibodies, containing two binding sites for antigen, influence the nanoscale organization (and also activation) of the receptors. Control experiments to study possible contribution of antibodies for the measured outcome should be performed to verify the main findings. One possible approach could be to use fluorescently tagged integrins available. Alternatively, integrins (or adapter proteins) could be tagged with small ligand and detected using monovalent binder.

      Only a limited number of integrin adapter proteins were investigated. Given the high number of identified adapter proteins, this is an understandable choice. However, it would be fascinating to understand if the nanoclusters of inactive integrins are dominantly bound with certain adapter protein, such as tensin.

      Comments on revision:

      The authors addressed the concern related to the use of antibodies and secondary antibodies by performing DNA-PAINT experiment, which revealed highly similar results as obtained with conventional antibodies.

    4. Author response:

      The following is the authors’ response to the original reviews.

      We have addressed all the concerns and recommendations by the reviewers, in particular, the requested control experiments using alternative super-resolution microscopy approaches and analysis of the data using Voronoi tessellation in addition to DBSCAN, as requested by reviewer 3. We also provide additional data on tensin3 as suggested by reviewer 2. Finally, to provide a first insight on the role of mechanical forces in the distribution of integrin nanoclusters inside FAs as recommended by reviewer 1, we have performed experiments at different cell seeding times where it is known that FA maturation over time requires mechanical forces.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      In recent years, it has become increasingly evident how beautifully intricate IAC are at the nanoscale. Studies like the one presented here that shed light on the precise inner organisation of IAC are thus quite important and relevant in order to obtain a better in-depth understanding of IAC functioning and the contribution of different integrin subtypes to cell adhesive and mechanotransductive processes.

      Interestingly, the authors found a distinct localisation of α5β1 and αvβ3 integrin nanoclusters within focal adhesion of human fibroblasts, with α5β1 integrin nanoclusters being at the periphery of IAC and αvβ3 integrin nanoclusters randomly distributed. Furthermore, a surprisingly high percentage of inactive integrins within IAC and relatively low spatial integrin colocalisation with adaptor proteins has been shown.

      Strengths:

      This is a very thoroughly performed STORM-based assessment of the nanodistribution of α5β1 and αvβ3 nanoclusters within IAC (and outside). The image quality is outstanding, and the authors have meticulously executed the experiments and the image analyses.

      We are grateful to the reviewer for acknowledging the strengths of our study.

      Weaknesses:

      The only weakness is maybe that the manuscript remains descriptive. However, the high quality of the "description" of the nano-organisation of IAC by this scrupulous study is really important to better understand the inner workings of IAC. It provides a very solid foundation to look deeper into the (patho)physiological implications of this organisation, see recommendations (which are rather suggestions in this case).

      We thank the reviewer for their feedback and have addressed their recommendations in our updated manuscript and accompanying reply (see recommendations to the authors). In summary, we have now performed experiments at different seeding times as FA maturation requires mechanical forces, and enquired whether forces might play a role in establishing the spatial distribution of the two different integrins within more mature IACs. The results are now shown as new Fig. 2 and discussed in pages 9 and 10. In addition, in order to get a first insight into the biological implications of our findings we performed dual-colour super-resolution experiments of tensin-3 and α<sub>5</sub>β<sub>1</sub> in FAs, as tensin-3 has been implicated in fibronectin fibrillogenesis. The results are now shown in Fig. S8 and we discuss their potential implications in pages 22 and 23 of the revised manuscript (see more details in the reply to the recommendation to the authors).

      Reviewer #2 (Public review):

      Summary:

      In this study, dual-color super-resolution microscopy analysis was performed to study the co-operation between integrins and focal adhesion proteins in human fibroblast cells. The study focused on two integrins which have been previously found to be mainly responsible for focal adhesions, namely α5β1 and αvβ3.

      Specifically, the study tried to shed light on the nanoclustering of integrins in focal adhesions.

      In the current study, more integrin nanoclusters were observed in focal adhesions compared to other cell-matrix adhesion structures. The study revealed that both α5β1 and αvβ3 form nanoclusters, and those appear segregated from each other. While αvβ3 nanoclusters organize randomly inside focal adhesions regardless of their activation state, α5β1 nanoclusters, and particularly the nanoclusters containing β1-integrin in active conformation, preferentially organized at the edges of focal adhesions. The nanoclusters formed by each integrin were similar in size.

      Cytoplasmic adapter proteins appeared less in nanocluster assemblies, suggesting that integrin nanoclusters are also forming without the studied cytoplasmic adapter proteins (talin, vinculin, paxillin). Active integrins were identified with the help of conformation-specific antibodies, and this enabled us to study the colocalization between integrins and their cytoplasmic adapter proteins. This analysis revealed that activated integrins are strongly engaged with adapter proteins.

      Strengths:

      The study stems from the thorough computational modelling of the nanoclusters, which enables quantification of the behavior of the clusters, including their mesoscale distribution.

      The study strengthens the view that α5β1 and αvβ3 have specific functions in focal adhesions, α5β1 nanoclusters localizing preferentially on focal adhesion edges. The study also revealed that nanoclusters localized at the edges of focal adhesion were enriched for talin and paxillin but not for vinculin.

      Analysis of adaptor protein nanoclusters (paxillin, talin, and vinculin) revealed that all adapter protein nanoclusters studied here close to active β1 nanoclusters are enriched on the focal adhesion edge region, whereas integrin adaptor nanoclusters far from active β1 appear to be more uniformly distributed.

      Importantly, the current study suggests that integrin subtype-specific nanoclusters are not only present at an early stage of adhesion formation, but integrin nanoclusters remain segregated from each other also in mature focal adhesions, maintaining their sizes and number of molecules.

      Interestingly, the study revealed that selected cytoplasmic adaptors (paxillin, talin, and vinculin), also form nanoclusters of similar size and number of single molecule localizations as the integrins, regardless of whether they locate inside or outside focal adhesions. The adapter nanoclusters are enriched in the focal adhesion "belt", colocalizing with the active α5β1 integrin nanoclusters.

      We are grateful to the reviewer for acknowledging the strengths of our study.

      Weaknesses:

      The current study is highly dependent on the antibodies. It is possible that antibodies containing two binding sites for antigen influence the nanoscale organization (and also activation) of the receptors. Control experiments to study the possible contribution of antibodies to the measured outcome should be performed to verify the main findings. One possible approach could be to use fluorescently tagged integrins available. Alternatively, integrins (or adapter proteins) could be tagged with a small ligand and detected using a monovalent binder.

      We understand the concern of the reviewer regarding the use of antibodies for imaging. Nevertheless, we would like to clarify that antibody labelling has always been performed after cell fixation, precluding potential cross-linking artefacts due to protein mobility and avoiding unwanted receptor activation.

      Nevertheless, and although it is highly unlikely to happen in fixed cells, there could be two potential sources of antibody (Ab) labelling artefacts. As the reviewer noted, a primary Ab containing two binding sites could bind to two adjacent proteins (within ~10 nm from each other), potentially underestimating the stoichiometry of the nanoclusters, i.e., number of receptors or proteins per nanocluster. However, in our manuscript we never attempted to provide an estimation of the nanocluster stoichiometry, as it is highly challenging (and prone to artefacts) to provide quantification of the number of proteins using super-resolution-based single-molecule localisation methods which rely on the stochastic blinking of individual fluorophores.

      A second source for potential artefacts comes from the use of the secondary Ab, which (albeit unlikely) could bind to two different primary Abs. To exclude this potential artefact, we performed super-resolution imaging using DNA-PAINT as a different imaging strategy. In this case, the DNA docking site is site-specifically coupled to one camelid single-domain Ab (sdAB), having a much smaller size as compared to a secondary Ab, reducing therefore linkage error and increasing the accessibility of primary Ab-labelled proteins. These new data are included now in Fig. S4. As can be observed, no differences in terms of nanocluster sizes and/or compositions were observed for any of the proteins investigated using DNA-PAINT as compared to our initial STORM data. These control experiments thus rule out any potential artefacts introduced by the secondary Ab (for more details, please see the reply to the recommendations for authors section).

      Only a limited number of integrin adapter proteins were investigated. Given the high number of identified adapter proteins, this is an understandable choice. However, it would be fascinating to understand if the nanoclusters of inactive integrins are dominantly bound with a certain adapter protein, such as tensin.

      We fully agree with the reviewer and have now performed dual-colour super-resolution STED microscopy of α<sub>5</sub>β<sub>1</sub> and tensin-3 on HFF cells seeded for 24 hours. Interestingly, instead of being an integrin inactivator, we found that tensin-3 is also highly enriched at the FA periphery where a large fraction of active β<sub>1</sub> integrins are located, suggesting that at these particular regions, active β<sub>1</sub> could be either engaged to talin (as shown in our original data) or to tensin-3 (our new data shown in Fig. S8). We provide more details of our answer in the section of “recommendation to the authors”. Additional experiments, which in our opinion fall outside of the scope of this work, would be necessary to identify other potential integrin inactivator partners, but certainly a topic of future interest to our group.

      Reviewer #3 (Public review):

      Summary:

      In their study, the authors reveal using dual-color super-resolution STORM microscopy modality and immunolabeling in fixed adherent cells, that β1 and β3 integrins as well as adaptors (paxillin, talin and vinculin) are all organized in nanoclusters of similar size (50nm) and molecular density (20 copy number) inside FAs but also outside. Using activityspecific immunolabeling of β1 and β3 integrins, they revealed that active integrin subpopulations were both clustered but in distinct exclusive nano-aggregates in agreement with Spiess et al. (2018). Once more, the "active" integrin nanoclusters displayed similar properties in terms of size and molecular density, suggesting that molecular organization in nanoclusters is an intrinsic property of integrins in plasma membrane multimerizing independently of their location (inside or outside FAs), their level of activation, or their connection to the cytoskeleton. Then the authors followed up by analyzing at the mesoscale how these "universal" nanoclustered adhesive units are distributed spatially. Inspecting the surface density of nanoclusters revealed that the density of integrin nanoclusters in FAs was 5x larger, compared to integrin nanoclusters outside adhesions. Interestingly, whereas the density of total integrin nanoclusters was 2-4x larger than adaptor nanoclusters, the density of "active" integrin nanoclusters stoichiometrically matches that of talin and vinculin nanoclusters, and was slightly outnumbered by paxillin nanoclusters. These findings suggest that inside FAs, among the total number of integrin nanoclusters, the subset of "active" integrin nanoclusters could be engaged with "adaptor" nanoclusters on a 1:1 ratio. Using analysis of the nearest neighbor distance (NND) between distinct integrin clusters and each of the adaptors, the authors report that they found negligible spatial colocalization of integrins with these adaptor proteins and that spatial segregation is essentially determined by the density of nanoclusters within the FAs. As authors reported that α5β1 and αvβ3 do not intermix at the nanoscale, the authors finally highlighted how α5β1 and αvβ3 distinct nanoclusters are differently organized and segregated inside FAs. Adapting the NND analysis in order to inspect how far the nanoclusters are from the edges of FAs they are located in, authors revealed that α5β1 but not αvβ3 integrin nanoclusters are enriched on FA edges and that similar FA edge-enriched distribution for "active" α5β1 and adaptor protein nanoclusters was found for talin and paxillin but not vinculin. The latter results suggest that FA edges could constitute multiprotein hubs for enhanced colocalization and activation for α5β1 integrin nanoclusters and adaptors such as talin and paxillin. Unfortunately NND analysis could not confirm this enhanced colocalization hypothesis.

      General Assessment:

      While the study presents some valuable findings, it reads currently as a compilation of intriguing but preliminary observations derived primarily from a single methodology (dual-color STORM and DBSCAN clustering analysis). As the initial findings often lack confirmation through additional data analysis (such as the NND analysis the authors used), there's a critical necessity to bolster the methodological approach. This should involve replicating the main findings using alternative single-molecule super-resolution techniques (such as quantitative DNA-PAINT) or employing different clustering analytical tools (such as voronoi-tessellation). Furthermore, the manuscript feels incomplete, focusing solely on describing molecular organization without offering substantial insights into how these observations correlate with the regulation, activation, and functionality of integrins at the cellular level.

      We appreciate the comment of the reviewer and have taken their recommendation to heart in order to validate our methodology. In summary, we have now performed extensive DNA-PAINT to replicate most of our initial findings obtained by STORM, as requested by the reviewer. In addition, as a different super-resolution imaging strategy, we have also used STED microscopy to confirm the nanoclustering of integrins and some of the adaptors demonstrating now, by means of three different super-resolution techniques, that both integrins and their adaptors form nanoclusters of similar size and composition, regardless of whether they are inside or outside FAs. We have included these data as Figs. S3 and S4 and discussed the results in pages 8-9 of the main manuscript.

      Regarding the use of an alternative analysis for the data, we have now used the Voronoi tessellation algorithm to re-analyse our STORM data, as requested by the reviewer. The results of the analysis, which render similar sizes and number of localizations as obtained by DBSCAN, are now included in Fig. S5 and mentioned in page 8 of the main manuscript.

      The manuscript presents extensive datasets and utilizes methodologies in which the investigators demonstrate expertise. Nevertheless, there's uncertainty regarding the novelty and broad appeal of the findings. For instance, the observation of integrin nanoclustering has been previously reported in several publications (e.g., Changede et al., Dev Cell 2015; Spiess et al., JCB 2018; Fujiwara et al., JCB 2023). Similarly, the accumulation of specific proteins at the periphery of FAs has been documented elsewhere (e.g., Sun et al., NCB 2016; Stubb et al., NatComm 2019; Nunes-Vicente TCB 2023), as well as the differential dynamic organization of α5β1 and αvβ3 integrins inside FAs (e.g., Rossier et al., NCB 2012). Beyond the universal organization of adhesive proteins, there's a need to identify novel insights that significantly advance the field. One potential avenue could involve pinpointing the molecular determinant controlling the FA edge enrichment of active α5β1 integrins and talin nanoclusters. For instance, could there be an interplay between α5β1 and αvβ3 integrin nanoclusters visible on one's organisation when suppressing the other using deletion (KO) or depletion (SiRNA)? Also, could KANK, which also exhibits enrichment and regulates talin activity (e.g., Sun et al., NCB 2016), play a role in this process? Identifying the molecular players that regulate even partially the mesoscale organization of nanoclusters of proteins would really benefit the breadth of this manuscript.

      We could not agree more with the reviewer and in fact, we are currently investigating the mechanisms that control the enrichment of α<sub>5</sub>β<sub>1</sub> and adaptors at the edges of FAs. However, considering the amount of work needed to determine the spatiotemporal organization of other molecular players using super-resolution imaging constitutes a major tour de force.

      To get a first insight into the process of active α<sub>5</sub>β<sub>1</sub> enrichment at the FA edges, we hypothesised that mechanical forces exerted by the actomyosin machinery could influence the lateral distribution of both integrin subsets (α<sub>5</sub>β<sub>1</sub> and α<sub>v</sub>β<sub>3</sub>) inside FAs. Since FA maturation and strengthening over time requires mechanical forces, we performed experiments at different cell seeding times (90 min, 3 hours and 24 hours) and used STORM imaging to follow the evolution of integrin nanoclustering in time as well as their spatial distributions inside FAs. Interestingly, while nanoclustering of both integrin sub-sets inside FAs is not influenced by seeding times, their lateral distribution was markedly different, with α<sub>5</sub>β<sub>1</sub> nanocluster distribution being already established at earlier seeding times, while α<sub>v</sub>β<sub>3</sub> nanocluster distribution appeared as rather random at earlier seeding times and progressively organized reaching a well-defined lateral spacing at 24 hours of spreading time. These initial data strongly suggest that mechanical forces might play a role in the distinct lateral distribution of both subsets of integrin nanoclusters over time. We have now included these data as new Fig. 2 of the revised manuscript and discuss the results in the associated text (pages 9 and 10). We also discuss potential avenues for further research along the directions suggested by the reviewer.

      In addition, since it has been recently shown that tensin-3 interaction with talin drives the formation of fibronectin-associated fibrillar adhesions (Atherton et al, J Cell Biol 2022) which are enriched in β<sub>1</sub> integrins, we performed dual-colour super-resolution STED microscopy of β<sub>1</sub> and tensin-3 on HFF cells seeded for 24 hours. Interestingly, our initial data show co-enrichment of both tensin-3 and active β<sub>1</sub> nanoclusters at the FA periphery, suggesting that at these particular regions, active β<sub>1</sub> could be either engaged to talin (as shown in our original manuscript) or to tensin. Our current working hypothesis is that α<sub>5</sub>β<sub>1</sub> enrichment at the FA periphery serves to facilitate the translocation of α<sub>5</sub>β<sub>1</sub> integrins from FAs to fibrillar adhesions, most probably in a talin-tensin-dependent manner. We have now included these data as Fig. S8 and accompanying discussion in pages 22 and 23 of the revised manuscript.

      Echoing the previous concern, the manuscript described a novel and rather surprising finding related to molecular clustering of adhesion proteins. Indeed, the fact that nanoclusters exhibit uniform size and molecular density regardless of the protein type, location, or activation level is indeed surprising and raises many questions about the methodology used to assess molecular clustering. I feel that the description and characterization of integrin nanoclusters appear incomplete and need to be expanded by comparing different analytical strategies for protein clustering. Furthermore, a lack of the manuscript in its actual form concerns the quantification of integrin numbers inside the observed nanoclusters. I agree that the path from optical microscopy to protein stoichiometry quantification is hard and full of drawbacks. But the authors do not fully address these issues that are extremely important when discussing protein nanoclustering. This quantitative aspect should be discussed.

      We appreciate the comment of the reviewer as indeed, the existence of “universal” nanoclusters is intriguing. Recently, together with Prof. S. Mayor we have written a short review in Curr. Opin. Cell Biol 2024 proposing that nanoclustering constitutes a molecular-scale organisation principle that governs cellular information flow at the plasma membrane. Our proposal is supported by an extensive number of recent papers showing that most cell membrane receptors and downstream signalling components are organized as pre-assembled nanoclusters. We posit that these nanoclusters serve as modular units whose concatenation in a specific spatiotemporal sequence leads to distinct signalling outputs. Thus, the existence of universal nanoclusters of integrin receptors and adaptors is indeed intriguing but not surprising to us.

      In any case, the concern of the reviewer is well-taken, and as mentioned above, we have used a different algorithm to detect and quantify nanoclustering, obtaining similar values using either Voronoi tessellation or DBSCAN approaches. These data are now included as Fig. S5 in the manuscript.

      Regarding the quantification of integrin numbers inside the observed nanoclusters, we agree with the reviewer that determining protein stoichiometry using single-molecule localization microscopy or STED remains a major technical challenge and is highly prone to artefacts. For this reason, we refrain from making claims about absolute protein numbers per nanocluster. Our relative comparison of nanoclustering among the different proteins investigated is thus exclusively based on the number of single-molecule localisations contained in each nanocluster which is a fair approach since we always use the same reporter fluorophore and maintain similar excitation conditions throughout our experiments. We have now included a few lines on page 9 regarding quantification of the absolute protein numbers inside the nanoclusters and further discuss in the revised manuscript the limitations of single-molecule localisation methods towards the stoichiometry determination of the nanoclusters (see page 20 of the revised manuscript).

      First, it is crucial for the authors to carefully examine and discuss in their manuscript whether there are any potential biases or limitations in the experimental techniques (dual-color STORM) or data analysis methods employed (DBSCAN). Second, the authors did not in the current manuscript, but should provide control samples to demonstrate the sensitivity and dynamic range of their experimental strategy.

      As already mentioned, we have validated the STORM data using both DNA-PAINT and STED and, validated our data analysis obtained with DBSCAN using the Voronoi tessellation algorithm. See Figs. S3, S4 and S5. In terms of sensitivity and dynamic range of our methodology: our set-up has single-molecule detection sensitivity which is demonstrated by the fact that we observe and detect discrete blinking events, a property of single-molecule fluorescence emission and key ingredient to super-resolution single-molecule localisation microscopy. The dynamic range (if we understand correctly the question of the reviewer) is given by the number of frames used to accumulate single-molecule localisations. In our case, we stop acquisition after we deplete most of the single-molecule spots in the imaging view, which typically occurred after 70,000 frames acquisition, as correctly mentioned in the material & methods section.

      In STORM images displayed in Figure S1, the authors highlighted localization clusters detected by DBSCAN as a signature for integrin nanoclusters. But the authors do not discuss the localization spots that were not detected by DBSCAN. Could they be individual integrins? And if so, they should also be considered as useful information? This brings me to another related technical question about how DBSCAN handles the case where fluorescent molecules are blinking. This is important as multiple emissions by a single fluorophore could be detected as a nanocluster of several molecules where it would be an artefact due to the photophysics of the fluorophore. Could the authors comment on these points?

      As mentioned in the original manuscript, between 20-30% of the localizations were not assigned to nanoclusters (Fig. S1H, I) since we imposed a minimum of ten localizations within the radius defined by DBSCAN to be considered as a true nanocluster. This essentially means that regions with less than 10 localizations were not considered in our nanoclustering analysis. However, we cannot be certain as to whether these lower number of localizations correspond to individual integrins, stochastic blinking of the fluorophore or small aggregates containing only a couple of integrins, for the same reasons that we cannot provide quantification of the absolute number of proteins included in each nanocluster: stoichiometry determination by means of single-molecule super-resolution methods is highly prone to artefacts.

      Regarding the concern of how DBSCAN handles fluorophore blinking, the reviewer is completely right as the photophysics of the fluorophore can influence the analysis of the data and the identification of true nanoclusters. To decouple the photophysics of the fluorophore we first assess the number of blinking events within the DBSCAN radius, i.e., number of localizations corresponding to individual antibodies sparsely distributed on the glass surface. In our case, the median values for the two activator-reporter pairs corresponded to 5 localizations for Alexa 405-Alexa647-conjugated Abs and 3 localizations for Cy3-Alexa 647-conjugated Abs (see Fig. 1E). Yet, despite these median values, the number of localizations per individual Ab naturally shows a distribution. Thus, to avoid any overestimation in the degree of nanoclustering, we impose an additional constrain to our analysis and consider true nanoclusters only those ones containing at least 10 localizations. We have now significantly extended the explanation in the main text (see page 6) as well as materials & methods so that it becomes clearer to the reader.

      Also, using isolated and stochastically physisorbed fluorophores (Ab coupled with activator /reporter pairs used in this study) on glass helped define the signature in STORM of a single isolated molecule. To obtain the signature of clustered fluorophores, the authors could use anti-donkey antibodies to cross-link those STORM-specifically labeled Ab as a means to artificially obtain clustered fluorophores. Ultimately, to avoid the bias effect of the glass surfaces on the photophysics of fluorophores and be in the same imaging conditions as for the described nanoclusters, the authors should use model systems composed of multimers of GFP vs. single GFP, immunolabeled with a GFP-binding monoclonal antibody. This will permit evaluation of the cluster signature obtained with DBSCAN analysis of STORM data for single vs. multimers of known stoichiometry. This would constitute an undisputable molecular stoichiometry ruler.

      We appreciate the suggestions of the reviewer. Regarding the potential bias effect of the glass surface on the photophysics of the fluorophores we would like to clarify that the “calibration” for the number of blinking events per individual Ab on glass were performed on the same sample containing the cells that we image, so that we maintain exactly the same experimental and imaging conditions avoiding any potential artefacts. To our understanding this approach is more accurate than performing the calibration on glass substrates and then moving to samples containing the cells. This information is now contained in page 6 of the revised manuscript and in the materials and method section. Once the number of blinking events from individual Abs on glass within the DBSCAN radius are determined, one can then determine the number of localizations within the same DBSCAN radius on other parts of the sample. More localizations within the same DBSCAN radius basically means more molecules, and thus nanoclusters. This approach has been extensively used by other experts in the field as we properly acknowledge in our manuscript (Pageon et al, Mol. Cell. Biol 2016; Spiess et al. J. Cell Biol 2022).

      Using anti-donkey antibodies to cross-link those STORM-specifically labelled Ab in order to artificially obtain clustered fluorophores, as suggested by the reviewer, is indeed a sound approach to retrieve signatures of clustering. Nevertheless, we have preferred not to use this approach because those artificially induced clusters would have very little resemblance to the real nanoclusters and would only allow us to validate the performance of DBSCAN for cluster recognition. As mentioned above, DBSCAN is a well-established algorithm and used by many different experts in the field and thus can be trusted by the community. Instead, and following the recommendation of the reviewer, we now provide results using an alternative cluster analysis algorithm (Voronoi tessellation) reaching similar conclusions regarding the existence of integrin and adaptor nanoclustering inside FAs.

      Finally, the suggestion of using monomeric vs multimeric GFPs to determine the stoichiometry of the nanoclusters is highly appreciated. Indeed, we have used this approach in the past to identify nanoclustering of the chemokine receptor CXCR4 in living T cells (Mol. Cell 2018 and PNAS 2022). However, these experiments are best performed at sub-labelling conditions, which inherently underestimate the degree of nanoclustering. Combining GFPs with PALM to enable super-resolution is another approach but also subject to artefacts regarding the photo-conversion efficiency of GFPs as we reported earlier (Nature Methods 2017) and leading to underestimation of nanocluster stoichiometry.

      In summary, providing nanocluster stoichiometry from single-molecule localisation images remains a major technical challenge and is highly sensitive to methodological assumptions. We have therefore focused here on providing robust evidence for the existence of integrin and adaptor nanoclustering, using three different superresolution approaches and two independent analytical methods for cluster determination.

      Due to the surprising finding of the nanoclusters' "universality", it is imperative for the authors to validate the findings through complementary methodologies and analytical tools. This should involve replication of results using alternative super-resolution techniques (quantitative DNA-PAINT) and exploring different clustering algorithms (VoronoïTesselation) to ensure the robustness and reliability of the observations.

      As already mentioned, we have now performed extensive DNA-PAINT to replicate most of our initial findings obtained by STORM, as requested by the reviewer. In addition, as a different super-resolution imaging strategy, we have also used STED microscopy to confirm the nanoclustering of integrins and some of the adaptors demonstrating now, by means of three different super-resolution techniques, that both integrins and their adaptors form nanoclusters of similar size and composition, regardless of whether they are inside or outside FAs. We have included these data as Figs. S3 and S4 and discussed the results in pages 8-9 of the main manuscript.

      Regarding the use of an alternative analysis for the data, we have now used the Voronoi tessellation algorithm to re-analyse our STORM data, as requested by the reviewer. The results of the analysis, which render similar sizes and number of localizations as obtained by DBSCAN, are now included in Fig. S5 and mentioned in page 8 of the main manuscript.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      This work already, as is, provides significant and novel information on IAC.

      The unexpectedly low spatial colocalisation of integrins with adaptor proteins might indeed be caused by the potentially quite long extension of talin upon force exposure and the ample zones of activity of IAC proteins, imaging the involved proteins in scale, as can be seen in Barnett and Goult (2022, doi: 10.3389/fncel.2022.1014629)? In super-resolution microscopy, this spatial separation might, in fact, become apparent. It would be interesting to see whether lowering the actomyosin contraction by different concentrations of blebbistatin lowers this separation. In general, it would also be interesting to understand whether lowering the forces can disrupt the nano-organisation and the strong separation of the two analysed integrin subtypes. It is true that nascent adhesion formation is force-independent, but maybe the forces play a role in establishing the particular integrin subtype nano-organisation within more mature IAC. I am also aware that a lot of work has already gone into the conclusion of this project.

      We thank the reviewer for these thoughtful comments and suggestions. Our most recent preliminary data (not yet included in this manuscript) indeed indicate that the physical separation of integrin nanoclusters (and adaptors) inside focal adhesions (FAs) is force-dependent. We are currently reproducing these experiments using lipid bilayers of varying viscosities and controlled ligand density to explore how ligand mobility (i.e., equivalent to force exerted from the extracellular side) controls the degree of IAC nanoclustering and their spatial segregation in FAs. This approach is more amenable to super-resolution microscopy as lipid bilayers are quite thin and optically transparent, yet the experiments are still challenging, time-consuming, and thus ongoing.

      To obtain a first hint as to whether forces might play a role in establishing the spatial distribution of the two different integrins within more mature IACs as the reviewer suggests, we have performed experiments at different seeding times (90 min, 3 hours and 24 hours). Our results show that even at earlier times (90 min), when a lower number of mature FAs are established, nanoclustering of integrins and main adaptors are similar to 24 hours. In contrast, and as suggested by the reviewer, the spatial distribution of the different subsets of integrin nanoclusters inside FAs is markedly different as a function of seeding time, with α<sub>5</sub>β<sub>1</sub> nanocluster distribution being already established at 90 min, while α<sub>v</sub>β<sub>3</sub> nanocluster distribution appears rather random at earlier seeding times and progressively organizes reaching a well-defined lateral spacing at 24 hours of spreading time. As FA strengthening over time requires mechanical forces, and α<sub>v</sub>β<sub>3</sub> is preferentially involved in FA strengthening (Roca-Cusachs et al PNAS 2009), these data strongly suggest that forces play a differential role in the lateral distribution of both integrin nanoclusters over time. We have now included these data as a new Fig. 2 in the revised manuscript and discuss the results in the associated text (pages 9 and 10). We also mention in the discussion additional experiments, as suggested by the reviewer, to further substantiate this hypothesis.

      Considering the high quality of the work and the new insight about the inner organisation of IAC, maybe the summary Figure 5 should be elaborated a bit, taking into account e.g. different lengths of extended talin proteins and also the various positions of vinculins on talin proteins (depending on opened cryptic binding sites), as well as the possibility that various actin filaments might be associated with single talins. What I mean is, the authors impressively demonstrate the complexity of IAC nano-organisation, which should be paid more tribute in the concluding figure. The quality of the figure should be adapted to the quality of the work.

      We have adapted Figure 5 (now Figure 6) as suggested by the reviewer.

      I would be curious to hear a bit more about the further speculations of the authors in the discussion, e.g., about why the integrin subunits are organised in this way. Why might the α<sub>5</sub>β<sub>1</sub> be preferentially located in the periphery? What is the potential physiological relevance of this organisation? Is this organisation different in other cell types (have the authors looked at other cells)? Is the organisation lost in pathophysiological situations, such as cancer?

      Although we do not know yet what drives the preferential location of α<sub>5</sub>β<sub>1</sub> nanoclusters to the FA periphery, it is known that Kank2 also exhibits enrichment at the FA periphery, regulates talin activity and it is involved in the formation of α<sub>5</sub>β<sub>1</sub>-enriched fibrillar adhesions (Sun et al, Nature Cell Biol 2016). Thus, it is highly probable that α<sub>5</sub>β<sub>1</sub> enrichment at the FA periphery is a necessary step for their translocation from mature FAs to fibrillar adhesions to then assemble fibronectin into the fibrillar networks as found and needed in connective tissues. Consistent with this idea, we have observed similar α<sub>5</sub>β<sub>1</sub> distribution on other fibroblast cell lines (MEFS), which are the primary cells that produce fibrillar adhesions. Thus, α<sub>5</sub>β<sub>1</sub> nanocluster distribution inside FAs might be physiologically important for the process of fibronectin fibrillogenesis.

      Since it has been documented that tensin is important for fibronectin fibrillogenesis (Pankov et al J Cell Biol 2000) and more recently, it has been shown that tensin-3 interaction with talin drives the formation of fibronectin-associated fibrillar adhesions (Atherton et al, J Cell Biol 2022), we thought to investigate the spatial distribution of tensin-3 and its relationship with α<sub>5</sub>β<sub>1</sub> inside FAs by means of dual colour super-resolution STED microscopy. Interestingly, our initial data on HFF cells seeded for 24 hours show both enrichment of tensin-3 and α<sub>5</sub>β<sub>1</sub> nanoclusters at the edges of mature FAs, supporting our working hypothesis that α<sub>5</sub>β<sub>1</sub> enrichment at the FA periphery serves to translocate α<sub>5</sub>β<sub>1</sub> integrins from FAs to fibrillar adhesions, probably in a talin-tensin-dependent manner. While these initial data are quite exciting, many more experiments that include simultaneous super-resolution mapping of α<sub>5</sub>β<sub>1</sub>, talin and tensin in mature FAs are required to fully validate our hypothesis. Yet, because of their relevance we consider it appropriate to include these data as Fig. S8 and discussing their potential implications in pages 22 and 23 of the revised manuscript.

      Reviewer #2 (Recommendations for the authors):

      (1) Perform control experiments to confirm that the nanocluster size/composition is not affected by the antibodies used.

      As explained in the response to the public reviews, antibody labelling has always been performed after cell fixation, precluding potential cross-linking artefacts due to protein mobility and avoiding unwanted receptor activation. In addition, we have performed super-resolution imaging using DNA-PAINT as a different imaging strategy. In this case, the DNA docking site is site-specifically coupled to one camelid single-domain Ab (sdAB), having a much smaller size as compared to a secondary Ab, reducing therefore linkage error and increasing the accessibility of primary Ab-labelled proteins. As can be observed in new Fig S4, no differences in terms of nanocluster sizes and/or compositions were observed for any of the proteins investigated using DNA-PAINT as compared to our initial STORM data. These control experiments thus rule out any potential artefacts introduced by the secondary Ab. Finally, we would like to highlight that our results on the nanoclustering of integrins in terms of their size and number of localizations is consistent with previous results obtained by other groups around the world using similar labelling protocols as us (Spies et al, J. Cell Biol 2022), or relying on halo-tag strategies, as suggested by the reviewer (see Fujiwara et al, J. Cell Biol. 2023). The consistency of these results amongst different groups gives us further confidence that the nanocluster size/composition are not affected by the antibodies used.

      (2) Extend the study by inspecting a set of integrin adapter proteins for their association with inactive integrins, focusing on adapters associated with the maintenance of the inactive state. Possible candidates would be tensin and filamin, for example.

      We thank the reviewer for the suggestion and have now performed dual-colour super-resolution STED microscopy of α<sub>5</sub>β<sub>1</sub> and tensin-3 on HFF cells seeded for 24 hours. Interestingly, instead of being an integrin inactivator, we found that tensin-3 is also highly enriched at the FA periphery where a large fraction of active β<sub>1</sub> integrins are located, suggesting that at these particular regions, active β<sub>1</sub> could be either engaged to talin (as shown in our original data) or to tensin-3 (our new data shown in Fig. S8). These results might be surprising at first, since tensin competes with talin for the same binding site to the cytoplasmic β-tail of integrins, and thus believed to act as integrin inactivator, as the reviewer indicates. Nevertheless, recent data has shown that tensin is capable to activate integrins (in particular if β<sub>1</sub> is phosphorylated) by interacting with the actin cytoskeleton, providing mechanical coupling for integrin activation (Georgiadou & Ivaska, Trends Cell Biol. 2017). We have now included these new data as Fig. S8 in the revised manuscript. Additional experiments, which in our opinion fall outside of the scope of this work, would be necessary to identify other potential integrin inactivator partners, but certainly a topic of future interest to our group.

      (3) While the methods are described in sufficient detail, it is important to ask if the findings are based on sufficient data. Table S5 provides detailed information about the number of samples studied, and it appears that only small numbers of samples were investigated for certain protein pairs. This should be discussed, and perhaps more data should be obtained to strengthen the data.

      We have now performed additional experiments using DNA-PAINT as alternative super-resolution imaging technique (as also requested by reviewer 3) which adds additional data to the whole manuscript.

    1. eLife Assessment

      This important study presents a convincing methodological approach to probe the structural features of the full-length human Hv1 channel as a purified protein. The method is supported by rigorous biochemical assays and spectral FRET analysis, which will interest biophysicists and physiologists studying Hv1 and other ion channels and membrane proteins. Overall, the work introduces an interesting labeling strategy and provides a methodology that is of value in investigating hHV1 in particular and can be extended to other ion channels. The authors also provide preliminary observations regarding conformational changes induced by zinc.

    2. Reviewer #1 (Public review):

      [Editors' note: this version has been assessed by the Reviewing Editor without further input from the original reviewers. The authors have addressed the comments raised in the previous round of review, shown below.]

      In this study, the noncanonical amino acid acridon-2-ylalanine (Acd) was inserted at various positions within the human Hv1 protein using a genetic code expansion approach. The purified mutants with incorporated fluorophore were shown to be functional using a proton flux assay in proteoliposomes. FRET between native tryptophan and tyrosine residues and Acd were quantified using spectral FRET analysis. Predicted FRET efficiencies calculated from an AlphaFold model of the Hv1 dimer were compared to the corresponding experimental values. Spectral FRET analysis was also used to test whether structural rearrangements caused by Zn2+, a well-known Hv1 inhibitor, could be detected. The experimental data provide a good validation of the approach, but further expansion of the analysis will be necessary to differentiate between intra- and intersubunit structural features.

      Interestingly, the observed rearrangements induced by Zn2+ were not limited to the protein region proximal to the extracellular binding site but extended to the intracellular side of the channel. This finding agrees with previous studies showing that some extracellular Hv1 inhibitors, such as Zn2+ or AGAP/W38F, can cause long-range structural changes propagating to the intracellular vestibule of the channel (De La Rosa et al. J. Gen. Physiol. 2018, and Tang et al. Brit J. Pharm 2020). The authors should consider adding these references.

      Since one of the main goals of this work was to validate Acd incorporation and the spectral FRET analysis approach to detect conformational changes in hHv1 in preparation for future studies, the authors should consider removing one subunit from their dimer model, recalculating FRET efficiencies for the monomer, and comparing the predicted values to the experimental FRET data. This comparison could support the idea that the reported FRET measurements can inform not only on intrasubunit structural features but also on subunit organization.

    3. Reviewer #2 (Public review):

      This manuscript by Carmona, Zagotta, and Gordon is generally well-written. It presents a crude and incomplete structural analysis of the voltage-gated proton channel based on measured FRET distances. The primary experimental approach is Förster Resonance Energy Transfer (FRET), using a fluorescent probe attached to a noncanonical amino acid. This strategy is advantageous because the noncanonical amino acid likely occupies less space than conventional labels, allowing more effective incorporation into the channel structure.

      Fourteen individual positions within the channel were mutated for site-specific labeling, twelve of which yielded functional protein expression. These twelve labeling sites span discrete regions of the channel, including P1, P2, S0, S1, S2, S3, S4, and the dimer-connecting coiled-coil domain. FRET measurements are achieved using acridon-2-ylalanine (Acd) as the acceptor, with four tryptophan or four tyrosine residues per monomer serving as donors. In addition to estimating distances from FRET efficiency, the authors analyze full FRET spectra and investigate fluorescence lifetimes on the nanosecond timescale.

      Despite these strengths, the manuscript does not provide a clear explanation of how channel structure changes during gating. While a discrepancy between AlphaFold structural predictions and the experimental measurements is noted, it remains unclear whether this mismatch arises from limitations of the model or from the experimental approach. No further structural analysis is presented to resolve this issue or to clarify the conformational states of the protein.

      The manuscript successfully demonstrates that Acd can be incorporated at specific positions without abolishing channel function, and it is noteworthy that the reconstituted proteins function as voltage-activated proton channels in liposomes. The authors also report reversible zinc inhibition of the channel, suggesting that zinc induces structural changes in certain channel regions that can be reversed by EDTA chelation. However, this observation is not explored in sufficient depth to yield meaningful mechanistic insight.

      Overall, while the study introduces an interesting labeling strategy and provides valuable methodological observations, the analysis appears incomplete. Additional structural interpretation and mechanistic insight are needed.

    4. Author response:

      The following is the authors’ response to the original reviews.

      We greatly appreciate both reviewers for their insightful comments. As noted by the reviewers and editors, this work validates a new approach to study conformational dynamics of full-length hH<sub>v</sub>1 using acridon-2-ylalanine (Acd), including its advantages and limitations for future structural studies. Accordingly, we have changed the manuscript to “Tools and Resources”. The incorporation of the reviewers’ suggestions has greatly improved the quality of the updated manuscript.

      Public Reviews:

      Reviewer #1 (Public review):

      We thank Reviewer #1 for the feedback and comments. We addressed the points the reviewer raised below.

      Interestingly, the observed rearrangements induced by Zn2+ were not limited to the protein region proximal to the extracellular binding site but extended to the intracellular side of the channel. This finding agrees with previous studies showing that some extracellular Hv1 inhibitors, such as Zn2+ or AGAP/W38F, can cause long-range structural changes propagating to the intracellular vestibule of the channel (De La Rosa et al. J. Gen. Physiol. 2018, and Tang et al. Brit J. Pharm 2020). The authors should consider adding these references.

      We added the suggested references to the Results section.

      Since one of the main goals of this work was to validate Acd incorporation and the spectral FRET analysis approach to detect conformational changes in hHv1 in preparation for future studies, the authors should consider removing one subunit from their dimer model, recalculating FRET efficiencies for the monomer, and comparing the predicted values to the experimental FRET data. This comparison could support the idea that the reported FRET measurements can inform not only on intrasubunit structural features but also on subunit organization.

      We calculated the predicted intrasubunit FRET efficiency and presented the results in the new Figure S10. Pearson’s coefficient decreased from 0.48 for the dimer to 0.18 for the monomer, suggesting the experimental FRET contains information about subunit organization. This was added to the text.

      Reviewer #2 (Public review):

      We appreciate the detailed revision, comments, and feedback from Reviewer #2. We addressed the reviewer’s major and minor points below, which we believe improve the manuscript.

      (1) Tryptophan and tyrosine exhibit similar quantum yields, but their extinction coefficients differ substantially. Is this difference accounted for in your FRET analysis? Please clarify whether this would result in a stronger weighting of tryptophan compared to tyrosine.

      We accounted for differences in the extinction coefficients of Trp and Tyr in our calculations, which are detailed in the Supplementary Text. The assumptions result in a stronger contribution from Trp than from Tyr.

      (2) Is the fluorescence of acridon-2-ylalanine (Acd) pH-dependent? If so, could local pH variations within the channel environment influence the probe's photophysical properties and affect the measurements?

      The acridone fluorescence, which is the fluorophore in Acd, is not pH-dependent between pH 2 and 9 (Stephen G.S. and Sturgeon R.J. Analytica Chimica Acta. 1977). This was added to the text.

      (3) Several constructs (e.g., K125Tag, Y134Tag, I217Tag, and Q233Tag) display two bands on SDS-PAGE rather than a single band. Could this indicate incomplete translation or premature termination at the introduced tag site? Please clarify.

      Yes, the additional bands in the WB are due to the termination of translation for the mentioned protein constructs. We added a note in the legend of Figure 2 regarding this point.

      (4) In Figure 5F, the comparison between predicted FRET values and experimentally determined ratio values appears largely uninformative. The discussion on page 9 suggests either an inaccurate structural model or insufficient quantification of protein dynamics. If the underlying cause cannot be distinguished, how do the authors propose to improve the structural model of hHv1 or better describe its conformational dynamics?

      We understand the confusion about this point. We are not planning to improve the structural model with FRET between Trp/Tyr and Acd. We modified the text to avoid confusion regarding this point. We plan to use Acd as a transition metal ion FRET (tmFRET) donor to study the conformational dynamics of hHv1 in the future (Discussion).

      (5) Cu<sup>2+</sup>, Ru<sup>2+</sup> and Ni<sup>2+</sup>, are presented as suitable FRET acceptors for Acd. Would Zn<sup>2+</sup> also be expected to function as an acceptor in this context? If so, could structural information be derived from zinc binding independently of Trp/Tyr?

      Transition metal ion FRET (tmFRET) uses a fluorophore as the donor and a transition metal ion chelator as the acceptor. For FRET to occur between these donor-acceptor pairs, the fluorescence spectrum of the donor must overlap the absorption spectrum of the metal ion (Zagotta et al., eLife. 2021; Zagotta et al., Biophys J. 2024; Gordon et al., Biophys J. 2024). Zn<sup>2+</sup> does not absorb visible light, so tmFRET cannot occur for this divalent metal.

      (6) The investigated structure is most likely dimeric. Previous studies report that zinc stabilizes interactions between hHv1 monomers more strongly than in the native dimeric state. Could this provide an explanation for the observed zinc-dependent effects? Additionally, do the detergent micelles used in this study predominantly contain monomers or dimers?

      Our full-length hH<sub>v</sub>1 in Anz3-12 detergent micelles is predominantly a dimer, as demonstrated in the new panel of Figure S5. From our data, we cannot compare the effects of zinc between monomers and dimers.

      (7) hHv1 normally inserts into a phospholipid bilayer, as used in the reconstitution experiments. In contrast, detergent micelles may form monolayers rather than bilayers. Could the authors clarify the nature of the micelles used and discuss whether the protein is expected to adopt the same fold in a monolayer environment as in a bilayer?

      We used Anzergent 3-12 detergent micelles, which stabilize hH<sub>v</sub>1 in solution. We indicated this in the Results and Materials and Methods sections. We are also intrigued by whether protein folding and conformational dynamics differ between detergent micelles and proteoliposomes, but our data do not provide an answer to this question. We found that the proteoliposomes used for measuring the hHv1 function don’t have enough Acd signals to record their spectra, preventing us from performing the same FRET measurements between Trp/Tyr and Acd in liposomes. Still, detergent-solubilized hH<sub>v</sub>1 is functional upon reconstitution, demonstrating that its functional folding is not irreversibly altered in micelles.

      Recommendations for the Authors:

      Reviewer #2 (Recommendations for the authors):

      (1) On page 9, the reference to Figure S11 should be corrected to Figure S10.

      We thank the reviewer for catching this mistake. It was corrected in the updated version.

      (2) On page 9, multiple prior studies describing zinc binding to hHv1 should be acknowledged, for example:

      Musset et al. (2010), J. Physiol., 588, 1435-1449; Jardin et al. (2020), Biophys. J., 118, 1221-1233.

      References were added to the text.

      (3) On page 11, the statement "with Acd incorporated ... we can interrogate its gating mechanism in unprecedented detail" appears overly strong relative to the data presented. Another phrasing might be appropriate.

      The sentence was changed. It now reads: “With Acd incorporated at multiple sites in full-length hH<sub>v</sub>1, it will be possible to interrogate conformational changes across the protein’s different structural domains using Acd as a tmFRET donor to understand its molecular mechanisms.”

    1. eLife Assessment

      This important study compares how different classes of drugs act on the SARS-CoV-2 main protease, a key antiviral target, and shows that many of them work by controlling whether the enzyme assembles into its active dimeric form. The evidence, based on a range of complementary biophysical methods, is convincing and points to the interface between the two protein protomers, including a newly found binding site, as a promising target for broad-spectrum antiviral drugs. This work will be of interest to biochemists and virologists working on treatments for coronaviruses.

    2. Reviewer #1 (Public review):

      Summary:

      Since dimerization is essential for SARS-CoV-2 Mpro enzymatic activity, the authors investigated how different classes of inhibitors, including peptidomimetic inhibitors (PF-07321332, PF-00835231, GC376, boceprevir), non-peptidomimetic inhibitors (carmofur, ebselen, and its analog MR6-31-2), and allosteric inhibitors (AT7519 and pelitinib), influence the Mpro monomer-dimer equilibrium using native mass spectrometry. Further analyses with isotope labeling, HDX-MS, and MD simulations examined subunit exchange and conformational dynamics. Distinct inhibitory mechanisms were identified: peptidomimetic inhibitors stabilized dimerization and suppressed subunit exchange and structural flexibility, whereas ebselen covalently bound to a newly identified site at C300, disrupting dimerization and increasing conformational dynamics. This study provides detailed mechanistic evidence of how Mpro inhibitors modulate dimerization and structural dynamics. The newly identified covalently binding site C300 represents novelty as a druggable allosteric hotspot.

      Strengths:

      This manuscript investigates how different classes of inhibitors modulate SARS-CoV-2 main protease dimerization and structural dynamics, and identifies a newly observed covalent binding site for ebselen.

      Weaknesses:

      None. The requested mutagenesis data have been provided in the revised manuscript, and all of my previous concerns have been satisfactorily addressed.

    3. Reviewer #2 (Public review):

      Summary:

      This manuscript presents a sophisticated investigation into the mechanisms by which different inhibitor classes affect the SARS-CoV-2 main protease (Mpro), a pivotal antiviral drug target. This study reveals that effective inhibition can be achieved by modulating the stabilization of the essential dimeric state. It also indicates the dimer interface could be a druggable allosteric site, which may offer a strategy for developing broad-spectrum anticoronaviral agents.

      Strengths:

      The identification of dimer interface stabilization/destabilization as distinct inhibitory mechanisms and the discovery of C300 as a potential allosteric site for ebselen are important contributions to the field. The experimental approach is modern, multi-faceted, and generally well-executed.

      Comments on revised version:

      The authors have very nicely addressed most of the previous comments raised. But one comment remains to be clarified relating to original point 5 and the authors' response:

      "We agree with the reviewer about the need for quantitative rigor in reporting HDX changes. We have calculated the fractional deuterium uptake difference for each peptide fragment discussed in the text between the inhibitor-bound and unbound states. These values, along with their statistical significance (p-values from a two-tailed t-test), have been provided in the revised manuscript (Legends for Figures 3 and 4). Although the HDX change of residues 296-306 is relatively small (<5%), this region showed a reproducible difference with low experimental variability and statistical significance (p < 0.05). Given its location within the C-terminal dimerization interface and its consistency with native MS, we interpret this change as a subtle local conformational perturbation."

      Two questions remain for the statements in line 376-380. First, while it is stated "residues 296-304 in the C-terminal region of Mpro were more flexible upon ebselen binding", the segment of 296-306 is shown Figure 4c. Second, the HDX change for this segment upon ebselen binding is very subtle in the figure (in contrast to the significant HDX change of the same segment in the protein upon PF-07321332 binding), thus making the strong conclusion that "This suggests that ebselen targeting C300 may induce structural changes in the C-terminal helical segment, weakening key hydrogen bonds at the dimer interface and ultimately inhibiting activity" not convincing. The reviewer would suggest the authors either delete this conclusion or largely tone it down.

    4. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      Since dimerization is essential for SARS-CoV-2 M<sup>pro</sup> enzymatic activity, the authors investigated how different classes of inhibitors, including peptidomimetic inhibitors (PF-07321332, PF-00835231, GC376, boceprevir), non-peptidomimetic inhibitors (carmofur, ebselen, and its analog MR6-31-2), and allosteric inhibitors (AT7519 and pelitinib), influence the M<sup>pro</sup> monomer-dimer equilibrium using native mass spectrometry. Further analyses with isotope labeling, HDX-MS, and MD simulations examined subunit exchange and conformational dynamics. Distinct inhibitory mechanisms were identified: peptidomimetic inhibitors stabilized dimerization and suppressed subunit exchange and structural flexibility, whereas ebselen covalently bound to a newly identified site at C300, disrupting dimerization and increasing conformational dynamics. This study provides detailed mechanistic evidence of how M<sup>pro</sup> inhibitors modulate dimerization and structural dynamics. The newly identified covalently binding site C300 represents novelty as a druggable allosteric hotspot.

      Strengths:

      This manuscript investigates how different classes of inhibitors modulate SARS-CoV-2 main protease dimerization and structural dynamics, and identifies a newly observed covalent binding site for ebselen.

      Weaknesses:

      The major concern is the absence of mutagenesis data to support the proposed inhibitory mechanisms, particularly regarding the role of the inhibitor binding site.

      We thank the reviewer for the recognition and comments. We agree that mutagenesis is critical for validating the proposed role of C300. We therefore generated the C300S and C300F mutants and characterized their oligomeric states and proteolytic activities. C300S was designed to remove the reactive thiol group while minimally affecting M<sup>pro</sup> structure and dimerization. C300F was introduced to mimic the steric perturbation associated with C300 modification and assess its impact on M<sup>pro</sup> dimerization. Native PAGE showed that WT and C300S M<sup>pro</sup> predominantly formed dimers, whereas C300F was mainly monomeric. Consistently, C300S retained approximately 70% of WT activity, whereas C300F retained only approximately 10%. Because C300F itself strongly disrupted dimerization, C300S was used as the principal mutant to evaluate the specific contribution of the C300 thiol to ebselen action. Native MS showed that ebselen could still bind to both monomeric and dimeric C300S M<sup>pro</sup> but did not markedly shift its monomer-dimer equilibrium toward the monomeric state. In parallel, ebselen reduced WT activity to approximately 53% of the untreated control, whereas C300S retained approximately 78% activity at the same 1:3 M<sup>pro</sup>-to-ebselen molar ratio. These results provide experimental support for the contribution of C300 to ebselen-induced dimer destabilization and functional inhibition, while the residual binding and inhibition observed for C300S suggest the involvement of additional C300-independent interactions. The corresponding revisions have been made to Methods (Lines 627–648), and Results (Lines 397–453) of the manuscript, together with the newly added figures (Figures S14–S16).

      Reviewer #2 (Public review):

      Summary:

      This is a mechanistic study that provides new insights into the inhibition of SARS-CoV-2 M<sup>pro</sup>.

      Strengths

      The identification of dimer interface stabilization/destabilization as distinct inhibitory mechanisms and the discovery of C300 as a potential allosteric site for ebselen are important contributions to the field. The experimental approach is modern, multi-faceted, and generally well-executed.

      We thank the reviewer for the positive comments and recognition of our study.

      Weaknesses:

      The primary weaknesses relate to linking the biophysical observations more directly to functional enzymatic outcomes and providing more quantitative rigor in some analyses. While the study is overall strong, addressing its weaknesses and limitations would elevate the impact and translational relevance of the current manuscript.

      We thank the reviewer for these comments, which have helped to iM<sup>pro</sup>ve the quality and impact of our manuscript.

      (1) Correlation with Functional Activity:

      The most significant gap is the lack of direct enzymatic activity assays under the exact conditions used for MS and HDX. While EC50 values are listed from literature, demonstrating how the observed dimer stabilization (by peptidomimetics) or dimer disruption (by ebselen) directly correlates with inhibition of proteolytic activity in the same experimental setup would solidify the functional relevance of the biophysical observations. For instance, does the fraction of monomer measured by native MS quantitatively predict the loss of activity? Also, the single inhibitor concentration used in each MS experiment needs to be specified in the main text and legends. A discussion on whether the inhibitor concentrations required to observe these dimerization effects (in native MS) or structural dynamics (in HDX-MS) align with EC50 values would be helpful for contextualizing the findings.

      We thank the reviewer for these important points. To link the biophysical observations more directly to function, we compared the oligomeric states and proteolytic activities of WT, C300S, and C300F M<sup>pro</sup>. C300F was predominantly monomeric and retained only approximately 10% of WT activity, whereas C300S remained predominantly dimeric and retained approximately 70% activity. We further evaluated ebselen inhibition using a matched 1:3 M<sup>pro</sup>-to-ebselen molar ratio. Ebselen reduced WT activity to approximately 53% of its untreated control but reduced C300S activity only to approximately 78%, demonstrating that removal of the C300 thiol significantly attenuated the functional effect of ebselen. These data support a relationship between C300-dependent dimer destabilization and reduced proteolytic activity. The Methods (Lines 627–648), and Results (Lines 397–453) have been revised accordingly, with Figures S14–S16 newly added, in the revised manuscript. We did not expect a linear relationship between the monomer fraction measured by native MS and enzymatic activity loss, because ebselen can modify multiple cysteine residues, and individual modification events may have distinct effects on M<sup>pro</sup> dimerization and catalytic function. The concentrations and molar ratios used in the native MS, HDX-MS, and activity assays have now been stated in the figure legends. The ebselen concentrations used for native MS and HDX-MS were optimized for biophysical characterization and comparison, and therefore, these concentrations might not be directly related to their IC<sub>50</sub> or EC<sub>50</sub> values. In these experiments, ebselen was applied at a 3-fold molar excess relative to M<sup>pro</sup>, consistent with the enzymatic assay. The observed dimer disruption and conformational changes were consistent with functional inhibition, supporting their mechanistic relevance.

      (2) For the two Cys residues found to be targeted by ebselen, what are their respective modification stoichiometry related to the ebselen concentration? Especially for the covalent binding site C300, which is proposed in this study to represent a novel allosteric inhibition mechanism of ebselen, more direct experimental evidence is needed to support this major hypothesis. Does mutation or modification of C300 affect the M<sup>pro</sup> dimerization/monomer equilibrium and alter the enzymatic activity? If ebselen acts as a covalent inhibitor linked to multiple Cys, why is its activity only in the μM range?

      We thank the reviewer for the insightful comments. Our LC-MS/MS data identified C44 and C300 as ebselen-modified residues, but they do not permit reliable site-resolved occupancy measurements because modified and unmodified peptides can differ in digestion efficiency and MS response. We have therefore clarified that these data provide qualitative site identification rather than absolute modification stoichiometry. To obtain direct functional evidence for C300, we generated C300S and C300F mutants. C300S preserved dimer formation and substantial activity, whereas C300F was mainly monomeric and showed severe activity loss. Importantly, although ebselen-bound C300S species were still detected by native MS, ebselen did not markedly redistribute C300S toward the monomeric state, and its inhibition was reduced from approximately 47% for WT to approximately 22% for C300S. These results indicate that C300 is an important contributor to ebselen-induced dimer disruption, while residual binding and inhibition indicate additional reactive sites. Corresponding revisions have been made to the (Lines 627–648), and Results (Lines 397–453) of the manuscript, together with the newly added figures (Figures S14–S16). The moderate micromolar potency of ebselen is consistent with its heterogeneous, multi-site covalent reactivity: modification occupancy and functional consequence are site-dependent, and not every adduct produces complete inhibition.

      (3) For the allosteric inhibitor pelitinib with low-μM activity, no significant differences in deuterium uptake of M<sup>pro</sup> were observed. In terms of the binding affinity, what is the difference between pelitinib and ebselen? Some explanations could be provided about the different HDX-MS results between the two non-peptidomimetic inhibitors with similar activities.

      We agree with the reviewer that the absence of significant HDX changes for pelitinib requires clarification. Different from ebselen that forms covalent bond with multiple cysteine residues of M<sup>pro</sup>, which could lead to sustained conformational changes that are more readily detected by HDX-MS, pelitinib non-covalently binds M<sup>pro</sup> and might not induce significant perturbations in backbone dynamics that are detectable at the peptide level by HDX-MS. These points have been integrated into the revised manuscript (Lines 333-337).

      (4) Native MS Quantification: 

      The analysis of monomer-dimer ratios from native MS spectra appears qualitative or semi-quantitative. A more rigorous and quantified analysis of the percentage of dimer/monomer species under each condition, with statistical replicates, would strengthen the equilibrium shift claims. For native MS analysis of each inhibitor, the representative spectrum can be shown in the main figure together with quantified dimer/monomer fractions from replicates to show significance by statistical tests.

      We thank the reviewer for the suggestion. We have performed a quantitative analysis of the monomer-dimer equilibrium based on triplicate native MS measurements for each condition. Representative spectra, quantified monomer/dimer ratios, and statistical analyses have been added to Figures 1 and S3. The quantitative results have also been described in the Results section (Lines 158–161, 165-168, 172-174, 177-179, 199-200).

      (5) Changes of HDX rates in certain regions seem very subtle. For example, as it states 'residues 296-304 in the C-terminal region of M<sup>pro</sup> were more flexible upon ebselen binding (Figure 4c)', the difference is barely observable. The percentage of HDX rate changes between two conditions (with p values) can be specified in the text for each fragment discussed, and any change below 5% or 10% is negligible.

      We agree with the reviewer about the need for quantitative rigor in reporting HDX changes. We have calculated the fractional deuterium uptake difference for each peptide fragment discussed in the text between the inhibitor-bound and unbound states. These values, along with their statistical significance (p-values from a two-tailed t-test), have been provided in the revised manuscript (Legends for Figures 3 and 4). Although the HDX change of residues 296–306 is relatively small (<5%), this region showed a reproducible difference with low experimental variability and statistical significance (p < 0.05). Given its location within the C-terminal dimerization interface and its consistency with native MS, we interpret this change as a subtle local conformational perturbation.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Major points:

      (1) The study lacks validation through inhibitor binding site mutagenesis assays, especially peptidomimetic inhibitor PF-07321332 and ebselen, which would strengthen the mechanistic conclusions.

      We appreciate this suggestion. For PF-07321332, the inhibitor forms a covalent interaction with the catalytic residue C145 and inhibits M<sup>pro</sup> activity through a distinct mechanism. Previous studies have shown that mutation of C145, such as C145A, completely abolishes M<sup>pro</sup> catalytic activity (Bhandari, D, et al. Communications Biology 2025, 8, 1061), making it difficult to directly evaluate the contribution of this residue to inhibitor-induced inhibition using enzymatic assays alone. This limitation and the relevant literature have now been discussed in the revised manuscript (Lines 272–277). Therefore, we focused on C300-dependent regulation of ebselen, which represents a distinct inhibitory mechanism involving modulation of M<sup>pro</sup> structural dynamics and dimer stability.

      To validate the role of C300 in ebselen-mediated regulation of M<sup>pro</sup>, we generated C300S and C300F mutants and performed additional biochemical and structural characterization. The enzymatic assay showed that the C300F mutation significantly affected M<sup>pro</sup> activity, and the inhibitory effect of ebselen on C300S M<sup>pro</sup> was markedly reduced compared with WT M<sup>pro</sup>. Furthermore, native MS analysis demonstrated that ebselen could still bind to C300S M<sup>pro</sup> but failed to induce a significant shift in the monomer-dimer equilibrium observed for WT M<sup>pro</sup>. These results indicate that C300 is not the only site involved in ebselen binding but is critical for mediating ebselen-induced structural perturbation and dimer destabilization. The manuscript has been revised accordingly for the (Lines 627–648), and Results (Lines 397–453), with new figures (Figures S14–S16) included, further supporting the functional contribution of C300 in ebselen-mediated M<sup>pro</sup> regulation.

      (2) MR6-31-2 is an ebselen derivative and exhibits a lower EC50 (1.78 μM) compared to ebselen (4.67 μM). It would be helpful to discuss why their activities differ, probably based on the assay conditions or binding behavior.

      We agree with the reviewer that the difference in antiviral activity between MR6-31-2 and ebselen requires further clarification. The lower EC<sub>50</sub> of MR6-31-2 may result from iM<sup>pro</sup>ved cellular properties, including compound stability, permeability, intracellular exposure, and potentially altered interactions with M<sup>pro</sup> and/or iM<sup>pro</sup>ved cellular properties. Although MR6-31-2 shares the ebselen scaffold, the modified chemical structure may affect its binding behavior and biological activity. However, EC<sub>50</sub> values obtained from cellular assays cannot directly reflect the biochemical inhibition potency against purified M<sup>pro</sup>. These points have been integrated into the revised Introduction (Lines 98–101).

      (3) In Figures 2, S1, S2, S4, S6, and S11, adding the drug name under each panel would make the data much clearer for readers.

      The corresponding drug names have been added to panels to iM<sup>pro</sup>ve figure clarity.

      Minor points:

      (1) Line 62-63 refers to the "long linker loop," while Figure 1a labels it as the "long loop linker." Please keep this consistent.

      The terminology has been unified as “long loop linker” throughout the manuscript.

      (2) Table 1 should be cited at line 80, and PDB code 7BAK should be included in Table 1.

      PDB code 7BAK has been included in Table 1, and Table 1 has been cited in the context, as suggested.

      (3) Figure 1a should include the corresponding PDB code in the figure legend.

      The corresponding PDB code has been added to the Figure 1a legend, as suggested.

      (4) It would be helpful to indicate in Figure 1a that the upper structure represents the dimer and the lower structure represents the monomer.

      The upper and lower structures in Figure 1a have been indicated as dimeric and monomeric M<sup>pro</sup>, respectively, as suggested.

      (5) In the Figure S1 legend, it should mention that some inhibitor structures (like ebselen and MR6-31-2) are not fully resolved. Also, the Se atom in ebselen should be shown in Figure S1f (PDB: 7BAK).

      The Figure S1 legend has been revised to indicate that some inhibitor structures, including ebselen and MR6-31-2, are partially unresolved, and the selenium atom of ebselen has also been shown in Figure S1f, as suggested.

      (6) Pelitinib is an allosteric, non-covalently binding inhibitor. However, in Figure S3, the native MS profile shows dimer species (13+ to 15+) compared with unbound M<sup>pro</sup> (14+ to 17+). Please clarify this difference.

      We thank the reviewer for raising this good point. Protein charge-state distributions can be influenced by solution-phase conformation, conformational flexibility, solvent properties, and electrospray droplet charging (Susa AC, et al. J Am Soc Mass Spectrom 2017, 28, 332-340). The observed shift in charge state distribution in native MS might suggest that the addition of pelitinib caused changes in the protein conformation, solvent property and electrospray droplet charging. The relevant literature and discussion have been added in the revised manuscript (Lines 200–204).

      (7) Line 172: "S1are" should be corrected to "S1 are."

      Corrected.

    1. eLife Assessment

      This is a fundamental study on the sensory roles of cerebrospinal-fluid-contacting neurons (CSF-cNs) in mammals, revealing how the apical extension is used as an amplifier of chemical changes in content of the CSF. Specifically, the authors show compelling evidence that PKD2L1 is predominantly a pH-sensing channel in CSF-cNs and link its apical localization to dual phasic and sustained responses underlying CSF chemosensation.

    2. Reviewer #1 (Public review):

      [Editors' note: this version has been assessed by the Reviewing Editor without further input from the original reviewers. The authors have addressed the comments raised in the previous round of review.]

      This study by Vitar et al. probes the molecular identity and functional specialization of pH-sensing channels in cerebrospinal fluid-contacting neurons (CSFcNs). Combining patch-clamp electrophysiology, laser-based local acidification, immunohistochemistry, and confocal imaging, the authors propose that PKD2L1 channels localized to the apical protrusion (ApPr) function as the predominant dual-mode pH sensor in these cells.

      The work establishes a compelling spatial-physiological link between channel localization and chemosensory behavior. The integration of optical and electrical approaches is technically strong, and the separation of phasic and sustained response modes offers a useful conceptual advance for understanding how CSF composition is monitored.

    3. Reviewer #2 (Public review):

      Summary:

      Cerebrospinal fluid contacting neurons (CSF-cNs) are GABAergic cells surrounding the spinal cord central canal (CC). In mammals, their soma lies sub-ependymally, with a dendritic-like apical extension (AP) terminating as a bulb inside the CC.

      How this anatomy-soma and AP in distinct extracellular environments-relates to their multimodal CSF-sensing function remains unclear.

      The authors confirm in the GATA3:GFP mice where these cells are labeled that CSFcNs exhibit prominent spontaneous electrical activity mediated by PKD2L1 (TRPP2) channels, non-selective cation channels with ~200 pS conductance modulated by protons and mechanical forces.

      They investigated PKD2L1 pH sensitivity and its effects on CSFcN excitability. They uncovered that PKD2L1 generates both phasic and tonic currents, bidirectionally modulated by pH with high sensitivity near physiological values.

      Combining electrophysiology (intact and isolated AP recordings) with elegant laser-photolysis, they show functional PKD2L1 channels localize specifically to the apical extension (AP).

      This spatial segregation, coupled with PKD2L1's biophysical properties (high conductance, pH sensitivity) and the AP's unique features (very high input resistance), renders CSFcN excitability highly sensitive to PKD2L1 modulation. Their findings reveal how the AP's properties are optimised for its sensory role.

      Strengths:

      This is a very convincing demonstration using elegant and challenging approaches (uncaging, outside out patch of the AP) together to form a complete understanding on how these sensory cells can detect so finely the changes of pH in the CSF.

    4. Author response:

      The following is the authors’ response to the previous reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      This study by Vitar et al. probes the molecular identity and functional specialization of pH-sensing channels in cerebrospinal fluid-contacting neurons (CSFcNs). Combining patch-clamp electrophysiology, laser-based local acidification, immunohistochemistry, and confocal imaging, the authors propose that PKD2L1 channels localized to the apical protrusion (ApPr) function as the predominant dual-mode pH sensor in these cells.

      The work establishes a compelling spatial-physiological link between channel localization and chemosensory behavior. The integration of optical and electrical approaches is technically strong, and the separation of phasic and sustained response modes offers a useful conceptual advance for understanding how CSF composition is monitored.

      Comments on revised version:

      I thank the authors for their extensive revisions and detailed responses to the reviewers' comments. The manuscript has been substantially improved, and most of the major concerns raised in the initial review have been adequately addressed. In particular, the additional analyses of PKD2L1 channel activity, the incorporation of physiologically relevant pH conditions, the clarification of ASIC involvement, and the expanded Discussion have significantly strengthened the study.

      Major scientific concerns largely addressed:

      Quantification of PKD2L1 channel activity

      The authors appropriately addressed my previous concerns regarding the use of Po as the sole measure of channel activity. The inclusion of additional parameters such as apparent Po, open time, nmax, holding current, and membrane charge provides a more robust assessment of PKD2L1 activity and substantially strengthens the conclusions.

      Physiological relevance of pH modulation

      The inclusion of experiments at pH 6.5 and the additional analyses of holding current and resting membrane potential are valuable additions. These experiments considerably improve the physiological relevance of the study.

      ASIC contribution

      The additional pharmacological experiments using ASIC blockers are helpful and support the conclusion that the photolysis-evoked response in the apical process is predominantly mediated by PKD2L1 channels.

      Functional implications

      The expanded Discussion regarding Ca2+-dependent signaling, neurosecretion, and the potential physiological roles of CSFcNs considerably improves the manuscript.

      Remaining concerns:

      Continued overstatement regarding "exclusive" localization and function:

      Although the authors softened some statements in the revised manuscript, the term "exclusive" remains in several key locations, including the title.

      For example:

      "PKD2L1 channels segregated to the apical compartment are the exclusive dual-mode pH sensor..."

      The data clearly demonstrate strong enrichment of functional PKD2L1 channels in the apical process. However, the available evidence does not fully justify the term "exclusive," particularly because:

      - PKD2L1 immunoreactivity is still detectable outside the apical process.

      - ASIC-mediated responses are present in CSFcNs.

      - The authors themselves use more appropriate terminology such as "predominantly located" in the Discussion.

      Therefore, I recommend replacing "exclusive" with more conservative terminology such as:

      - predominant

      - predominantly localized

      - enriche

      - functionally segregated

      throughout the manuscript, including the title, Abstract, Introduction, Results, and Discussion.

      We agree with the reviewer that the world “exclusive” is misleading and should be replaced. Following the reviewer’s suggestions, we have deleted the word “exclusive from the title, which now reads: “PKD2L1 channels segregated to the apical compartment are the functional dual-mode pH sensors in cerebrospinal fluid-contacting neurons.”

      In addition, the word “exclusive” has been changed with more conservative terminology in other parts of the text: lines 80, 420, 466 and 551.

      Use of the term "tonic current"

      The manuscript continues to use the term "PKD2L1 tonic current."

      While the dibucaine-sensitive holding current is clearly present, the precise mechanism generating this current remains uncertain. Indeed, the authors themselves acknowledge in the Discussion that:

      - an alternative conducting state may exist, or

      - unresolved brief channel openings may account for the current.

      Therefore, the data support the existence of a sustained PKD2L1-associated current, but do not yet definitively establish a distinct tonic gating mode of the channel.

      I therefore recommend replacing:

      "tonic current" with a more neutral expression such as:

      - sustained current

      - PKD2L1-associated holding current

      - sustained PKD2L1-mediated current throughout the manuscript.

      Continued use of "off-current" and "off-response":

      The revised manuscript has improved considerably in this regard. However, the terms "off-current" and "off-response" still remain in portions of the text and figure legends.

      Because the manuscript itself demonstrates that the response reflects recovery from transient acidification rather than a separate OFF signaling mechanism, these terms remain potentially misleading.

      I recommend replacing them with terminology such as:

      - photolysis-evoked PKD2L1 current

      - recovery current

      - proton-removal-induced current

      throughout the manuscript, including figure legends.

      We apologize, as the word “tonic” and the terminology “off-current” should have completely disappeared after the first round of revisions. We have now replaced those all along the text. “Tonic” has been replaced by “sustained”.

      “Off-current” or “off-response” have been replaced by appropriate terms in lines: 339, 342, 544, 545, 546, 549, 552, 555, 556, 560, 576, 580, 585, 588, 807 and 929. We have nevertheless conserved the term “off-current” in line 552 as we are referring to terminology used by other authors.

      Minor editorial corrections

      Figure 1Bd Please change: "po" to "Po" for consistency with standard channel physiology nomenclature.

      Figure 1Ca Please add units (mV) to the voltage labels shown on the left side of the traces.

      Figure 3E Please change: "Norm po" to "Norm Po".

      Figure 4Fb Please replace: "sec" with "s" to conform with SI unit conventions.

      Done.

      The authors have addressed the majority of my previous concerns and the manuscript has been substantially improved. The remaining issues are primarily related to terminology and overinterpretation rather than experimental deficiencies.

      Reviewer #2 (Public review):

      Summary:

      Cerebrospinal fluid contacting neurons (CSF-cNs) are GABAergic cells surrounding the spinal cord central canal (CC). In mammals, their soma lies sub-ependymally, with a dendritic-like apical extension (AP) terminating as a bulb inside the CC.

      How this anatomy-soma and AP in distinct extracellular environments-relates to their multimodal CSF-sensing function remains unclear.

      The authors confirm in the GATA3:GFP mice where these cells are labeled that CSFcNs exhibit prominent spontaneous electrical activity mediated by PKD2L1 (TRPP2) channels, non-selective cation channels with ~200 pS conductance modulated by protons and mechanical forces.

      They investigated PKD2L1 pH sensitivity and its effects on CSFcN excitability. They uncovered that PKD2L1 generates both phasic and tonic currents, bidirectionally modulated by pH with high sensitivity near physiological values.

      Combining electrophysiology (intact and isolated AP recordings) with elegant laser-photolysis, they show functional PKD2L1 channels localize specifically to the apical extension (AP).

      This spatial segregation, coupled with PKD2L1's biophysical properties (high conductance, pH sensitivity) and the AP's unique features (very high input resistance), renders CSFcN excitability highly sensitive to PKD2L1 modulation. Their findings reveal how the AP's properties are optimised for its sensory role.

      Strengths:

      This is a very convincing demonstration using elegant and challenging approaches (uncaging, outside out patch of the AP) together to form a complete understanding on how these sensory cells can detect so finely the changes of pH in the CSF.

      Weaknesses:

      Not weaknesses, there are only minor requests to complete the beautiful study.

      (1) The apical extension's response to removal of acidification is nicely illustrated in Figure 4C,G. There's something puzzling there: while the response to Glutamate is immediate, the channel responses to H+ is extremely delayed by 100ms - 2s, and even sometimes came in bursts separated by few hundreds of ms. H+ diffuse even faster than glutamate. Why is that?

      I don't quite understand how the response is so delayed & how to explain the recurring bursts of channel opening in the figure panel ?

      The kinetic of the response to proton uncaging is analyzed in Figure 4E, where the charge of the current traces is plotted against time. What this analysis shows is that the response lasts a few hundred ms (τ 250 ms) and then the PKD2L1 activity increase subsides to baseline. The peak of the response is at 100 ms (Figure 4G), but the increase in activity happens as soon as the uncaging pulse ends (Figure 4D, G and H). This behavior has already been shown in expression systems, where the channel activity is blocked by protons and the blockage is released when the acid is withdrawn. In an intact cell as the CSFcNs studied here, the exact kinetics of the recovery response are probably more complex (and variable) than in expression systems. Indeed, it is known that the recovery of this current depends, for example, on pH and extracellular calcium. Also, PKD2L1 are inhibited by intracellular calcium (de Caen et al, eLife 2016) but are themselves permeable to Ca<sup>++</sup> ions. The interaction of these effects could give rise to the “bursts” that are observed in some cases. However, this is merely speculative at this point.

      - The authors should show in Fig 4C,G the traces for 1-2 s before uncaging occurs so we can appreciate whether such events occur as well in baseline and discuss this further in revisions.

      Following the reviewer’s suggestion, we have added a trace in Figure 4C (upper blue trace) showing the spontaneous activity of the cell, prior to uncaging, as it is already shown for another example in Figure 4D.

      - Could the authors use a fluorescent pH sensor to monitor pH in the extracellular space and in the cell ?

      This is an important point that was already addressed by the reviewing editors in the previous round of revisions. Indeed, we have attempted to perform pH calibrations in the setup using the pHsensitive dye pyranine (or HPTS: 8-Hydroxypyrene-1,3,6-trisulfonic acid). HPTS is a very useful tool for pH calibrations in the physiological range: its pKa value is close to 7.2 and it can be used as a ratiometric dye (its fluorescence is pH-independent at 405–410 nm and pH-dependent at 450 nm). Unfortunately, the calibration under the conditions of a real experiment is not possible because the photolysis in the slice occurs in a tiny volume (approximately 1 µm³ in a total bath volume of more than 1 ml). In these conditions, the 405 nm uncaging pulse bleaches the dye in the photolysis spot and any useful information is lost. In addition, our imaging system is not fast enough to follow the pH change. As discussed in the Materials and Methods section, subsection “Estimation of the pH drop induced by photolysis” (line 791), the fast protonation of bicarbonate indicates that the pH change induced by the photolysis recovers in the submillisecond range.

      - Could the authors investigate whether in the apical extension, PKD2L1 channels are mainly at the outer membrane in the apical extension OR whether many channels are located in inner membranes ?

      PKD2L1 channels are probably subject to a high rate of turnover, and they are certainly localized in the plasma membrane of the apical process as well as in the inner membranes. Although this is a very interesting point, we believe it is out of the scope of this work.

      (2) Suppl Fig 4 is very cool and should be moved to main figure. The coupling of Soma and AP is very tight, yet there is a clear difference in targeting of channels that respond to cues in the CSF. In the context of an intact spinal cord, we can wonder how and when the contribution from ASIC in the some would be relevant to physiology. Can the authors think of experiments with an intact central canal to test the sensitivity and condition of recruitment of pH sensing in the soma (ASIC) versus the apical extension (PKD2L1)?

      We have followed the suggestion of the reviewer and have made Supplementary Figure 4 a main figure.

      The fact that the normal interphase between the spinal cord parenchyma and the cc is lost is already acknowledged in the discussion, lines 486 to 489. As the reviewer suggests, PKD2L1 and ASIC channels seem both to be important in the response of CSFcN to pH changes. However, both channels are activated in very different physiological contexts, as is discussed in the section “The involvement of ASIC channels”. Keeping the central canal intact in order to be as close as possible to physiological conditions, as suggested by the reviewer, would be ideal. However, as CSFcNs are in the middle of the cord, it would require the use of optical techniques that allow to penetrate deep into the tissue (e.g., 2-photon microscopy) that unfortunately are not available in our labs.

      (3) The Reissner fiber is missing after slicing the spinal cord. From our observations in fish, the fiber being under tension triggers lots of activity in CSF-cNs (Bellegarda et al Elife 2023) that also relies on PKD2L1 (Bohm et al NC 2016; Sternberg et al NC 2019). Could the authors discuss the contribution of the Reissner fiber to the PKD2L1 mediated modulation of CSFcN excitability ? Could the authors conceive a way to slice along the anteroposterior axis (sagitally) the spinal cord to keep the Reissner fiber in the central canal when recording CSF-cN apical extension ?

      - The authors should show in Fig 4C,G the traces for 1-2 s before uncaging occurs so we can appreciate whether such events occur as well in baseline and discuss this further in revisions.

      As discussed in the previous point, the in vitro slice preparation has technical limitations that are mainly related to the alterations of the normal structure of the tissue. Although keeping the Reissner fiber intact in a sagittal slice seems possible, accessing the CSFcNs with electrophysiological methods would still be a challenge.

      We have now added a sentence in the Discussion, lines 561 to 564, where we discuss that CSFcN excitability is modulated by the Reissner fiber and that it remains to be explored whether in rodents the gating of PKD2L1 channels is modulated by the Reissner fibre, as has been shown in zebrafish.

    1. eLife Assessment

      This important study reports that neural activity in the auditory cortex (field L) of singing male zebra finches can be modulated by the presence of a female conspecific. These findings extend recent work showing that the activity of dopaminergic neurons in songbirds is also affected by an audience. Solid evidence is presented for the importance of singing context in modulating auditory processing during vocal production, but the study does not yet fully establish that these effects arise specifically from audience-dependent modulation of auditory feedback, as opposed to possible acoustic, temporal, or recording-related confounds. This work should be of interest to researchers studying the context dependence of sensory processing, the role of auditory feedback, and vocal communication during courtship behavior.

    2. Reviewer #2 (Public review):

      This study asks whether auditory responses in the songbird auditory pallium/field L during singing are modulated by social context. Specifically, the authors examine neural responses to delayed auditory feedback during male zebra finch song produced either alone or in the presence of a female. This is an interesting and important question because the evaluation of self-generated vocal output may differ when the song has a dedicated social function.

      The main strength of the work is that it addresses auditory feedback processing during natural vocal behavior and does so across two naturalistic contexts. The revised manuscript is strengthened by additional analyses of spike waveform similarity, response significance, response latency stability, and exclusion of motifs overlapping with female calls. These additions make the reported context-dependent response differences more credible and help address some concerns about recording stability and contamination by female vocalizations.

      The results show that some auditory pallium neurons respond differently to feedback perturbations during directed and undirected song. This finding is potentially significant because it suggests that auditory processing during vocal production is not rigid but could subserve a social function that depends on the listener. If robust, this would add an important dimension to models of song monitoring and sensorimotor control.

      However, the strength of evidence remains moderate rather than conclusive. Several alternative explanations are not fully ruled out. Directed and undirected songs may differ acoustically in ways that could influence neural responses, and it is not yet clear that relevant song features were directly compared or controlled across contexts. The experimental sequence also appears to be ordered, with undirected song recorded before directed song, which makes it difficult to fully separate social-context effects from time-dependent changes in recording quality or neural responsiveness. The added waveform analysis is useful, but does not completely establish continuous unit stability across long recording sessions. In addition, possible song changes around the delayed-feedback target point, including compensatory modifications before or after feedback, remain an important potential confound. Finally, clarification of the time-warping and spike-alignment procedures is important because condition-specific alignment could affect comparisons between directed and undirected song.

      Overall, the data support context-dependent differences in neural responses in some neurons, but do not yet fully establish that these differences arise specifically from audience-dependent modulation of auditory feedback processing rather than from acoustic, temporal, or recording-related confounds. The work is likely to be useful to researchers interested in vocal communication, auditory feedback, and social modulation of sensorimotor processing, particularly as a foundation for future experiments using counterbalanced designs and more direct controls of song structure across contexts.

    3. Reviewer #3 (Public review):

      In this study, Jones et al. examine how neural activity in auditory regions (the auditory pallium) of singing male songbirds is modulated by the presence or absence of an audience (a female conspecific). They test whether activity in auditory pallium differs between conditions in which the male is singing to a female (directed song) or alone (undirected song) and whether response to distortions of auditory feedback (DAF) differ between these conditions. Previous work has shown that in other parts of the songbird brain, sensory-motor activity can differ between directed and undirected song, and that responses to DAF are attenuated when males sing directed song versus undirected song. These prior results raise the interesting question of the extent to which such modulations of activity by the presence of an audience are already present in primarily auditory areas within the pallium. This possibility is also motivated by prior work that has shown that activity in the auditory pallium is not exclusively explained by auditory input, but can also be modulated by the bird's state - whether it is singing or not.

      Against this background, the questions asked here are of interest for two inter-related reasons:

      (1) The authors address whether the presence of an audience (a female conspecific) alters activity in an auditory region during singing. Primary songbird auditory areas such as Field L, and analogous mammalian thalamo-recipient cortical regions such as A1, are often thought of as responding very specifically to the features of sensory stimuli, but are also understood to be modulated by a variety of factors including the attentional and behavioral state of the animal. For audition, such modulation includes whether or not animals are vocalizing and listening to themselves or listening to playback of their own vocalizations. Cited works from Keller (2009) as well as Eliades and Wang (2008) have indicated that the act of vocalizing can modulate auditory responses to self-generated feedback in primary auditory areas relative to those arising from playback of the same sounds. Here, the question is whether responses to self-generated feedback differ between conditions of singing alone versus singing to a female audience. A demonstration that the presence of an audience matters to responses in auditory pallium would add to a general understanding of how it is that non-auditory factors can modulate activity within regions that are considered primarily sensory.

      (2) The authors address the possible source of an audience-dependent modulation of responses to feedback perturbation in the VTA previously reported by Goldberg and colleagues (2023). In the VTA, responses to perturbations during singing are consistently attenuated when males are singing to females versus when they are singing alone, but the underlying mechanisms of this modulation are unknown. Here, the authors test the possibility that such modulation by an audience is already present at the level of auditory pallium. The previously reported attenuation in VTA is a nice example of how neural processing can differ with varying behavioral priorities. Understanding whether this modulation of responses to DAF arises already in auditory areas would further a mechanistic understanding of an intriguing example of state-dependent modulation of sensory processing and behavior and lend broad insight into related phenomena.

      The authors report 1) that activity in the auditory pallium differs between directed and undirected singing at many individual recording sites, but that these changes are heterogeneous, with both increases and decreases in activity, so that there is no consistent change across the population and 2) that modulation of activity by DAF can differ between directed and undirected song, but that there is no consistent attenuation of response (as observed in the VTA) and instead heterogeneous increases and decreases in response to DAF so that there is no net change at the population level.

      These findings are important and of general interest; while they do not readily explain the source of the audience-dependent attenuation of auditory responses to DAF in the VTA, the demonstration of audience-dependent modulation of self-generated feedback and its disruption in the auditory pallium provides an opportunity for further investigation of how changes in social context influence brain and behavior.

      Additional comments and suggestions:

      The authors have done a good job of addressing many of the issues that were raised in the initial round of reviews. There is additional analysis that strengthens the study, including 1) applying a stability criterion to assess the quality of unit isolation, 2) shifting away from a categorical identification of units as "retuning" or not, to an analysis that presents a continuum of changes to neural firing between conditions, 3) use of non-parametric statistics for the assessment of significance of differences in response measures between conditions and 4) exclusion from analysis data from motifs during which females were observed to be vocalizing.

      The authors also have added to the text several important clarifications, and modified language in several ways that improve the presentation and interpretation of results. This includes 1) noting that differences between neural activity during singing with and without DAF does not necessarily reflect "error detection" but could instead reflect how neurons with fixed auditory receptive fields might respond differently to the distinct auditory inputs present between these conditions, 2) clarifying that the recordings were not specifically restricted to Field L, but were distributed more broadly across the auditory pallium, and 3) discussing some of the mechanisms whereby tuning might change due to various sensory, motor and internal factors associated with differences between singing alone and singing to a female.

      I only have a couple of areas of remaining concern that I think could be addressed with further analysis, or some additional discussion, according to the authors' preferences.

      (1) Stationarity of neural response

      My main residual concern relates to the issue raised in the previous round of review of how much of the observed difference in activity between morning sessions when the male is alone and later sessions when the male is singing to a female could reflect changes in neural response properties (non-stationarity) due to the passage of time (sometimes at least several hours) rather than specifically due to the presence or absence of an audience.<br /> The authors restriction of data to recordings that passed a stability criterion for unit waveforms is helpful in addressing whether the same units are 'held' over the course of the experiment. However, even with well isolated units, the response properties or tuning of units can change over time due to a variety of factors that include changes in internal state, circuit excitability, up and down states, neural plasticity, etc.

      The previous review noted several examples of data from the manuscript that illustrated this concern - instances where response properties of neurons appeared to change over time within a given condition. Any such changes in response properties that occur in the absence of a change in audience would tend to contribute to the reported "retuning" of responses.

      One thing that the authors could do to address this issue would be to discuss potential contributions of non-stationarity of responses over time as a potential confounding variable and then editorialize about why they think this seems unlikely to explain many cases in which response properties change between conditions. See comments to authors for one specific suggestions along these lines.

      Alternatively, the authors could carry out additional analyses to evaluate this issue more quantitatively. For example, by measuring the magnitude of "spontaneous" changes in responsiveness observed within conditions (such as by comparing the motif aligned activity for the first n examples within a condition against the activity during the last n examples) and comparing that with the magnitude of changes observed across conditions.

      Another approach would be to carry out some sort of "change point analysis" on the motif-related activity for each experiment in order to establish how often the most abrupt changes in activity occur at the transition between conditions versus spontaneously at other times.

      Lastly, while the experimental design didn't specifically include interleaved blocks of undirected (alone) singing and female directed singing, the methods indicate that the female directed singing data were collected by repeatedly introducing females for 10 minutes at a time. If there are even a couple of cases where the males produced song alone in the periods between female presentation, it would be worth testing whether modulation of neural firing tracked these interleaved conditions.

      Previously published work such as the interleaved recordings of Hessler and Doupe indicate a close and reversible tracking between modulation of neural activity in sensorimotor song system nuclei and switches between singing alone and singing to a female. With respect to the possibility raised in the rebuttal of whether males continue to sing 'female directed song' even after the removal of a female, these and other published data suggest that this is not likely to be the case. But if this were a concern in interpreting any data, the authors could directly assess the male's song for previously described changes in acoustic variability that also track changes in the presence of an audience.

      (2) Further discussion of how an audience might influence responses.

      With respect to the mechanisms whereby an audience might modulate neural responses, a somewhat expanded discussion of possibilities with reference to relevant literature would be helpful. This could include reference to evidence for various neuromodulatory systems participating in modulating singing related activity in song system nuclei based on presence or absence of a female - do these neuromodulatory systems project to the relevant regions of the auditory pallium or its lower-level inputs within the ascending auditory pathway such that they could concurrently act on auditory circuitry?

      In addition to possibility that the presence or absence of an audience affects auditory circuitry via a change in attention, alertness, or motivation, might efference signals associated with singing or locomotion/dancing reach and influence auditory pathways? Given that premotor activity and acoustic structure of song differ between conditions, might any singing-related efference copy activity that reached auditory regions also differ between conditions? A related interesting possibility that could be worth noting is that the presence of a female generally elicits increased locomotion and dancing on the part of the male that accompanies female directed song. Several studies have noted that general forms of locomotion can also result in efference copy signals reaching and influencing auditory regions (e.g. see Schneider and Mooney, Annual Review, 2018; Han et al. "Locomotion-induced neural activity independent of auditory feedback in the mouse inferior colliculus" iScience 2026 - the latter reference is interesting in that it appears to indicate bi-directional modulation of neural activity as observed across units in the current study).

      Minor:

      (1) The authors describe some units as showing "Activation by the absence of DAF." Because the birds in the study have extensive experience with DAF on a subset of trials, it is possible that the increased responses in the absence of DAF reflect a positive deviation from expectation of distortion (as seems to be the case for VTA neurons in previous work). But it is also possible that the broadband DAF stimulus drives inhibition of auditory responses in some cases, and the greater responses in the interleaved trials with normal feedback simply reflect the absence of that inhibition (rather than a positive deviation from a learned expectation). In keeping with the authors shift away from the use of "error detection" elsewhere in the manuscript, it might also be good to use less interpretive language here instead of "activation by absence of DAF".

      (2) At the authors discretion, it would be interesting to know if there is any relationship between the way in which changes in audience affect activity with normal auditory feedback versus with DAF. For example, if normal singing responses are attenuated in the female directed condition, are the responses to DAF also attenuated?

      (3) In figure 2, the vertical dashed lines associated with the rasters indicate the onset and offset of motifs. For several of the figure panels, the spectrograms show motifs that are not aligned with these onsets and offsets. Please clarify or modify (are the rasters from time-warped data but the spectrograms are not -time warped?).

      (4) The authors equate peaks in activity before the onsets of motifs with premotor activity: ["A previous study recording from Field L in zebra finches reported neural activations prior to the onset of singing, consistent with premotor signaling (Keller and Hahnloser, 2009). We tested for context dependent changes in premotor activity by examining peaks in neural activity aligned to motif onsets. Across the population neurons did not exhibit significant changes in the timing of motif onset-aligned activity (Figure S4)."]<br /> However, the spectrograms as shown in Figures 1 and 2 indicate that each motif is often preceded immediately by other song syllables such as introductory notes or syllables from the end of the preceding motif. Further analysis would be required in the current study to demonstrate that the activity present before motif onsets reflects premotor activity rather than auditory responses to the proceeding syllables. Please soften the claim that this reflects premotor activity or provide additional analysis or argument.

    4. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public Review):

      Summary:

      This study examines the context-dependent modulation of auditory cortical neurons in response to expected sensory input, either self-generated sounds or expected perturbations of self-generated sounds. Specifically, using songbirds, the authors ask whether social context (the presence of a female conspecific) affects 1) the response of auditory cortical neurons to the bird's own song when he is singing; and 2) the response of neurons to perturbations of auditory feedback that the bird has been trained to expect.

      Strengths:

      First, the authors report that across the population, the responses of the neurons does not differ when a male bird sings alone or if he sings to a female. A fraction of auditory cortical neurons, however, do show significant differences in the firing rate, precision, and/or degree of burst firing when males sing alone vs. when they sing to females. This finding is broadly consistent with the literature showing that sensory neurons (visual, auditory, somatosensory, etc.) can be rapidly reconfigured into different "information processing modes" depending on behavioral state (e.g., quiescence vs. vigilance).

      For the perturbation experiments, the authors trained birds to expect distorted auditory feedback during a particular syllable. They found that some neurons showed greater responses during perturbation when a female was present (compared to when males were alone) while other neurons had smaller responses during perturbation when a female was present. In addition, the response of a small number of auditory cortical neurons were not affected by behavioral state. These results contrast with their prior report that the responses of midbrain dopaminergic neurons that project to the basal ganglia are "uniformly reduced" in the presence of a female, raising a question of how an evaluation signal is transformed in the circuit from the primary sensory region to the midbrain.

      Weaknesses:

      While the experiments and analysis are solid, the finding that social context can alter responses of auditory cortical neurons in a multitude of ways (increase, decrease or no change) raises several questions that can be examined with additional analysis. For example, do context-dependent differences in auditory responses derive from context-dependent differences in the songs? Are context-dependent differences present in all classes of neurons and throughout the auditory system?

      The observed heterogeneity in the firing properties of auditory cortical neurons, both in response to self-generated sounds and during perturbations of auditory feedback, raises the question of which neurons are sensitive to social context (which likely can be addressed by the authors in a revision). The authors should provide additional details about the recordings:

      (a) What are the locations of the recording sites? Prior work has shown that there is an organized map of spectrotemporal features of sounds in the auditory cortex of songbirds; spectral tuning widths change along the medial-lateral axis and temporal tuning widths differ between the input and output layers of Field L. Were the recordings primarily in Field L2 (thalamo-recipient region), L1 or L3? Were some recordings lateral to Field L in secondary auditory regions? Were the neurons that showed context-dependent changes in firing properties localized or distributed throughout Field L (i.e., were the context-dependent differences in neural responses truly brain-wide)? At a minimum, the authors should include a schematic showing the different regions of Field L and a summary of the location of the recording sites. Images of the processed tissue with electrolytic lesions would also be helpful.

      We agree that the anatomical targeting and limits of localization should be made explicit. In the original manuscript, we referred broadly to recordings from "Field L" and described targeting coordinates in the Methods. In the revised manuscript, we have softened the anatomical claim from "Field L" to "auditory pallium" where appropriate, while explicitly stating that electrodes were aimed at Field L. We also added anatomical caveats and a new supplemental figure.

      The revised title and abstract now reflect this more conservative anatomical framing. For example, the abstract now states: "Here we recorded neural activity from the auditory pallium in zebra finches practicing singing alone and directing courtship songs to females." In the Introduction, we now explicitly state both the intended target and the limitation: "We targeted our recording electrodes to Field L, a primary auditory pallial area that projects into multiple higher auditory areas that, in turn, project to VTA."

      We then added the caveat: "Field L is composed of multiple subdivisions and surrounds the interfacial nucleus and because the implanted wire bundles spread in a small radius of up to ~0.5 mm, our recordings likely included large territories of the auditory pallium (Figure S1)."

      We also added mechanistic/anatomical context for why these recordings may reflect activity shaped by broader auditory forebrain circuitry: "Although Field L is classically described as a primary auditory thalamorecipient region, its activity may also be shaped by contextual signals related to the courtship context, potentially via recurrent interactions with higher-order auditory forebrain regions such as the caudal mesopallium (CM) and caudomedial nidopallium (NCM) (Bauer et al., 2008; Figure S1C)."

      Changes made in revision: We added Figure S1, which includes anatomical subdivisions, an example histological slice showing the area where the cannula was implanted, and auditory pathway connectivity. We also revised the wording throughout the manuscript from "Field L neurons" to more conservative phrasing such as "auditory pallium neurons" or "pallial auditory neurons" when appropriate. We did not claim layer-specific localization, because the revised manuscript explicitly states that we cannot make such claims.

      (b) Was the context-dependent modulation limited to a particular class of neurons (distinguished by spike waveform shape, spontaneous firing rate, or other feature)?

      We agree that identifying whether context-dependent modulation is associated with specific neuronal classes is important. In the revision, we added analyses examining relationships between spike width and various firing characteristics. We also looked for potential relationships between mean rate and DAF response, IMCC and DAF response, and found no clear trend. Overall, we did not observe any clear relationship between DAF response, DAF response modulation, and metrics like spike width or mean firing rate.

      The revised manuscript states: "Action potential width of individual neurons has previously been used to classify putative interneurons or putative principal cells in the zebra finch auditory pallium (Calabrese and Woolley, 2015)."

      We then describe the new analysis: "We tested if DAF-response scores, mean firing rates during singing, burst fraction, IMCC, and the change in all of these between undirected and directed singing was correlated with spike half width (spike half-width measured as peak-to-trough time; Figure S5)."

      The revised result is: "DAF response in either condition, the change in DAF response across conditions, and the change in firing rate, burst fraction, and IMCC were not significantly correlated with spike width (Figure S5A-B)."

      We also report that some general firing properties did correlate with spike width: "Consistent with previous literature, firing rate was significantly correlated with spike width (Pearson's correlation, p=0.003; Figure S5C). Interestingly, burst fraction (p=8.3x10-4) and IMCC (p=5.4x10-5) were also significantly correlated with spike width (Figure S5C)."

      Changes made in revision: We added Figure S5 and associated text analyzing whether context-dependent changes in DAF response, firing rate, burst fraction, and IMCC were correlated with spike half-width. These analyses did not support the conclusion that context-dependent DAF modulation was restricted to a waveform-defined neuronal class.

      (a) Prior work has shown that songs of zebra finches differ slightly when males sing alone compared to when they sing to females: songs are faster; pitch is less variable; and the number of introductory elements is greater when males sing to females. Do some of the observed social context-dependent differences in the responses of auditory neurons reflect differences in the songs in the two conditions? Did the authors of this study also find premotor activity in Field L, and if so, did it differ between the two social contexts? Might differences in Field L responses reflect motor/song differences?

      The revised manuscript now addresses the issues of context-dependent changes in song in several ways. First, we explain why motif-aligned comparisons are meaningful: "The acoustic structure of undirected and directed motifs is highly similar in adult finches, enabling singing-related neural activity to be precisely aligned and compared across contexts."

      Second, we added analysis and discussion of premotor-related activity. Changes made in revision: New results paragraph and new Fig S4. "A previous study recording from Field L in zebra finches reported neural activations prior to the onset of singing, consistent with premotor signaling (Keller and Hahnloser, 2009). We tested for context-dependent changes in premotor activity by examining peaks in neural activity aligned to motif onsets. Across the population neurons did not exhibit significant changes in the timing of motif onset-aligned activity (Figure S4)."

      (b) For the perturbation experiments, this raises a question of whether perturbation amplitude is different when a male is alone and when a female is present. It would be useful to know if (and how much) perturbation amplitude varied depending on the location inside the cage as well as whether the sound pressure level of the underlying song was higher (e.g., Lombard effect).

      We previously calibrated the perturbation amplitude in Roeser et al., 2023, in an identical recording setup. Two speakers deliver the feedback on either side of the bird's home cage. We acknowledge the possibility that the position and orientation of the bird can affect the way the sound hits either of the bird's eardrums and thus potentially affect a neural response. However, neural activations following the absence of distortion playbacks were a major feature of the dataset and were context-dependent in some cases.

      The Methods state: "DAF was implemented with a custom LabVIEW acquisition program that analyzed song syllables in real-time and delivered syllable-targeted feedback." and "DAF (50 ms broadband noise bandpass filtered at 1.5-8 kHz to match frequency range of zebra finch song) was played over speakers in the recording chamber on top of a specific target syllable randomly on 50% of motif renditions."

      The revised manuscript also makes clear that experiments occurred in the bird's home cage: "Experiments were carried out in the male's home cage, which was inside a sound isolation chamber."

      Importantly, the revised Results show that not all DAF-related responses were simple activations to additional sound. Some neurons were activated by the absence of distortion: "Unexpectedly, some neurons were not activated by the song distortion but rather by the lack of target syllable distortion." and "These activations following undistorted renditions could also depend on the courtship context."

      Changes made in revision: We clarified the DAF stimulus and recording setup in Methods and added Figure 3 showing neurons activated following undistorted renditions. These data argue that context-dependent responses are not simply explained by DAF sound amplitude, although we do not claim that position-dependent acoustic variation was fully eliminated.

      Finally, it would be helpful if the authors could include a model and/or more discussion of how the uniform attenuation in midbrain dopaminergic neurons may arise given the heterogeneous responses in Field L.

      The revised manuscript provides evidence for context-dependent retuning upstream of VTA, but does not offer a direct mechanistic explanation for the uniform attenuation seen in dopaminergic neurons. The revised Discussion states: "Because the main goal of this study was to test if courtship-associated reduction in DAF signaling, recently observed in VTA DA neurons (Roeser et al., 2023), resulted from a local process in VTA or reflected a retuning of auditory responsiveness, we explicitly tested for changes in DAF responsiveness between alone and female-directed singing."

      It then explicitly contrasts auditory pallium and VTA: "Surprisingly, we discovered that Field L neurons could retune at the transition from lone to courtship singing in diverse ways, consistent with a more widespread process in the brain that does not fully explain the uniform DAF-signal attenuation observed in VTA."

      Changes made in revision: We expanded the Discussion to explicitly state that auditory pallium retuning is heterogeneous and therefore does not fully explain the uniform attenuation observed in VTA. We do not present a formal circuit model, but we now more clearly frame the result as evidence for broader sensory retuning that is likely transformed downstream.

      Reviewer #2 (Public Review):

      Summary:

      In the manuscript, Jones and Goldberg study auditory cortex in male zebra finches. They explore song-related responses in two different contexts, when the male is either alone or in the presence of a female. They find a heterogeneity of responses, in line with auditory cortical neurons computing the social modulation of responses found in VTA.

      Weaknesses:

      Stability of responses has not been studied: some neurons seem to have responses that slowly drift in time, which could lead to observed differences between alone and with-female conditions. Also, possible motor confounds and sound-of-audience confounds should be addressed. The language is often imprecise.

      Stability and Reversal: It is a bit unfortunate that stability of effects seemingly has not been studied by reversing experimental conditions. The work would be much stronger if authors could show that audience-dependent tuning is robust in individual cells. Did they record from some neurons during reversal back to the alone condition?

      We agree that recording stability is essential. A reversal experiment was not feasible for this dataset, as it is difficult to confirm whether song motifs produced immediately following female presence represent undirected singing or are directed to an unseen but recently present female. Instead, the revised manuscript adds a strict unit-stability criterion based on waveform similarity across conditions.

      The revised Results state: "Importantly, because these neural recordings were performed over long time courses (~2-8 hours), a strict threshold for stability was imposed." The exact criterion is: "A Pearson's correlation coefficient of at least 0.99 between the average neural waveform during undirected and directed singing was required for a unit to be considered stable (Dickey et al., 2009; Figure S2)."

      Changes made in revision: The strict waveform-stability inclusion criterion and new Figure S2 directly showcase unit stability across the time course of the experiments.

      Motor responses: Does DAF playback change song? If so, especially if it applies only in one of the two conditions (audience/no audience), then the observed response differences could be motor-related rather than auditory responses.

      We agree that motor confounds must be minimized. We previously found that DAF did not affect the acoustics of the subsequent syllable (Gadagkar et al., 2016). The revised manuscript clarifies that DAF and undistorted trials were randomly interleaved and analyzed by comparing matched renditions within conditions. Importantly, we only analyzed motif-aligned activity, ensuring that all syllables within the song motif are the same.

      Changes made in revision: We clarified the DAF analysis framework and added a more conservative permutation-based analysis comparing distorted and undistorted trials within each context, then comparing those DAF-response vectors across contexts. We do not claim that all possible motor consequences of DAF are eliminated, but the analysis directly tests neural responses to randomly interleaved distorted versus undistorted renditions.

      Similarly, motif-aligned spiking activity was time warped to the median duration of undirected or directed motifs. Could the shorter motifs during directed song lead to alignment differences that would account for the different error responses in alone/with-female conditions?

      We agree this is an important technical point. The time-warping we conducted, standard in the field, compensates for the tempo differences between directed and undirected song. Importantly, our main analysis of change in error response no longer uses a 100 ms response window, but rather includes all windows in the motif.

      Changes made in revision: We clarified that the revised DAF response analysis uses motif-aligned, time-warped spike trains. Importantly, the revised analysis moves away from relying on a single scalar response window and uses bin-wise permutation tests with family-wise error correction.

      Audience versus sound of audience: Is it truly the audience that causes the difference in error responses or is it the sounds the audience makes?

      We agree that the sensory cues defining "audience" cannot be fully separated in this experiment. The reviewer raises an important point that female zebra finches occasionally call at the male. We have excluded all song motifs from analyses that include an overlapping female call.

      The revised Methods now explicitly state that motifs overlapping with female calls were excluded: "Any motifs that had overlapping time with a female call in directed motifs was excluded from analysis."

      We also revised the Discussion to treat the mechanism by which auditory pallium receives information about the female as an open question: "An open question is how auditory pallium receives information about whether a female is present, and how this information influences neural activity."

      Changes made in revision: We excluded motifs overlapping with female calls and added discussion explicitly acknowledging that how female presence is represented in auditory pallium remains unresolved. We do not claim to distinguish visual, auditory, social, or motivational components of the female-present condition.

      Reviewer #3 (Public Review):

      Summary:

      In this study, Jones et al. examine how neural activity in a primary auditory area (field L) of singing male songbirds is modulated by the presence or absence of an audience (a female conspecific). Prior work has demonstrated that the presence of an audience attenuates the responses of dopaminergic neurons to distortions of auditory feedback (DAF). Here the authors report that even in a region that is primarily considered sensory, responses to DAF are also modulated by the audience, although in a heterogeneous manner. However, to be fully persuasive, additional analyses will be required to address how much of the apparent modulation by audience may be explained by other factors such as changes in recorded neurons or their properties over time.

      (1) A central concern relates to whether the main reported effects associated with differences in singing directed versus undirected song reflect only those changes in conditions, versus contributions from changes in unit isolation or response properties over time.

      We completely agree that unit stability is critically important in this study. To address this concern, we now quantify stability and apply strict inclusion criteria adopted from a study that assessed unit stability over days (Dickey et al., 2009). Additionally, we now include average waveform overlays for all example units across conditions as supplemental Figure S2.

      Changes made in revision: We added: "Importantly, because these neural recordings were performed over long time courses (~2-8 hours), a strict threshold for stability was imposed." and "A Pearson's correlation coefficient of at least 0.99 between the average neural waveform during undirected and directed singing was required for a unit to be considered stable (Dickey et al., 2009; Figure S2)."

      (2) A second concern has to do with the categorical definition of 'error neurons'. The authors define a subset of neurons as error responsive only if their responses to DAF exceed a specific threshold (2.5 standard deviations). The problem is that for some neurons categorically defined as being responsive to DAF in only one condition, there is almost certainly not a significant difference in the actual responses to DAF between conditions.

      We overhauled our analyses characterizing DAF responses. Rather than relying only on a 2.5 z-score threshold, we now use a more conservative permutation-based approach that directly tests DAF responsiveness and context-dependent changes in DAF responsiveness.

      The revised Results state: "Statistical tests defining auditory neurons as DAF-responsive or not in a binary fashion may not be suitable if the underlying population of DAF-related responses exist on a continuum from responsive to non-responsive."

      The updated result is: "This more conservative approach identified 48/147 neurons as DAF-responsive in at least one condition, with 13 of those neurons exhibiting a significant modulation in their DAF response between undirected and female-directed singing."

      (3a) Some discussion of what is already known about the auditory tuning of Field L, and the extent to which responses associated with distortion of feedback may reflect the frequency tuning of Field L neurons versus something that might be construed as more specifically as detecting an error in perceived feedback.

      We agree that DAF-related changes in firing do not necessarily imply that neurons are explicitly detecting an "error" between predicted and actual feedback. Field L neurons can have spectrotemporal receptive fields and frequency tuning such that a broadband DAF stimulus could drive excitation or inhibition simply because the stimulus overlaps with excitatory or inhibitory regions of a neuron's receptive field. We therefore revised the manuscript to use more cautious language and to describe these responses as DAF-related or feedback-related signals rather than categorically as "error responses".

      Changes made in revision: The title was changed from "Auditory cortical error signals retune during songbird courtship" to "Auditory cortical feedback signals are modulated during songbird courtship". We also added a sentence to the Discussion: "However, it is important to note that DAF-related changes in firing in auditory neurons do not necessarily imply that neurons compute sensory prediction errors. DAF-related responses could arise from ordinary auditory tuning to the broadband distortion stimulus."

      (3b) It would also be useful to discuss further previous work on differences in auditory tuning or responses between conditions when subjects are vocalizing, versus when vocalizations are played back, and to what extent efference copy signals might contribute to the processing of feedback distortions.

      We agree these are important points. Our experimental design did not include sufficient passive bird-own-song (BOS) playback trials to permit quantitative comparisons with vocalizing conditions, and we therefore cannot draw firm conclusions about the contribution of efference copy signals to the DAF responses described here. We did observe robust motif onset-associated neural activations, including some activity preceding motif onset, which were present across both social contexts (see new Figure S4). These observations are consistent with prior reports of premotor-related signals in Field L (Keller and Hahnloser, 2009), but whether such signals contribute differentially to DAF processing across contexts remains an open question that we now acknowledge in the Discussion.

      (3c) To what extent did the current study control for any vocalizations or other sounds produced by females during the directed singing, and could this have contributed to differences in Field L activity between conditions?

      Please see response R2.4 above, in which we describe the exclusion of all song motifs that overlapped in time with a female call. This exclusion criterion was applied throughout all analyses of directed singing.

      Figure 1D: In the directed condition there are no spikes at all following the first handful of motif renditions. Were the directed and undirected recordings interleaved here?

      Undirected and directed trials were not interleaved. The raster plots are presented in chronological order; however, for each behavioral condition, rows are sorted with the earliest renditions at the bottom and the most recent at the top. We have clarified this in the figure legend.

      A minor issue: the raw example trace with male alone does not seem to have a corresponding set of points in the raster plot. For panel E, I also cannot find rasters that correspond to the example recordings shown at top.

      In the original version, we randomly downsampled the condition with more trials to equalize trial counts across conditions in the example rasters, while performing all analyses on the full set of recorded trials. As a result, the example spike shown in the raw trace was drawn from one of the downsampled trials not displayed in the raster.

      Changes made in revision: For greater transparency, we now include all trials from both conditions for each example neuron in Figure 1.

      Figure 2A also shows a neuron that looks like it has non-stationarity; for the alone condition without altered feedback, the main peak has no spikes for the bottom half of the rasters.

      In the original version, example neurons were selected to illustrate the DAF-response scoring method, which in some cases highlighted neurons with less stable response profiles. In the revised manuscript, we have replaced this example with neurons that exhibit more robust and stable DAF-related responses, and we now provide a broader set of example neurons illustrating both increases and decreases in DAF responsiveness across conditions.

      Other figures show firing rate distributions that appear to be very non-Gaussian, with some motifs during which there is a lot of activity, and others in which there is little activity. Please consider applying non-parametric tests as appropriate.

      We agree. In general, some neurons exhibited non-uniform firing rate distributions across trials. All of our main analyses are now conducted using non-parametric permutation tests, which do not assume a Gaussian distribution of trial-by-trial firing rates.

      Approaches to addressing the non-stationarity issue could include more specifically indicating examples in which recordings from the alone condition and directed condition are interleaved and exhibit reversible changes in the pattern of responses.

      Unfortunately, nearly all of our undirected and directed recording periods were not interleaved, as the experimental design required a block of undirected singing followed by directed singing with female presence. We find it informative, however, that DAF-response modulation was observed in both directions, with some neurons losing DAF responsiveness during directed song and others gaining it, a pattern that is difficult to attribute to a simple unidirectional drift in recording quality. We now provide additional examples illustrating both directions of modulation in Figures 2 and 3.

      The methods and/or raster plots should include some further explanation of the time periods over which recordings were made in the alone versus directed conditions, and the extent to which they are interleaved or not.

      We have clarified this in the revised Methods. In brief, recording began when the home cage lights came on each day, with the male left to sing alone until at least 40 undirected song motifs were collected. A female was then introduced in approximately 10-minute intervals until at least 40 directed song motifs were collected. The total recording duration on a given day ranged from 0.56 to 10.27 hours, reflecting variability across birds in the time required to elicit sufficient singing in each context. We have added this information to both the Methods and relevant figure legends.

      It would be most helpful to assess the stability of waveforms and unit isolation across time.

      We now apply strict inclusion criteria based on waveform stability, as described in R3.1 above. SNR was quantified as Vpp/(2*sigma_noise), where Vpp was the peak-to-peak amplitude of each filtered spike waveform and sigma_noise was estimated from the median absolute deviation of the filtered voltage trace. This combines the peak-to-peak normalization used by Nordhausen et al. (1996) with the robust noise estimator described by Rey et al. (2015). Waveform overlays for all included example units are provided in Figure S2.

      It would be reassuring to see that significant differences between conditions are equally or more prevalent under the conditions of greatest unit isolation and recording stability.

      The average SNR of neurons ultimately included in the analysis was 9.47 +/- 3.57, with a minimum of 4.69. Neurons that exhibited significant DAF-response modulation did not have a significantly different SNR than neurons that did not exhibit significant modulation (Wilcoxon rank-sum test, p=0.38). The mean SNR for significantly modulated neurons was 8.70, compared to 9.5 for non-modulated neurons, indicating that the detection of context-dependent modulation was not systematically biased toward neurons with lower recording quality.

      One other way that the authors might be able to address the main concern would be to look at the stability of firing patterns within conditions.

      We agree that stability of firing patterns within conditions is an important consideration, and this concern directly motivated the adoption of the permutation-based analysis described above. In this framework, the observed DAF-response difference between conditions is compared to a null distribution generated by shuffling condition labels across trials. This approach inherently accounts for within-condition trial-by-trial variability and does not assume stationarity of firing rates.

      It would be helpful to have additional explanations of the criteria used for counting spikes, and assessing stability of recordings.

      Spike waveforms were visually inspected for consistency using our custom MATLAB GUI on a 12-second file basis. Interspike interval violations below 1 ms were explicitly checked as an indicator of multi-unit contamination. Detection thresholds were manually set, and each recording file included in the analysis was independently inspected. We have added a more explicit description of these procedures to the Methods section.

      For the specific examples shown in figures, it would be useful to indicate by small tick marks or otherwise which spikes were counted as single units.

      We appreciate this suggestion. In the revised figures, we have improved the clarity of the example raw voltage traces by annotating the detection threshold and, where multiple units were present on a channel, indicating the waveform amplitude range corresponding to the isolated single unit. We believe this provides sufficient transparency regarding spike identity without requiring tick marks on every individual spike, which would substantially reduce legibility of the example traces.

      What were the criteria for determining multi-unit versus single-unit activity?

      In the context of this manuscript, "multi-unit activity" refers to channels on which no single neuron could be reliably distinguished from others based on waveform shape and amplitude. Units ultimately included in the study were those for which a single, consistent waveform cluster could be identified and isolated in the custom GUI. In cases where a second distinguishable unit was present on the same channel, it was manually excluded from the sorted single-unit record. We have clarified this distinction in the Methods.

      Categorical scores: This definition results in cases where responses of 2.45 vs 2.55 are described as 'retuned', even if these responses are not significantly different. Retuning would be more persuasively demonstrated if the authors could provide a test of whether or not the responses for individual neurons differ significantly between conditions.

      We completely agree, and thank the reviewer for motivating us to develop a more rigorous statistical approach. Our revised analysis uses a non-parametric permutation test that explicitly tests for significantly different DAF responses between undirected and directed singing conditions, with correction for multiple comparisons. This replaces the previous threshold-based categorical classification and directly addresses the concern that neurons near the threshold boundary were being treated as categorically different.

      Recommendations for the authors:

      Reviewer #1 (Recommendations For The Authors):

      Minor comments:

      (1) Please include a schematic of the brain, including the different subregions of Field L and the connections between auditory regions and the midbrain.

      Done. Figure S1 has been added, including a schematic of Field L subdivisions and auditory pathway connectivity.

      (2) The authors should include some additional information about the recordings, such as the proportion of Field L neurons that exhibited singing-related changes in firing rate. It would be helpful to include some examples of spontaneous activity when the bird is quiescent in Figs. 1-2, especially for cells that do not show firing locked to song.

      We appreciate this suggestion. Given the scope of the current revision and the primary focus on DAF-response modulation, we have elected not to add spontaneous activity examples to Figures 1-2 at this time. We agree this would be a valuable addition in future work and have noted it as a limitation in the Discussion.

      (3) Methods, p. 10: Surgery and awake-behaving electrophysiology: "The of the cannula" - this is the only mention of a cannula. Do the authors mean the ends of the probes?

      Cannula placement and wire bundle extension from the end of the cannula has been clarified in the Methods.

      (4) Bottom of p. 10: Fix reference for biorxiv paper: "ref andreas paper"

      Fixed.

      (5) Methods, p. 12: Redundant sentences regarding significant error response criteria.

      Fixed. The redundant sentences have been removed.

      Reviewer #2 (Recommendations For The Authors):

      (1) The abstract is too vaguely formulated. Authors should try to quantify the statements already in the abstract.

      We have reworded the abstract to align with the revision's more conservative claims regarding social context modulation of auditory feedback, and have added specific quantitative statements where possible.

      (2) Authors repeatedly refer to 'perceived song errors' without performing experiments or reporting on behavioral readouts of how birds perceive the jamming sounds. The wording should be changed to something more neutral, e.g. 'DAF responses'.

      We revised the manuscript throughout to use more neutral language centred on "DAF-related" or "feedback-related" responses rather than "perceived errors" or "mistakes". The title was changed from "Auditory cortical error signals retune during songbird courtship" to "Auditory cortical feedback signals are modulated during songbird courtship". We similarly revised the abstract and all relevant passages in the Results and Discussion.

      (3) Authors write that 33 neurons were DAF responsive in both conditions. How should we interpret this overlap relative to independence and identity assumptions?

      We agree that the original presentation made the interpretation of overlap across conditions unclear. The observed overlap is greater than expected under a strict independence assumption but smaller than expected if responsiveness were identical across conditions, consistent with partial but incomplete sharing of DAF responsiveness across social contexts. In the revised manuscript, however, we have moved away from this binary classification framework because DAF responsiveness appears to vary continuously across neurons. The permutation-based analysis now directly tests for changes in DAF responsiveness across contexts without requiring categorical assignment.

      (4) Only 10 neurons were not affected by courtship state or only 10 error responsive neurons were not affected? I suggest authors do a multivariate analysis or use a mixed effect model and summarize the result as a table.

      We agree that the categorical accounting of neurons across conditions was difficult to follow in the original manuscript. In the revised manuscript, we clarified the distinction between neurons responsive to DAF within a condition and neurons exhibiting significant modulation of DAF responsiveness across conditions. We now explicitly report: "This analysis identified 71/147 neurons as DAF responsive in at least one behavioral condition, whereas 76/147 were not responsive in either condition." and "This more conservative approach identified 48/147 neurons as DAF responsive in at least one condition, with 13 of those neurons exhibiting a significant modulation in their DAF response between undirected and female-directed singing."

      (5) It would help if authors could define 'z-scored difference'. Better known is d prime, is this the same?

      For each neuron, the z-scored DAF response was computed as the z-scored firing rate difference between distorted and undistorted trials. Importantly, our revised main analysis avoids any normalization such as z-scoring, and instead uses a permutation-based approach applied directly to spike counts.

      (6) Is the 'retuning' assessment a bit conservative? Neurons could also retune by showing error scores greater than 2.5 in both conditions but a shifted response time.

      We agree that neurons could retune by shifting the latency of DAF responses. Although potential latency shifts are beyond the scope of the current study, we did observe suggestive evidence of possible latency changes in some example neurons across conditions. We have noted this as an interesting direction for future analysis.

      (7) Could the stability of DAF response across trials be described? E.g. as the ratio between intra versus inter condition variability?

      We agree that stability of DAF responses across trials is an important concern. In addition to imposing strict waveform stability requirements, our permutation-based statistical test explicitly accounts for trial-by-trial variability by constructing null distributions from within-condition trial shuffles. We have also replaced the previously shown unstable example neuron with neurons that exhibit more consistent DAF-related responses across trials, and provide additional examples in Figures 2 and 3.

      Minor:

      (8) 'significant increase in burst fraction': specify effect size of t test in results section.

      We now specify in the main text: "A small but significant increase in burst fraction was observed (paired t-test, p=9.3x10-6, n=138 neurons, mean +/- SEM: 0.11 +/- 0.006 vs 0.15 +/- 0.007, Figure 1J)."

      (9) The IMCC parameter should be specified in the main text.

      The Gaussian smoothing parameter (20 ms) has now been specified in the main text.

      (10) Fig. 2: indicate the windows within which error scores are computed.

      This is no longer applicable, as the revised permutation-based analysis does not rely on scoring error responses within a fixed window.

      (11) In Fig. 2A, the neuron has an error score of -2.54 (significant), but the red and blue curves look almost the same.

      We agree that the previous error score quantification did not always capture firing rate differences in an intuitive way. This example neuron has been replaced in the revised manuscript, and the new analysis avoids scalar error scores in favor of the permutation-based approach.

      Reviewer #3 (Recommendations For The Authors):

      Minor points:

      (1) "(ref andreas paper)." Add reference here?

      Fixed.

      (2) Hessler and Doupe 1999 is a good reference for premotor signal re-tuning during courtship.

      We agree. The reference has been included in the revised manuscript.

      (3) Page 5: "discharge depended on courtship state, using" - should this be "depending"?

      The original wording was intentional: "we tested how discharge depended on courtship state." We have verified this reads correctly in context and made no change.

      (4) Page 9: "consistent with a brainwide process" - what is meant here?

      We have revised this wording. The revised manuscript replaces "brainwide process" with clearer language describing a distributed modulation of auditory responsiveness that is not confined to a single nucleus.

    1. eLife Assessment

      In this fundamental manuscript, Richter et al. present a thorough anatomical characterization of the Drosophila melanogaster larval pharyngeal sensory system, which is involved in taste-guided behaviors. This study fills a major gap in the larval sensory map, providing a compelling neuroanatomical foundation for future investigations into sensory circuits and behavior. The exceptional data will significantly enrich the field of Drosophila neurobiology.

    2. Reviewer #2 (Public review):

      Summary:

      The authors wanted to achieve a detailed ultrastructural reconstruction of the gustatory sensory organs in the Drosophila pharynx. Using serial EM and the associated bioinformatics tools they have achieved their goal.

      Strengths:

      Given the dataset, finding presented are solid and will be an important work of reference for the future.

      Comments on revised version.

      The authors have well responded to my previous comments and added text and figure material.

    3. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The authors provide a detailed ultrastructural analysis of the larval pharyngeal sensory organs, including the dorsal pharyngeal sensilla, dorsal pharyngeal organ, ventral pharyngeal sensilla, and posterior pharyngeal sensilla. Using electron microscopy and 3D reconstruction, Richter et al., present a comprehensive mapping and classification of pharyngeal sensory structures, defining the morphological type of pharyngeal sensilla based on ultrastructure and generating a neuron-to-sensillum map. These findings significantly advance our understanding of internal larval sensory systems and establish a robust framework for future functional studies in coordination with external sensory systems.

      Strengths:

      The application of high-resolution electron microscopy and 3D imaging analysis successfully overcomes technical challenges associated with visualizing deep internal structures. This enables an unprecedented level of anatomical detail of the larval pharyngeal sensory system. Thus, the study complements and completes existing maps of larval sensory circuits, contributing a comprehensive neuroanatomical characterization of larval sensory input pathways. These insights will inform future studies on larval behavior, sensory processing, and may also have applied relevance for insect control strategies.

      Weaknesses:

      While the manuscript is concise, clearly written, and methodologically rigorous, it primarily addresses a specialized readership with expertise in insect neuroanatomy.

      We thank the reviewer for the positive assessment of our study and for the helpful suggestions. In response, we have clarified the visual presentation of the pharyngeal sense organs in Figure 1, expanded the discussion of adult pharyngeal sensory systems, briefly broadened the comparison to other insect species, checked and corrected the scale bars, and added further methodological detail where appropriate.

      Reviewer #2 (Public review):

      Summary:

      This manuscript documents the structure of the pharyngeal nervous system of the Drosophila larva. The authors wanted to achieve a detailed ultrastructural reconstruction of the gustatory sensory organs in the Drosophila pharynx. Using serial EM and the associated bioinformatics tools, they have achieved their goal. The paper is written clearly and illustrated beautifully with 3D models and annotated sections. The data will significantly enrich the field of Drosophila neurobiology.

      Strengths:

      Given the dataset, the findings presented are solid and will be an important work of reference for the future.

      Weaknesses:

      Previous work, including EM, on the pharyngeal sensory organ is not sufficiently referenced and used for comparison with the data presented in this study.

      We are grateful for the reviewer’s thoughtful comments and for the suggestion to strengthen the historical and comparative context of the work. We have revised the introduction to better acknowledge and discuss the relevant previous EM-based literature on adult and larval internal gustatory sensilla, clarified the organization of the shared pore structure in T1–T3, highlighted the DPO multidendritic neurons more explicitly, and added a comparison that emphasizes the added value of the complete serial EM dataset.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) For improved clarity, highlight the pharyngeal sense organs in Figure 1B. Consider using the color schemes to differentiate between peripheral and internal sensory organs.

      We thank the reviewer for this helpful suggestion. We have revised Figure 1 to more clearly separate the pharyngeal sense organs from the external sense organs in the head region. This revision improves visual clarity and accessibility for readers.

      (2) In reference to lines 80-84, expand the discussion to address how future studies could explore the conserved morphological and functional characterization of the adult pharyngeal sensory system.

      We appreciate this suggestion and have expanded the discussion accordingly. We now briefly address how future work could compare the larval and adult pharyngeal sensory systems to examine conserved morphological and functional features.

      (3) To broaden the manuscript's appeal and emphasize its relevance beyond Drosophila, briefly discuss similarities, differences, or conserved roles of pharyngeal sensory systems in other insect species.

      Thank you for this valuable recommendation. We have added a paragraph placing the Drosophila pharyngeal sensory system in a broader insect context, including similarities, differences and potential conservation across species.

      (4) Recheck the scale bars in all figures, including the supplemental material.

      We thank the reviewer for pointing this out. We carefully rechecked all scale bars across the main and supplemental figures and corrected the missing ones.

      (5) Consider including additional details on image processing or provide appropriate citations for further reading.

      We appreciate this suggestion. We have expanded the methods section to include additional information on technical details and provide the relevant reference for further reading.

      Reviewer #2 (Recommendations for the authors):

      (1) Line 57ff: The previous literature describes internal gustatory sensilla in considerable detail.

      (a) Adult: These sensilla form three complexes, the labral sensory organ, and the ventral and dorsal cibarial sensory organ (Nayak & Singh, 1983, 1985; Singh, 1997; Stocker & Schorderet, 1981; Kendroud et al., 2017). The work by Nayak and Sing includes TEM and presents detailed EM-based schematics. This should be referenced and discussed.

      (b) Larva: Gendre et al. 2004, describes the internal gustatory organs and relates them to their adult counterparts:

      - Dorsal pharyngeal sense organ (DPS) and dorsal pharyngeal organ DPO) are the forerunners of adult labral and ventral cibarial sensory organs

      - Posterior pharyngeal sensory organ (PPS) is the forerunner of the adult dorsal cibarial sensory organ

      - Ventral pharyngeal sensory organ (VPS), derived from the labial segment, undergoes apoptosis during metamorphosis

      This work, connecting larva and adult (and containing detailed diagrams comparing adult and larval pharyngeal sensilla) should be presented in the introduction.

      We thank the reviewer for this important comment. We have revised the introduction to better cite and discuss previous EM-based studies of internal gustatory sensilla in both adult and larval stages, and we now place our findings more explicitly in the context of this prior work.

      (2) Line 180: the relationship between the ending of T1-T3 in one shared pore, and the individually wrapped sensilla should be explained; maybe a simple diagram would help. I did not understand how it works. Normally, in a gustatory sensillum, you have one or more sensory neurons, surrounded by thecogen, trichogen, and tormogen cells. The trichogen generates the shaft with the pore at its tip. Now here, in T1-T3, you have three sets of thecogen/trichogen/tormogen. Do all three trichogen cells somehow participate in the shaft with the common pore? Or only a single one, and the other two generate no shaft? It is possible this cannot be resolved, but the authors should address the problem and suggest a possible scenario.

      We appreciate the reviewer’s concern and agree that this point required clarification. We have revised the relevant text to better explain the organization of T1-T3 and their shared pore and the organization of the support cells.

      (3) Line 205: the DPO multidendritic neurons with dendrites into the hemolymph should be shown; in Figure S4G, I could see only cell bodies. These MD neurons in the gustatory system are, I believe, a true novelty and should be emphasized more if the material allows (text figure!)

      Thank you for highlighting this point. We have revised the results and supplementary material to show these neurons more clearly and to emphasize their novelty and potential relevance to the pharyngeal sensory system.

      (4) A somewhat detailed comparison between the ultrastructure of the DPS as extracted from the serial EM stack of this study, and the conclusions of Nayak and Singh 1983 as depicted in their diagram Figure 7a would be productive. The idea being: what additional details can (only) a complete EM stack provide, compared to conventional EM.

      We appreciate this suggestion. Rather than directly comparing larval and adult structures in detail, we now emphasize what the complete serial EM dataset adds beyond conventional single-section EM, namely a more comprehensive and complete reconstruction of the sensory organs and associated cell types (multidendritic neurons, papilla sensilla, and chordotonal organs that were not described before, organization of support cells)

      (5) To round off the work and connect it to the previously published analysis of gustatory terminal arborizations and connectivity in the brain (Miroschnikow et al.,2018), it would be helpful to add an analysis of the distribution of axons from the different sensilla in the nerves. Miroschnikow analyzes the central terminations of the same sense for which the peripheral structure is described here, only that in their L1 connectome, the periphery was cut off. Do the findings of the current study match their predictions, as to the number of sensory neurons, etc? It should be possible to follow, even at the lower resolution of the dataset presented here, to follow axons of sensory neurons through the nerves to the neuropil entry, and thereby make the connection. I consider this to be of great importance for the field, for authors who want to use the data of this study, and the Miroschnikow et al analysis, for their own studies.

      We thank the reviewer for this thoughtful and constructive suggestion. We fully agree that linking the peripheral sensory anatomy described in this study to the central projections analyzed by Miroschnikow et al. would be highly valuable and of broad interest. However, a systematic analysis of axon distributions from the different sensilla through the nerves to their neuropil entry points is beyond the scope of the present work. Owing especially to the dataset’s resolution and inherent limitations, tracing the connections from sensory organs through the nerves to their projections in the brain is technically highly challenging and extremely time-consuming, since much of the process would need to be performed manually. We therefore do not include a detailed comparison with the predictions from Miroschnikow et al. in this manuscript. Nevertheless, we appreciate that such an analysis would be an important next step for the field and a useful resource for future studies.

      We are grateful for the reviewers’ thoughtful feedback, which has helped us improve the manuscript substantially. We hope that the revised version addresses the concerns raised and better conveys the significance of our work.

    1. eLife Assessment

      This important study provides mechanistic evidence for how the tea-adapted Kanzawa spider mite, Tetranychus kanzawai, overcomes the catechin-based defenses of green tea plants. The work identifies the horizontally transferred dioxygenase DOG15 as a key contributor to host adaptation and supports a two-step model involving evolutionary modification of enzyme activity together with strong inducible upregulation upon feeding on tea. The evidence is convincing because comparative behavioral and toxicological assays, transcriptomic and proteomic analyses, RNAi-mediated functional validation, and recombinant enzyme assays converge to link DOG15 activity and expression with improved performance on tea. The revised manuscript appropriately acknowledges that the products of catechin cleavage have not yet been characterized and that additional detoxification pathways may contribute to tea adaptation, providing a balanced interpretation of the otherwise strong mechanistic evidence.

    2. Reviewer #1 (Public review):

      Summary:

      This study investigates the molecular mechanisms allowing the KSM mite to infest tea plants, a host that is toxic to the closely related TSSM mite due to high concentrations of phenolic catechins. The authors utilize a comparative approach involving tea-adapted KSM, non-adapted KSM, and TSSM to assess behavioral avoidance and physiological tolerance to catechins. The main finding is that tea-adapted KSM possesses a specific detoxification mechanism mediated by an enzyme, TkDOG15, which was acquired via horizontal gene transfer. The study demonstrates that adaptation is a two-step process: (1) structural refinement of the TkDOG15 enzyme through amino acid substitutions that enhance enzymatic efficiency against catechins, and (2) significant transcriptional upregulation of this gene in response to tea feeding. This enzymatic adaptation allows the mites to cleave and detoxify tea catechins, enabling survival on a toxic host plant.

      Strengths:

      A multiomics approach (transcriptomics and proteomics) provided a compelling cross-validation of its findings. Functional bioassays, such as RNAi and recombinant enzyme assays, demonstrated that the adapted mite has higher activity against catechins via TkDOG15. Other methodologies, like feeding assay using a parafilm-covered leaf disc, were effective in avoiding contact chemosensation.

      Comments on revised version.

      The authors have satisfied all previous concerns through necessary text revisions and clarified discussions. The manuscript is now well-balanced and scientifically sound.

    3. Reviewer #2 (Public review):

      Summary:

      The fascinating topic of the host range of arthropods, including insects, and the detoxification of host secondary metabolites has been elucidated through studies of the host specificity of two closely related species. The discovery that key genes were acquired from fungi through horizontal gene transfer (HGT) is particularly significant.

      Strengths:

      (1) The discovery that the TkDOG15 enzyme, acquired through HGT from fungi, plays a key role in the detoxification of green tea catechins in the Kanzawa mite, revealing a new mechanism of plant-herbivore interactions, is highly encouraging.

      (2) The verification of this finding through various experiments, including behavioral, toxicological, transcriptomic, and proteomic analyses, RNAi-based gene function analysis, and recombinant enzyme activity assays, is also highly commendable.

      (3) By proposing a two-step model in which amino acid substitutions and expression regulation of a specific enzyme gene (TkDOG15) enable host adaptive evolution, this study contributes significantly to our understanding of the evolutionary mechanisms of speciation and plant defense overcoming.

      Comments on revised version.

      I believe the manuscript has been significantly refined since the initial draft was submitted.

    4. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment:

      This important study provides mechanistic evidence that tea-adapted two-spotted spider mite overcomes green tea catechin defenses via the horizontally transferred dioxygenase TkDOG15, supporting a two-step adaptation model, combining enzyme refinement and inducible upregulation. The evidence is convincing because multi-omics signals converge with functional validation (RNAi knockdown and recombinant enzyme assays) and well-controlled behavioral/toxicity assays to link TkDOG15 activity and expression to survival and feeding on tea.

      We thank the editors and reviewers for this positive assessment of the importance of our study and the strength of the evidence. We would like to point out one factual correction. The assessment describes the tea-adapted mite as the "two-spotted spider mite" (TSSM, Tetranychus urticae), but the species adapted to tea in this study is the Kanzawa spider mite (KSM, Tetranychus kanzawai). We would suggest revising "tea-adapted two-spotted spider mite" to "tea-adapted spider mite" or "tea-adapted Kanzawa spider mite" accordingly.

      Reviewer #1 (Public review):

      Summary:

      This study investigates the molecular mechanisms allowing the KSM mite to infest tea plants, a host that is toxic to the closely related TSSM mite due to high concentrations of phenolic catechins. The authors utilize a comparative approach involving tea-adapted KSM, non-adapted KSM, and TSSM to assess behavioral avoidance and physiological tolerance to catechins. The main finding is that tea-adapted KSM possesses a specific detoxification mechanism mediated by an enzyme, TkDOG15, which was acquired via horizontal gene transfer. The study demonstrates that adaptation is a two-step process: (1) structural refinement of the TkDOG15 enzyme through amino acid substitutions that enhance enzymatic efficiency against catechins, and (2) significant transcriptional upregulation of this gene in response to tea feeding. This enzymatic adaptation allows the mites to cleave and detoxify tea catechins, enabling survival on a toxic host plant.

      Strengths:

      A multiomics approach (transcriptomics and proteomics) provided a compelling crossvalidation of its findings. Functional bioassays, such as RNAi and recombinant enzyme assays, demonstrated that the adapted mite has higher activity against catechins via TkDOG15. Other methodologies, like feeding assay using a parafilm-covered leaf disc, were effective in avoiding contact chemosensation.

      Weaknesses:

      Although TkDOG15 is assumed to "detoxify" catechins by ring cleavage, the study doesn't identify or characterize the breakdown metabolic products. If the metabolites are indeed non-toxic compared to the parent catechins, that would strengthen the detoxification hypothesis. Also, the transcriptomic and proteomic analyses identified other potential detoxification enzymes, such as CCEs, UGTs, and ABC (Supplementary Tables 3-1 & 3-2), which were also upregulated. The manuscript focuses almost exclusively on TkDOG15, potentially overlooking a multigenic adaptation mechanism, where these other enzymes might play synergistic roles, although it was mentioned in the discussion section.

      Reviewer #1 (Recommendations for the authors):

      There is no need for additional experiments, but I suggest revising the discussion section to mention the weaknesses pointed out above.

      We thank the reviewer for the positive assessment and helpful suggestions. We have revised the Discussion (L276-283) to address both points as limitations. First, we now note that we did not characterize the products of TkDOG15-mediated catechin cleavage, and that confirming their reduced toxicity relative to the parent catechins would further support its detoxification role. We note this as a direction for future work. Second, we note that DOG15 in KSM on tea was the only enzyme upregulated at both the mRNA and protein levels, whereas the CCEs, UGTs, ABC transporter, and other DOGs were enriched in only one dataset. We now state that tea adaptation in KSM may be multigenic, with these enzymes potentially acting synergistically with DOG15 and warranting functional validation.

      Minor corrections below:

      (1) Figure 1a: For better readability, I recommend adding "KSM" and "TSSM" to the two pictures, respectively.

      Done.

      (2) L165: tetur20g01790 refers to a TSSM gene, while TkDOG15 refers to a TSM protein. Revise it accordingly. (Same for L442 and L481).

      The reviewer is correct that tetur20g01790 is the TSSM gene ID. As the KSM genome is not yet available, we identified the TkDOG15 gene, the KSM ortholog of tetur20g01790, by de novo assembly of our RNA-seq reads. We have revised L169 and L455 accordingly.

      (3) L264: Supplemental Table 3-2.

      Done (L271).

      Reviewer #2 (Public review):

      Summary:

      The fascinating topic of the host range of arthropods, including insects, and the detoxification of host secondary metabolites has been elucidated through studies of the host specificity of two closely related species. The discovery that key genes were acquired from fungi through horizontal gene transfer (HGT) is particularly significant.

      Strengths:

      (1) The discovery that the TkDOG15 enzyme, acquired through HGT from fungi, plays a key role in the detoxification of green tea catechins in the Kanzawa mite, revealing a new mechanism of plant-herbivore interactions, is highly encouraging.

      (2) The verification of this finding through various experiments, including behavioral, toxicological, transcriptomic, and proteomic analyses, RNAi-based gene function analysis, and recombinant enzyme activity assays, is also highly commendable.

      (3) By proposing a two-step model in which amino acid substitutions and expression regulation of a specific enzyme gene (TkDOG15) enable host adaptive evolution, this study contributes significantly to our understanding of the evolutionary mechanisms of speciation and plant defense overcoming.

      Weaknesses:

      While transcriptome/proteome analyses reported changes in the expression of other detoxification-related enzymes, including CCEs, UGTs, ABC transporters, DOG1, DOG4, and DOG7, it is regrettable that the contribution of each enzyme, including its interaction with TkDOG15 and the functional analysis of each enzyme within the overall catechin detoxification system, was not investigated.

      We thank the reviewer for the encouraging assessment and this comment. We agree that the contributions of the other detoxification-related enzymes, including their interaction with DOG15, remain to be investigated. As this point overlaps with a comment from Reviewer 1, we have revised the Discussion (L276-283) to note that DOG15 was the only enzyme upregulated at both the mRNA and protein levels, whereas the CCEs, UGTs, ABC transporter, and other DOGs were enriched in only one dataset. We now state that tea adaptation in KSM may be multigenic, with these enzymes potentially acting synergistically with DOG15, and that the functional analysis of their individual and combined contributions warrants future work.

      Reviewer #2 (Recommendations for the authors):

      The manuscript titled "Adaptation of an Herbivorous Arthropod to Green Tea Plants by Overcoming Catechin Defenses" presents a well-designed, mechanistically insightful study that advances our understanding of herbivore adaptation to plant chemical defenses. The work is scientifically sound and of potential interest to a broad readership in chemical ecology and evolutionary biology.

      However, before the manuscript can be considered for acceptance, the authors must adequately address the comments outlined below regarding clarity, presentation, and interpretation across the manuscript.

      We thank the reviewer for the positive evaluation of our study. We have carefully addressed each of the specific comments below regarding clarity, presentation, and interpretation, and we believe these revisions have substantially improved the manuscript.

      Specific comments on each section:

      (1) Abstract

      (a) The authors are encouraged to add a concise concluding sentence summarizing the broader significance of the study and indicating potential future research directions or limitations, which would strengthen the impact of the abstract.

      We have added a concluding sentence to the Abstract summarizing the broader significance of the study and indicating future directions (L38-40).

      (b) The authors may consider adding representative quantitative results to the abstract, as this would enhance clarity and increase the impact and interpretability of the study for readers.

      We have added representative quantitative results to the Abstract. Specifically, we now state that the mRNA and protein levels of DOG15 in tea-adapted T. kanzawai are up to 31.6 and 12.1 times higher, respectively, than in T. urticae fed on tea plants (L30-32). For consistency, we now refer to the gene as "DOG15" throughout the Abstract (L29, L30, and L36).

      (2) Introduction

      (a) While the paragraph is informative, it reads more like a summary of the main results than a statement of study objectives. The authors are encouraged to reframe this section to explicitly define the study's aims and hypotheses.

      We have reframed the final paragraph of the Introduction to explicitly state the study's aims and hypotheses rather than to summarize the results (L73-81).

      (b) The authors should avoid excessive citation of multiple references for a single thematic statement when one key reference is sufficient. Where appropriate, inclusion of more recent literature is encouraged.

      We have reduced multiple citations for single statements to the most representative references: Cabrera et al. (2006) for the health benefits of catechins (L46) and Grbić et al. (2011) and Dermauw et al. (2013) for the DOG gene count (L66-67).

      (3) Materials and Methods

      (a) The Materials and Methods section is comprehensive and technically sound; however, its length and density reduce overall clarity. The authors are encouraged to streamline descriptions of standard or well-established protocols and rely on appropriate citations where possible.

      We agree that clarity can be improved by removing redundancy. The Materials and Methods are intentionally detailed to allow independent replication of our protocols, so we have retained this detail and instead removed the overlapping methodological descriptions from the figure captions, where the same information was repeated (see our response to comment 6a).

      (b) Greater consistency is needed in reporting biological and technical replicates across different experiments (e.g., performance assays, transcriptomics, proteomics, and enzymatic activity assays) to enhance reproducibility.

      We have standardized the reporting of replicates across all experiments to the format "x independent experimental runs (n = y per run)." Throughout the manuscript, "independent experimental runs" denotes biological replicates, with technical replicates specified separately where applicable (three technical replicates for qRT-PCR).

      (c) The authors should provide brief justification for key methodological parameters, such as catechin concentrations, exclusion criteria in behavioral assays, and thresholds used for defining DEGs and DEPs, to improve transparency and interoperability.

      We have added brief justifications for the three parameters. 1) The catechin concentration range was chosen to encompass the individual catechin levels measured in fresh tea leaves (L340-341). 2) In the behavioral assays, inactive mites were excluded because their movement was insufficient to determine chemo-orientation behavior, and escaped mites were excluded because they did not complete the assay (L365-367). 3) The thresholds for DEGs and DEPs follow criteria commonly applied in mite transcriptomic studies (Vidal-Quist et al., 2025, newly added to the references) (L414-416) and are consistent with our previous spider mite proteomic analysis (Arai et al., 2025) (L444-445).

      (4) Results

      (a) While significant differences in survival and fecundity are reported, briefly indicating the magnitude of these differences (e.g., percentage or fold change) would improve clarity and strengthen the presentation (Lines 91-96).

      We have added the magnitude of the differences (L94-97). The revised text now states that after 10 days, almost 90% of tea-adapted KSM survived, compared with about 5% of non-adapted KSM and 33% of TSSM, and that tea-adapted KSM laid up to about 2 eggs/surviving female daily, whereas the other two populations laid almost no eggs.

      (b) The final sentences include interpretative and concluding statements regarding catechins as key metabolites and mite adaptation. These statements would be more appropriate for the Discussion section rather than the Results (Lines 127-130). Follow the same for the rest of the Results section also.

      Following the reviewer's suggestion, we have removed the interpretive and concluding statements from the end of the Results section, so that it now reports only the observations (L129-130). The interpretation regarding the multiple modes of action of catechins and the insensitivity of tea-adapted KSM is already presented in the Discussion (L206-213 and Conclusions), so we did not duplicate it there. We also reviewed the remaining Results subsections and confirmed that they report the experimental observations and their direct conclusions without broader interpretation.

      (c) The comparison among catechin classes is clear; however, briefly listing the mean concentrations of each catechin (as shown in Figure 2a) in the text would improve readability without duplicating the figure (Lines 135-139).

      We have added the approximate mean concentration of each catechin to the text (L136-137).

      (d) Please clarify in the Results whether the same exposure concentration and duration were applied for all catechins and mite species, or explicitly direct readers to the Methods section (Lines 141-142).

      We have clarified in the Results section that all four catechins were tested at the same concentration series (0, 10, 10<sup>2</sup>, 10<sup>3</sup>, 10<sup>4</sup>, and 10<sup>5</sup> ppm) and the same exposure duration (24 h) for both mite populations (L143).

      (e) The phrase "lower sensitivity" should be explicitly linked to LC<sub>50</sub> estimates to ensure that the basis of comparison is immediately clear to readers (Lines 143-144).

      Following the reviewer's suggestion, we have linked the sensitivity comparison to the LC<sub>50</sub> values (L143-147). The comparison is now stated relative to TSSM based on the LC<sub>50</sub> estimates, and for ECg and EC we note that the LC<sub>50</sub> of tea-adapted KSM exceeded the highest concentration tested.

      (f) This section clearly identifies TkDOG15 as a key gene underlying tea adaptation in KSM; however, the authors are encouraged to briefly clarify the criteria used to define "highly enriched" mRNAs and proteins (e.g., fold-change and statistical thresholds) in the Results text or by explicitly directing readers to the Methods. This would improve transparency and facilitate interpretation of the multi-omics comparisons (Lines 147-173).

      We have added the criteria used to define the enriched mRNAs and proteins (log2 fold change ≥ 1 with adjusted p-value < 0.05 for mRNA and p-value < 0.05 for protein) and referred readers to the Materials and Methods (L159-160).

      (g) The enzymatic comparison between TkDOG15 and TuDOG15 is well presented; however, the authors are encouraged to briefly discuss whether the two amino acid substitutions (Q127A and T203A) were individually or jointly responsible for the increased catalytic efficiency, or to acknowledge this as a limitation and potential direction for future functional studies (Lines 176-190).

      We have added a brief discussion of whether the two substitutions (Q127A and T203A) act individually or jointly (L254-257). We note that T203A is adjacent to the active-site residue Y202 and may contribute more directly to catalytic efficiency, and we acknowledge that dissecting their individual contributions by site-directed mutagenesis is a direction for future work.

      (5) Discussion

      (a) The authors appropriately acknowledge that the molecular basis of chemosensory insensitivity and the contribution of additional detoxification enzymes remain unresolved. To further improve clarity, these statements could be explicitly framed as hypotheses or future research directions to clearly distinguish them from experimentally supported mechanisms (Lines 205-208; 266-270).

      We have reframed the statements on chemosensory insensitivity (L209-213) and the contribution of additional detoxification enzymes (L272-274) as hypotheses and future directions, distinguishing them from the experimentally supported mechanisms.

      (b) While DOG15 is convincingly identified as a key contributor to tea adaptation, a brief clarification of its relative importance compared with other upregulated detoxification enzymes would strengthen interpretative balance, even if the roles of these enzymes remain unresolved (Lines 259-265).

      DOG15 was the only enzyme upregulated at both the mRNA and protein levels (Figure 3d,e), and the only enzyme functionally validated in this study, by RNAi silencing (Figure 3f) and recombinant enzyme assays (Figure 4c). We have established that DOG15 contributes to tea adaptation, but because the other upregulated enzymes were not functionally tested, their relative contributions cannot be determined at this stage. As we note in the Discussion, tea adaptation in KSM may be multigenic, with these enzymes potentially acting synergistically with DOG15 (L277-283). We therefore did not add further text, to avoid duplication.

      (c) The discussion linking host plant adaptation to reproductive isolation and ecological speciation is interesting and well contextualized; however, these evolutionary implications should be slightly tempered or explicitly framed as potential long-term outcomes beyond the immediate scope of the present study (Lines 271-281).

      We have tempered the evolutionary implications (L292-294). The revised sentence now frames the link to reproductive isolation and ecological speciation as a potential outcome over longer evolutionary timescales rather than a direct finding of the present study.

      (6) Figure captions

      (a) The figure captions (Figures 1-4) are exceptionally detailed and, in several places, repeat methodological information already described in the Materials and Methods. The authors are encouraged to shorten the captions by retaining only information necessary to interpret the figures, while referring readers to the Methods for experimental details.

      We have shortened the figure captions (Figures 1-4) by removing methodological details that are described in the Materials and Methods, retaining only the information needed to interpret each figure. Where appropriate, readers are now referred to the Materials and Methods or to Supplemental Figure 1-2 for the full experimental procedures.

      (b) Several captions contain long, multi-sentence descriptions that may hinder readability. The authors may consider simplifying the wording, grouping related panels more concisely, and removing procedural details (e.g., extraction conditions, exposure durations, and instrument settings) to improve clarity and visual accessibility.

      As described in our response to comment 6a, we have simplified the figure captions by removing procedural details such as extraction conditions, exposure durations, and instrument settings, and by grouping related panels more concisely. These details are retained in the Materials and Methods.

      (c) In Figure 1, the panel labels (a-h) do not appear in a clear sequential order. For consistency with the other figures and to improve readability, the authors should ensure that panel lettering is arranged in a logical, sequential order throughout the manuscript.

      We appreciate the reviewer's attention to panel ordering. In the current layout, the panel lettering follows the order in which the panels are first cited in the text. Arranging the panels in a strict left-to-right, top-to-bottom sequence would require reducing the size of several panels, including the HPLC chromatogram in panel (e) and the survival and fecundity time courses in panels (c) and (d), which would compromise their readability. We have therefore retained the current arrangement, in which related panels are grouped together and the larger panels are kept at a legible size. We hope the reviewer finds this acceptable.

    1. eLife Assessment

      This valuable study presents a real-time system for identifying multiple unrestrained marmosets in a home cage setting using a combination of facial features and color-coded beads. While there is solid evidence that the system has a precision comparable to human experimenters in the tested scenarios, there is limited evidence that this would generalize to unconstrained multi-animal environments

    2. Reviewer #1 (Public review):

      The manuscript by Yang, Wang, and Cléry presents a pipeline for real-time identification of common marmosets in a laboratory setting. Models were trained and evaluated on data derived from a family of three closely related adults and a set of juvenile twins. Freely moving animals entered an enclosed space fixed to the housing cage door, which permitted the entry of individual animals for data acquisition. Utilizing YOLOv8-nano, identification was improved through the introduction of uniquely colored collar beads. Analyses of facial similarity showed close morphological relatedness amongst individuals and highlighted the need for highly discriminative classification. The authors demonstrate that combining facial detection with visual markers enables adequate identity assignment under controlled laboratory conditions with minimal cross-individual misclassification.

      The main strengths are that the proposed pipeline offers a solution for real-time identity tracking in common marmosets. Its lightweight design enables deployment across a wide range of hardware configurations. Furthermore, if similar strategies are employed, this methodology is likely adaptable for other species with minimal modification. Additionally, evaluation of closely related individuals provides a necessary stress test for the discrimination of facial identity tracking. However, the main weakness is the pipeline's reliance on controlled animal isolation and small visual markers, which raises questions about the approach's generalizability to unconstrained multi-animal environments. The authors justify the use of beads, but the dependency of facial recognition on the beads needs to be described more clearly, as it is unclear how independent facial recognition performance truly was. The overall utility of this approach therefore remains to be seen.

    3. Reviewer #2 (Public review):

      Summary:

      In this study, Yang et al. develop a real-time system for automatic face detection and identification of multiple unrestrained common marmosets in a home cage setting.

      Strengths:

      The study aims to address an unmet need in behavioral neuroscience: the ability to non-invasively identify animals is crucial to the automated and rigorous study of neural behaviors; this is especially true for common marmosets, which are rapidly becoming a model system of choice for the study of complex social cognition. By using a YOLOv8 backbone, the study achieves human level performance, both in terms of precision and recall of the trained models.

      Weaknesses:

      The robustness of the system is not clear from the limited datasets presented.

      Comments on revised version.

      The authors have adequately addressed my comments from the previous round, and I have no further comments

    4. Reviewer #3 (Public review):

      Summary:

      In the revised manuscript, the authors provide additional details and evidence regarding the robustness and utility of their method.

      Strengths:

      (1) The authors provide a very precise automatic identification of marmosets in their home cage, to levels comparable to animal health professional.

      (2) This method is robust across lightning, camera angles etc but importantly is able to identify marmosets in naturalistic conditions, which can be of tremendous value to neuroscientists and to ecological or behavioral studies.

      (3) Easy to use and implement, requiring minimal settings. Phone videos can even be used.

      Weaknesses:

      While the manuscript improved tremendously from the previous version, given the nature of the paper, it is still a strenuous read.

      Comments on revised version.

      The authors did a good job of addressing my previous concerns and I don't have more comments.

    5. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary: 

      The manuscript by Yang, Wang, and Cléry presents a lightweight pipeline for real-time identification of common marmosets in a laboratory setting. Models were trained and evaluated on data derived from a family of three closely related adults and a set of juvenile twins. Freely moving animals entered an enclosed space fixed to the housing cage door, which permitted the entry of individual animals for data acquisition. Utilizing YOLOv8-nano, identification was improved through the introduction of uniquely colored collar beads. Analyses of facial similarity showed close morphological relatedness amongst individuals and highlighted the need for highly discriminative classification. Overall, the authors offer a framework for identity tracking that prioritizes real-time inference. The authors demonstrate that combining facial detection with visual markers enables adequate identity assignment under controlled laboratory conditions with minimal cross-individual misclassification. 

      Strengths: 

      (1) The proposed pipeline offers a solution for real-time identity tracking in common marmosets. Its lightweight design enables deployment across a wide range of hardware configurations. Furthermore, if similar strategies are employed, this methodology is likely adaptable for other species with minimal modification. 

      (2) Evaluation of closely related individuals provides a necessary stress test for the discrimination of facial identity tracking. 

      Weaknesses: 

      (1) The pipeline's reliance on controlled animal isolation and small visual markers raises questions about the approach's generalizability to unconstrained multi-animal environments. The provided confusion matrices (Figures 6-8) indicate that the most common misclassifications are background-related, possibly suggesting that detection specificity is the primary source of error. All things considered, these findings raise concerns about performance in its use in socially dynamic and visually complex environments. 

      Thank you for the comment. The background column of the confusion matrix can be explained by several occasions: a) the model detects an object where there is no object, b) there is more than one prediction label for the same object, or c) an object appeared in the image but not manually labeled, however the program was able to detect that object. The value of the background column does not necessarily mean that the detection is incorrect, as the precision score for the detection labels are good. We have rephrased the relevant sections for clarification to include the sources of the increased value in background columns in confusion matrices, as follows:

      “The background class of the confusion matrix showed frequently predictions as marmoset faces and collar beads for the training (Figure 6A) and validation set (Figure 6B). However, it does not necessarily indicate incorrect predictions or misclassifications. Instead, these values were mostly explained by multiple detections of the same object class. For instance, additional marmoset faces were predicted when multiple animals were present within a single video frame. The long collar structure or motion blur of the marmosets could also cause multiple detections of beads that belong to the same collar. This also corresponded to the high precision and recall scores observed across prediction classes (Figure 5D), suggesting that the increased background false positives were mainly related to the object-count discrepancies, instead of poor detection performance.”

      Prediction misclassification is one source of the background false positive. The misclassification could not be avoided in automatic prediction algorithm, but we included the manual filtering and majority-voting during our real-time classification to reduce this effect. Multiple detection of the same class may also be considered as the background, since only one object may be labelled in the ground truth, such as multiple collar beads or automatic face extraction. In addition, blurry objects were not labeled manually during training but can be detected during prediction, which also resulted in background false positive. It was clarified in the main text as follows:

      “The normalized confusion matrices showed high accuracy and consistency of most marmoset faces and collars detection in training (Figure 8A) and validation (Figure 8B) tests, with some exceptions. Particularly, the background was frequently identified as the collar of Young2 marmoset. This elevated background score was likely contributed by the multi-color design of the Young2 marmoset collar, making it more difficult to distinguish compared to collars with a single bead color. In this occasion, if one bead is occluded, blurred, or outside the field of view, the other visible collar bead could affect the prediction and lead to an incorrect identification from the ground truth.”

      (2) The manuscript claims performance comparable to that of human experimenters but provides no explicit evidence to support these claims. While it is plausible that human experimenters may be less accurate in facial recognition tasks involving closely related marmosets, the authors don't provide evidence. Moreover, while that might be the case, the color-coded beads provide a salient identity cue for the model, which complicates the interpretation of this comparison grounded in facial recognition. 

      Thank you for pointing out this concern. The aim of the facial recognition tool is to collect data from marmosets without having experimenters to check the identity continuously. The program is not aimed at outperforming the experimenters’ role but avoid having constant human intervention that can disrupt a more ecological in cage data collection. It is also essential for having more flexibility to collect data in case a specific experimenter is not here and thus to not disrupt the project. Human experimenters have extensive experience closely working with marmosets, having the unique collar beads associated to each marmosets allows human experimenters to hardly make mistakes identifying marmosets and to do it quickly. We collected identification accuracy of human experimenters by presenting 10 clips of the five marmosets involved in the manuscript (2 clips per marmosets), with 2 random clips repeated twice. The results were plotted by each experimenter. The identification accuracy of the experimenters correlates with the time spent with the animals, as the animal health technicians (responsible for daily health check and husbandry) achieved 95.83% average accuracy in identifying the marmosets. We clarified those points in the main text as follows:

      “Its automated pipeline substantially reduces the time and work required for traditional manual identity labeling, while maintaining an expert-level human performance and reproducibility across experimenters (95.83% average accuracy for animal health technicians, responsible for daily health check and husbandry while lab experiments ranges between 25 to 80% of accuracy depending on the amount of time spent with each animal, Supplementary figure 1). The tool’s advantages are particularly efficient for large datasets and longitudinal studies, where manual identity labeling becomes difficult, as variability and errors increase along with dataset size and experimenter number.”

      We filtered the prediction of the collars, and the identification result solely based on the faces for the 2 young marmosets was correct. The prediction results were plotted on Video 7, Video 7—video supplement 1, and Video 7—video supplement 2 and added to the Results section as follows:

      “We tested the prediction performance without collar and its longitudinal application, using the face-only prediction on the young marmosets at 11 months and 16 months (Video 7, Video 7—video supplement 1, Video 7—video supplement 2). The face classifier correctly identified the young twin marmosets solely based on their facial features, indicating that facial identity classification was performed independently of collar information and that the collar beads acted only as an auxiliary confirmation rather than the main classifier of the system (Video 7).”

      Explanation for classification of marmoset faces and collar beads in the Discussion section:

      “Facial features serve as the main and intrinsic biometric identifier for each marmoset, providing a unique source of individual recognition. Since collar-based confirmation could be affected by visibility limitations, we implemented the uniquely color-coded bead collars as an auxiliary cue to provide additional confirmation in identity prediction. For example, this issue can be caused by identical or similar bead colors between individuals (Video 5, 6) and beads that are occluded by fur (Video 3 - 6). In addition, collar beads may change over time or not be worn by all animals.”.

      “With one separated model trained per family unit, our system can utilize distinct collar colors as an additional identifier when available, while facial features performed as the main biometric marker. Even though multiple marmosets with visually similar faces may present close to the camera, the additional collar information can improve confidence in identity prediction without replacing facial recognition as the primary mechanism of identification (Video 7).”

      Reviewer #2 (Public review):

      Summary: 

      In this study, Yang et al. develop a real-time system for automatic face detection and identification of multiple unrestrained common marmosets in a home cage setting. 

      Strengths: 

      The study aims to address an unmet need in behavioral neuroscience: the ability to non-invasively identify animals is crucial to the automated and rigorous study of neural behaviors; this is especially true for common marmosets, which are rapidly becoming a model system of choice for the study of complex social cognition. By using a YOLOv8 backbone, the study achieve human level performance, both in terms of precision and recall of the trained models.

      Weaknesses: 

      The robustness of the system is not clear from the limited datasets presented. The use of color-coded beads undercuts the study's premise that the system achieves truly non-invasive tracking. Although the system achieves good performance in face detection, it does not perform as well for classification using faces alone (especially when the faces are similar, as in twin animals). Here, too, the color-coded beads play a key role in identity discrimination. The stated goals of the study and the actual results presented are therefore at odds.

      Thank you for the comment. First, we would like to clarify the role of the collar beads in our system. Compared to the faces, a unique identity marker, the collar beads were not used as the main identity classifier but rather as an external visual marker. The color-coded bead was not used solely for the purpose of marmoset video classification; it was also used as an additional source of identification for one marmoset. As the marmosets usually move very fast inside the cage, it is mainly used as a visual marker for experimenters to recognize them in a distance in a short time.

      The mislabelling is more frequent with the young twins not only due to their face similarity, but also due to the limited number of images being used for the model training, as discussed in the paragraph #4 of the Discussion section. Collar beads are small and less frequently detected by the camera, since it could be occluded by the marmoset fur. In addition, it was invisible to the camera if the marmoset turned sideways or was far from the camera. Therefore, higher weight was assigned to the beads due to their small size and less frequent detections compared to face labels, such that it was only an element to confirm the identity, instead of the main classifier.

      The inclusion of the collar beads doesn’t affect the prediction results of the marmoset faces. The model achieved a good precision/recall score for the identity labeling in the manuscript. In the revision, with the majority-vote strategy, we filtered the detection of all collar beads and showed that the model was able to correctly identify the marmosets solely by their faces. The Results section has been modified as follows:

      “We tested the prediction performance without collar and its longitudinal application, using the face-only prediction on the young marmosets at 11 months and 16 months (Video 7, Video 7—video supplement 1, Video 7—video supplement 2). The face classifier correctly identified the young twin marmosets solely based on their facial features, indicating that facial identity classification was performed independently of collar information and that the collar beads acted only as an auxiliary confirmation rather than the main classifier of the system (Video 7).”

      Explanation for classification of marmoset faces and collar beads in the Discussion section:

      “Facial features serve as the main and intrinsic biometric identifier for each marmoset, providing a unique source of individual recognition. Since collar-based confirmation could be affected by visibility limitations, we implemented the uniquely color-coded bead collars as an auxiliary cue to provide additional confirmation in identity prediction. For example, this issue can be caused by identical or similar bead colors between individuals (Video 5, 6) and beads that are occluded by fur (Video 3 - 6). In addition, collar beads may change over time or not be worn by all animals.”.

      “With one separated model trained per family unit, our system can utilize distinct collar colors as an additional identifier when available, while facial features performed as the main biometric marker. Even though multiple marmosets with visually similar faces may present close to the camera, the additional collar information can improve confidence in identity prediction without replacing facial recognition as the primary mechanism of identification (Video 7).”

      Reviewer #3 (Public review):

      Summary: 

      In this manuscript, Yang et al introduce a new method for automatically identifying marmosets in their home cage using a supervised deep learning method that recognizes the face and colored beads on marmoset collars. The authors show a high precision rate of identifying marmosets to levels comparable to a human experimenter. The method overall seems robust at identifying marmosets at different life stages and different settings; however, given the current form, I'm struggling to see the generalizability and experimental utility of this method. 

      Strengths: 

      (1) The authors provide a near-perfect automatic identification of marmosets in their home cage. 

      (2) This method is robust across lightning, camera angles, etc., making it potentially useful for marmoset (and other NHP) identification outside the housing cage as well 

      Weaknesses: 

      (1) Despite the almost perfect precision, in its current form, I'm failing to see how this method can be useful to other labs. 

      Thank you for your comment. This Tools & Resources paper mainly described the development of the marmoset identification program and methods. Future work will focus on extending the program application on identification from different housing conditions, in combination with various behavioral tasks such as in-cage touchscreen system or manual tasks, and in the wild that precludes handling or isolation of marmosets for collecting behavioral data. The program solely requires a camera, a computing device, and marmosets, as there are no hardware restrictions. In addition, we are currently collaborating with other labs on the marmoset identification from videos taken from other setups. The program achieved effective face extraction from the marmoset in the video, without the need for additional program modifications.

      (2) This is a nice methods manuscript, but the authors do not present results to show how their method can be used outside of identifying marmosets inside their home cages in a small field of view. 

      Thank you for your feedback. The method developed was applied in combination with other touchscreen behavioral tasks, aiming to extract data without human intervention. This approach was discussed in paragraph #6 in the Discussion section. While this manuscript focuses on the methods of close-view face identification when marmosets perform behavioral tasks, the identification and automatic face extraction program could also be applied to marmoset videos taken from a larger view, including phone cameras. Even though the marmosets are still housed in their home cage, the example videos presented the program’s application in a larger field of view. We have added examples of the videos/photos from a different experimental setup to respond to this comment in the Discussion section as follows:

      “The motivation for this real-time marmoset identity recognition program was to develop an easy-to-use, generalizable pipeline that could be applied across different marmosets and lab environments, such as using larger field of view or phone cameras (Figure 10).”

      (3) Reading the manuscript is strenuous, given its repetitive nature. Consolidating and shortening the results, as well as adding some definitions to the results section, would be helpful. 

      Thank you for pointing this out and your suggestions. We have rephrased the Results section for simplification to facilitate the understanding of the manuscript.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) The weight of the color-coded beads was increased to improve identification accuracy. From a brief look at the code provided on GitHub, the weight assigned to the beads seems substantial. This calls into question the need to use facial recognition in the identification strategy. As the code currently stands, facial identity appears to serve primarily as a fallback when bead detection fails to register. To strengthen the methodological justification, the paper would benefit from the authors providing a rationale for choosing this weighting scheme and, if available, supplemental figures showing performance across a range of different weights to demonstrate why that specific value was assigned in the algorithm.

      We have added the model prediction results without the collar beads showing that the facial recognition algorithm works even without the collar beads and that those collar beads are not the main classifier. The Methods section has been modified as follows:

      “For each detected bounding box, the scripts returned a corresponding label of marmoset face and collar bead color. We assigned the detected collar beads as the corresponding marmoset identity with a higher weight, which improved the detection confidence across frames.”

      The Discussion section has been modified as follows:

      “Facial features serve as the main and intrinsic biometric identifier for each marmoset, providing a unique source of individual recognition. Since collar-based confirmation could be affected by visibility limitations, we implemented the uniquely color-coded bead collars as an auxiliary cue to provide additional confirmation in identity prediction. For example, this issue can be caused by identical or similar bead colors between individuals (Video 5, 6) and beads that are occluded by fur (Video 3 - 6). In addition, collar beads may change over time or not be worn by all animals.”

      “With one separated model trained per family unit, our system can utilize distinct collar colors as an additional identifier when available, while facial features performed as the main biometric marker. Even though multiple marmosets with visually similar faces may present close to the camera, the additional collar information can improve confidence in identity prediction without replacing facial recognition as the primary mechanism of identification (Video 7).”

      The weight assigned to the beads in the GitHub page is the highest weight that we would suggest. The actual weight can be customized by the experimenters based on the actual experimental setup. For example, we used the weight of 2 in our real-time version of the marmoset face identification, while marmosets were presented with their corresponding tasks once identified. The GitHub page has been edited to clarify this point.

      (2) The overall utility of this approach, other than the real-time detection component, needs more clarification. It is currently unclear why this approach, and in which specific experimental or observational settings, is particularly advantageous compared to existing methods for assigning animal identity.

      In addition to the advantages mentioned in paragraph #1 of the Introduction and paragraphs #1-3 in the Discussion, we have added more details in the Discussion section:

      “While existing marmoset identification approaches usually utilize visible markers, Radio Frequency Identification (RFID), or observation, the manual works and human interventions involved can impact animal behaviors, especially during their behavioral task performance. The facial identification tool aims to collect data from marmosets without having experimenters to check the identity continuously, instead of outperforming the experimenters’ role.”

      (3) Although it appears that performance based on faces and color beads was evaluated separately, this was not clearly presented, leading to confusion about whether face detection performance also benefited from color beads on the animals.

      Prediction of the different labels in the same model is independent, so the prediction of color beads is not affecting the prediction results of marmoset faces. Correct identity classification could be achieved without depending on the color beads, as we have filtered out the color beads detection class. The Results section has been modified as follows:

      “We tested the prediction performance without collar and its longitudinal application, using the face-only prediction on the young marmosets at 11 months and 16 months (Video 7, Video 7—video supplement 1, Video 7—video supplement 2). The face classifier correctly identified the young twin marmosets solely based on their facial features, indicating that facial identity classification was performed independently of collar information and that the collar beads acted only as an auxiliary confirmation rather than the main classifier of the system (Video 7).”

      Reviewer #2 (Recommendations for the authors):

      (1) I found the paper quite confusing as written. The term "model" is overused and highly conflated: there are the YOLOv8 pre-trained models, the "face classification" model, and the "automatic facial and identity extraction" model. The flowchart in Figure 2A is equally confusing. The mapping from the flowchart to the results is not straightforward, and I needed several passes to grasp it. I would recommend that the authors simplify the terminology and the mapping of the methods to the results.

      Thank you for pointing this out! The YOLOv8 pre-trained models, the "face classification" model, and the "automatic facial and identity extraction" model were indeed separately trained object detection models. They all have different weights/parameters but share the same YOLOv8 architecture/backbone. We removed some of the “model” term in the manuscript and replaced them with “classifier/framework/pipeline” to avoid misunderstanding. This information has been clarified in the revised manuscript of the Methods section and Figure 2, which provides an overall clearer explanation of the workflow of the methodology of the program.

      (2) It is not clear how robust these results are, given the limited data sets analysed.

      We agree that only five marmosets were involved in this manuscript, this unfortunately limited the robustness of the prediction results. Indeed, the limited number of animals that can be used per study has been a main limitation in non-human primate research, as they are very valuable animal models. However, we included approximately 3400 images in the training dataset, which were collected across days. New videos and photos that were captured from different devices were also used in the testing to ensure that the program can be used on new marmosets, different housing cage, and from different recording devices as indicated here:

      “The motivation for this real-time marmoset identity recognition program was to develop an easy-to-use, generalizable pipeline that could be applied across different marmosets and lab environments, such as using larger field of view or phone cameras (Figure 10). The pipeline was designed to have no specific hardware requirements and can be implemented for any standard recording device, including any commonly available cameras, primate chair system, and computer-based device.”

      (3) There are two paradoxes regarding the stated motives of the study:

      (a) If the objective was to truly use non-invasive methods for the identification of animals, then why use the color-coated beads?

      As mentioned previously, identity detection can be made without collar beads, still with correct prediction results as indicated here:

      “We tested the prediction performance without collar and its longitudinal application, using the face-only prediction on the young marmosets at 11 months and 16 months (Video 7, Video 7—video supplement 1, Video 7—video supplement 2). The face classifier correctly identified the young twin marmosets solely based on their facial features, indicating that facial identity classification was performed independently of collar information and that the collar beads acted only as an auxiliary cue rather than the main classifier of the system (Video 7).”

      The color-coated beads are used for easier and quick marmoset identification during daily care, health check, for weekend staff, training or handling.

      (b) If the objective was to achieve high identification performance, and the color-coated beads are sufficient for this purpose, then why bother with faces at all?

      Collar beads are small compared to the face, and less visible due to fur occlusion and motion blur. Moreover, it is possible that some marmosets do not have collar beads due to their young age or when involved in other procedures such as imaging scans. The collars need to be checked and changed regularly in growing marmosets and it is not always convenient (some marmosets do not support the collar, some can have sensitive skin that would lead to abrasion) thus the need to develop a facial recognition system. Furthermore, marmosets who are from other labs or in the wild might not wear a collar with colored beads, thus face is the main classifier in this model to be more generally applicable. It is highlighted here:

      “Facial features serve as the main and intrinsic biometric identifier for each marmoset, providing a unique source of individual recognition. Since collar-based confirmation could be affected by visibility limitations, we implemented the uniquely color-coded bead collars as an auxiliary cue to provide additional confirmation in identity prediction. For example, this issue can be caused by identical or similar bead colors between individuals (Video 5, 6) and beads that are occluded by fur (Video 3 - 6). In addition, collar beads may change over time or not be worn by all animals.”

      (4) I was puzzled by the face similarity results in Figure 9. It appears that the face similarity measures were stronger (higher cosine similarity, lower Euclidean distance) for the adult data set compared to the twin data set. If so, why was it more challenging for the system to handle the twin data set?

      Face similarity can only be compared within models (therefore within adults and within twins). As this is calculated from different models, the adult face similarity cannot be compared with twins’ face similarity. It has been clarified in the Methods section as follows:

      “Statistical tests were performed only within the face classifier of each marmoset family, as embedding spaces may vary in scaling, learned features, and baseline metrics making cross-model comparison of inter-individual face similarity unreliable (Bollegala, 2017).”

      And in the Results section as follows: “We performed the statistical tests only on the face classifier for the adult marmoset family, as the twin marmoset model only involved two individuals and thus not valid for within-model statistical analysis (Table 2, 3).”

      The twin dataset aims to represent a test for the program utility in new marmosets, especially for testing if the program can still distinguish between the marmosets with similar faces. Thus, the number of twin data collected is less than the adult dataset, as explained by Discussion paragraph #2 “While comparing between the adult and young marmoset datasets, we found that the adult marmosets’ face classifier, trained with a larger number of varied images, showed more reliability and efficiency in marmoset identity recognition.” This explains the challenge the system faces when differentiating the twins, while increasing the training dataset is required to solve this issue.

      Reviewer #3 (Recommendations for the authors):

      Major issues:

      (1) My main issue is regarding the utility of this method in scientific experiments. This manuscript is a "methods paper" introducing a face recognition method to identify a single marmoset in their home cage in a very specific and confined field of view. This comprises a limitation on what experiments can be performed using this method. On the contrary, if (a) the authors can show that this method can be used for a bigger field of view, where the social structure/interactions can be studied for neuroethological, cognitive or social studies that will make this method significantly more robust; or (b) design an experiment that can be performed using the current method to show that this method in its current form is sufficient.

      (a) Our method worked in larger home cage (larger view) with videos taken inside the cage / outside the cage, with multiple marmosets moving around, while the camera and its fixation are also moving. A new figure (Figure 10) has been added to highlight this wide application:

      “The motivation for this real-time marmoset identity recognition program was to develop an easy-to-use, generalizable pipeline that could be applied across different marmosets and lab environments, such as using larger field of view or phone cameras (Figure 10). The pipeline was designed to have no specific hardware requirements and can be implemented for any standard recording device, including any commonly available cameras, primate chair system, and computer-based device.”.

      (b) We are currently using this method to collect in-cage touchscreen data with multiple marmosets without the need to isolate such animals to acquire the data, avoiding social separation. The collection of data in nonhuman primates is still a long process, so we wanted to share the facial recognition system first, aligned with our commitment towards open science, to benefit the broader community (we have already been contacted by two labs since the publication of this preprint) while we keep collecting data for the scientific project. We have added the touchscreen application as example in the Discussion section as follows:

      “Once trained, the system operates automatically to collect real-time identity and can work to present subject-specific behavioral or cognitive tasks based on the identity of the detected animal, with no work or presence needed on the user’s end. This tool has already been implemented in touchscreen-based marmoset cognitive tasks, including pairwise visual discrimination paradigm.”

      (2) The authors claim a longitudinal identification of marmosets, yet I think the data to fully support this are deficient. This might be a result of unclarity of this experiment. How was this experiment done? Was the training done on the 7 months and then applied to the 11 months? Are there more continuous data that track the precision of the identification in time? For example, how does the twin identification evolve in time?

      This Tools & Resources paper mainly described the development of the marmoset identification program and methods. Ongoing work in the lab, the main research focus of which is the longitudinal assessment of cognitive functions, either during neurodevelopment or in preclinical ageing model, is benefiting from such algorithms to help identifying the animals to collect in cage behavioral data. As such, we have done some testing in one young cohort. The training of the young marmosets’ identification was done only on the 7-month data, and then we applied the identification program to the videos of the same marmosets when they were 11 months old and 16 months old (for the no-collar results) as indicated as follows in the Methods section:

      “Moreover, we evaluated the model performance and its generalization across developmental stages using new videos: 1) from the adult marmosets and 2) from the same young marmosets at 11 months, which were not involved during initial program training”.

      And in the Results section as follows: “We tested the prediction performance without collar and its longitudinal application, using the face-only prediction on the young marmosets at 11 months and 16 months (Video 7, Video 7—video supplement 1, Video 7—video supplement 2). The face classifier correctly identified the young twin marmosets solely based on their facial features, indicating that facial identity classification was performed independently of collar information and that the collar beads acted only as an auxiliary cue rather than the main classifier of the system (Video 7).”

      The identification program was shown to correctly identify the marmosets; however, we found that “While comparing between the adult and young marmoset datasets, we found that the adult marmosets’ face classifier, trained with a larger number of varied images, showed more reliability and efficiency in marmoset identity recognition.”

      The mislabeling was more frequent with the young twins not only due to their face similarity, but also due to the limited number of images being used for model training. The face images used in the identification model training were less compared to the adult model, which contributed to a less accurate prediction result. As the marmoset is still developing before adulthood, their face features will become more different as they age. By increasing the number of training images from different ages of the young marmosets, this could be solved as it is therefore possible to build efficient identification program for longitudinal study. Thus, instead of the current classifiers presented in this manuscript, we suggested that the method/tool could be beneficial for longitudinal studies, not restricting to the individual-based identification program mentioned in this manuscript, as “The tool’s advantages are particularly efficient for large datasets and longitudinal studies, where manual identity labeling becomes difficult, as variability and errors increase along with dataset size and experimenter number.”

      (3) How does this method compare to other methods that were used in the past?

      The advantages were mentioned in paragraph #1 of the Introduction and paragraph #1-3 in the Discussion. Current approaches for marmosets are usually visible markers (ear dye, collar, etc.), RFID, or observation, of which manual works and human interventions are required. These methods usually need continuous adjustment due to tighter collar, dye fading, etc. This can affect marmoset behaviors, especially during their behavioral task performance, as mentioned as follows:

      “While existing marmoset identification approaches usually utilize visible markers, Radio Frequency Identification (RFID), or observation, the manual works and human interventions involved can impact animal behaviors, especially during their behavioral task performance. The facial identification tool aims to collect data from marmosets without having experimenters to check the identity continuously, instead of outperforming the experimenters’ role.”

      (4) Did the authors think of adding a continuity or a space constraint? For example, video 6 shows misidentification of the twins; in this specific case, adding a probabilistic continuity or space constraint that will limit identity switches might be useful. This can also be using a retroactive correction - for example, video 3.

      We would like to thank the reviewer for this suggestion. We agree that these approaches will be valuable improvements for future offline analysis.

      The probabilistic continuity constraint can indeed help decrease identity switches. In our current application, we have implemented a temporal smoothing through majority voting across a 30-frame (1 second) window, of which the program outputs the most frequent prediction of identity. With this strategy, we could reduce the occasional frame misprediction and maintain the real-time performance. Our animals are free-moving and may appear in any location within the camera field of view and housing cage. Therefore, position is not strongly associated with the identity of individual.

      We agree that retroactive correction could improve the detection consistency for offline analysis by correcting past detection by future prediction results. However, the current pipeline is incorporated with behavioral tasks, meaning that the prediction results aim to be transmitted with minimal time delay. As additional frame analysis and extra computational power may be needed for retroactive correction, the increased latency can be limited to the utility of real-time system and task control.

      Minor issues:

      (5) In its current form, I think the manuscript can be significantly shortened and the results/figures can be consolidated (confusion matrices with validation figures for example).

      We have followed the reviewer’s suggestion and shortened the Methods and Results sections.

      (6) The term "unseen" that the authors use in their results is confusing. Are the authors referring to monkeys that are hidden from their view, or "unseen" before by the model? The video indicates the latter, but I think the term can be changed to something less confusing, like novel, new, etc.

      We have changed the term “unseen” by “new” to avoid confusion.

      (7) Can the authors add information about the relationship between the number of manually labeled images and the identification precision?

      The relationship between number of manually labeled images and identification precision has been described in the Discussion section as: “Moreover, the performance of the system is strongly dependent on the amount and variability of the training data, with identity classification improving as more marmoset images are involved in the model training.”

      This means that more manually labeled images (i.e. larger training datasets) could improve the identification precision. However, model performance will plateau regardless of training dataset size, referred in the Results: “Each of the models was trained until reaching the early stopping criteria (i.e. no improvement within the last 100 training epochs).”

      More manually labeled images could help improve the variability of model prediction, but too many of these images are also risky for overfitting. In this case, overfitted model might not be able to make valid predictions on new videos/images.

      (8) Can the authors expand a bit about the difference between YOLO Nano, small, and medium in the methods?

      Thank you for the suggestion. We have added a brief description of the pre-trained models in the methods: “These pre-trained models share the same object detection backbone but differ in number of parameters and computing power. Larger models, such as YOLOv8 medium, provide higher detection accuracy but require greater computational resources and longer inference time. In contrast, smaller models prioritize the computational efficiency.”

      (9) Clear and short definitions of what IoU, Recall, F1, and other terms represent should be added to the results section (not formulas, short sentences).

      The definitions and formula for the evaluation metrics were described in detail in the Methods section. To help readers while avoiding repetitions with the detailed methodology definitions, we have added a brief description of these terms at their first mention in the Results section: “Model performance included the precision (the proportion of correct positive predictions), recall (the proportion of corrected predicted ground-truth labels), and mAP@50–95 (the average detection accuracy across different IoU object localization thresholds; see Methods and Materials section for detailed definitions).”.

      (10) In the methods-"video collection" section, can the authors please include more information? Is this a motion-sensitive camera? Otherwise, what's the size of the data that is collected? This will help in reproducibility and system requirements. If this is not continuously collected data, discuss what can be done to make an online identification tool.

      We have included the camera as industrial color/RGB camera, which functions like any webcam and has no motion-sensitive functions. The size of data collected was described in detail in the Methods section that we slightly modified for clarity for: “Three adult marmosets from one family were recorded for 1 hour across each of the 5 recording days, with unrestricted voluntary access to the primate chair space. For the adult marmosets, the housing cage door was opened at the beginning of the recording session allowing them to enter and exit freely into the primate chair space for food rewards and observation (Figure 1C). Two young marmosets were briefly isolated and recorded separately for testing and improving the automatic face extraction program. We recorded them at two developmental time points, 7 months old and 11 months old (an additional time point at 16 months old has been added for one marmoset to test the identification without collar). During video collection, a sliding panel and an in-cage box were positioned near the housing cage door to temporarily isolate individual marmosets from other family members. Individual isolation was kept brief (approximately 10 minutes) to prevent disturbance and potential stress due to family separation. “To capture sufficient variability in postures, individuals, and lighting conditions, clips were sampled throughout the adult marmoset videos (approximately 5 hours) (Figure 2B).”

      The size of the training dataset was also described in detail in the Methods section, referred as: “To minimize image computations and data storage, we created a dataset of 2498 annotated images from the three adult marmosets. All images were manually annotated to label marmoset faces, individual identities, and the collar bead colors (Figure 2C). The annotated images were used for training models of multi-marmoset face classification and the automatic identity extraction, which can automatically detect, localize, and identify marmoset faces (Figure 2A, Step 1 – 4). We created another dataset of two young marmosets at 7 months old (total images = 502) for testing the automatic facial and identity extraction (Figure 2A, Step 4 – 5). For both adult and young marmoset datasets, images were randomly divided into a training set and a validation set at a ratio of 8:2.”

      Our manuscript is not describing an online identification tool (i.e. the described program does not require connection to internet). Instead, once trained and the program is performing well with new marmoset videos, we could use the trained weights for real-time marmoset identification (no need to collect new training data) as we are doing it for our touchscreen data collection in cage. It is referred in the Discussion as follows: “Once trained, the system operates automatically to collect real-time identity and can work to present subject-specific behavioral or cognitive tasks based on the identity of the detected animal, with no work or presence needed on the user’s end. This tool has already been implemented in touchscreen-based marmoset cognitive tasks, including pairwise visual discrimination paradigm.”

      This program can be used for online applications if you are using a camera that is connected 24/7. For now, we are only using it while using our behavioral testing chair due to limitations issues (safety recordings, removing all electrical apparatus during night, overheating of camera if use continuously).

      (11) Would increasing the size of the beads help with their identification? In the images included in Figure 2C, it's very difficult to see these beads.

      Yes, collar beads can be occluded by fur, blurred by motions, or outside the field of camera view. We have now included increasing the collar beads size in the Discussion, referred as: “An alternate experimental solution is to improve collar visibility, including using distinct color code across individuals within a family, increasing the size of the beads, or increasing collar beads number to reduce occlusion.” However, the size of the beads needs to be appropriate to avoid being inconvenient and disruptive to the animals to ensure their welfare.

      (12) It's difficult to understand the setup of the camera in regard to the housing cage. Can the marmosets go into the primate's chair at any point (from the videos, it seems so, but the dashed line in 1C might indicate otherwise), or is the primate chair there only to mount the camera? Consider redrawing 1C in a more clear way.

      The figure 1C is a simplified drawing of the photo 1A, we have edited the drawing of Figure 1C to highlight the free access when the chair is mounted to the cage as stated here: “For the adult marmosets, the housing cage door was opened at the beginning of the recording session allowing them to enter and exit freely into the primate chair space for food rewards and observation (Figure 1C).”

      (13) Given the repetitiveness of the figures, an icon atop each figure specifying what's being tested will be helpful.

      We have added an icon for each figure 3-8.

      (14) I feel the supplemental figures for Figure 9 are more compelling than the main figure. Consider including some of the panels in the main figure.

      Thank you for the suggestion. We included heat maps of the adult family relationships in Figure 9, including the cosine similarities and Euclidean distances. The current legend of Figure 9 is changed as follows: “Figures 9. Across-model visualization of the face similarity between marmoset pairs. Four types of family relationships (mother-father, father-son, mother-son, and twin1-twin2) were compared, based on the training results of adult and young marmosets. The similarity was calculated using (A) cosine similarity and the (B) Euclidean distance. Heat map of the (C) cosine similarity and (D) Euclidean distance was plotted between the 3 relationship pairs in the adult family. The cosine similarity score ranged from 0 (very different) and 1 (exactly same) for marmoset faces. The Euclidean distance score ranged from 0 (exactly same) and 1 (very different) for marmoset faces.”

      (15) Related to this, I feel that the display of Figure 9 obscures the differences that the authors report.

      The display of Figure 9 has been improved following the above reviewer’s suggestions.

      (16) In Figure 9, given the large effect sizes but non-significant p-values, will adding more training points/epochs improve the differences.

      For the trained recognition programs, we have already implemented the “early stopping criteria” as shown in the Results section: “Each of the models was trained until reaching the early stopping criteria (i.e. no improvement within the last 100 training epochs).” This means that the program has already plateau with its performance.

      (17) Line 30: either "within" or "in".

      This has been corrected.

    1. eLife Assessment

      This important study investigates how distinct honeybee viruses differentially alter flight performance through interactions with octopamine signaling pathways. The combination of behavioral flight assays, pharmacological perturbation, and transcriptomic analyses provides solid evidence that virus-specific effects on flight are associated with octopamine signaling. The data presented also establish a framework for additional analyses that may strengthen the proposed mechanistic model, including quantification of endogenous octopamine levels and receptor specificity studies.

    2. Reviewer #1 (Public review):

      Summary:

      Kaku and Flenniken investigate the mechanistic pathways through which specific viral infections alter the flight capabilities of honeybees. Building on their previous discovery that DWV impairs flight while SBV unexpectedly enhances it, the authors hypothesized that these behavioral shifts are driven by interactions with the insect's octopamine (OA) signaling pathway, which is responsible for the "fight-or-flight" neurohormonal stress response and energy mobilization. To test this, the authors experimentally infected adult honeybees with DWV or SBV and pharmacologically manipulated the OA pathway using either octopamine supplementation or epinastine (EP), an OA-receptor antagonist. They then evaluated the bees' flight performance (distance, duration, and speed) on custom flight mills and profiled their gene expression using qPCR and RNA sequencing.

      Strengths:

      A major strength of this study Is the high prevalence of preexisting background DWV and SBV infections in the honeybee cohorts, which meant there were no completely "virus-free" control groups. However, the authors successfully mitigated this limitation by rigorously quantifying viral RNA copies for every individual bee via qPCR and utilizing these viral abundances as continuous variables in powerful linear mixed-effect models.

      Weaknesses:

      The primary weakness lies in the methodology used for targeted pharmacological manipulations, as well as the lack of OA quantification across different treatments. Thus, their claims are not sufficiently supported by the current data.

      Comments on revised version.

      I appreciate the authors' efforts to address the reviewers' concerns and to revise the wording of the manuscript. The revised version is more cautious than the original, and some of the discussion has been appropriately toned down. However, I remain unconvinced that the key mechanistic conclusions are sufficiently supported by the current evidence.

      (1) The specificity of epinastine remains insufficiently demonstrated.<br /> The authors argue that AmOARβ2 is the predominantly expressed octopamine receptor subtype in their RNA-seq dataset and therefore the physiological effects of epinastine are most likely mediated through this receptor. However, I do not find this argument fully convincing.

      First, relatively low transcript abundance of other OA receptor subtypes does not exclude their physiological contribution. Even receptors expressed at lower levels may play important functional roles, particularly in specific neuronal populations or flight-related tissues. Therefore, the possibility that epinastine affects multiple OA receptor subtypes cannot be excluded.

      Second, although epinastine is widely used as a pharmacological tool to inhibit octopamine signaling, its receptor pharmacology has not been comprehensively characterized. The study by Roeder et al. primarily employed radioligand binding assays, which provide information on receptor affinity but not on functional antagonism or subtype selectivity. Without systematic functional characterization across the insect octopamine receptor family, it remains difficult to exclude contributions from other OA receptor subtypes or potential off-target effects.

      A more convincing pharmacological strategy would be to demonstrate similar results using an additional chemically distinct octopamine receptor antagonist. Concordant phenotypes obtained with two independent antagonists would substantially strengthen the conclusion and reduce concerns regarding off-target effects.

      (2) The OA supplementation experiments should be interpreted more cautiously.<br /> The authors correctly acknowledge that exogenous octopamine produces only transient elevations in signaling. However, I do not find the comparison with synthetic agonists entirely appropriate.

      Although synthetic agonists such as amitraz generally produce more prolonged receptor activation than endogenous octopamine, the more fundamental difference lies in their physicochemical properties. Octopamine is a highly polar endogenous amine that exhibits limited tissue penetration and is rapidly cleared through uptake and metabolic pathways. Consequently, exogenously administered OA is unlikely to efficiently reach relevant target tissues or receptor populations in a manner comparable to endogenous neurotransmitter release. In contrast, the greater lipophilicity of amitraz facilitates its distribution into target organs and enables more sustained receptor engagement following systemic administration.

      More importantly, the observation that OA supplementation partially rescues flight behavior does NOT necessarily establish that altered endogenous OA signaling is the primary mechanism underlying the virus-induced phenotypes. Such rescue experiments demonstrate that pharmacological enhancement of octopaminergic signaling can modulate the phenotype, but they do NOT provide direct evidence that endogenous OA levels or OA signaling are altered by viral infection. Therefore, these experiments should be interpreted as supportive rather than mechanistic evidence.

      (3) Direct quantification of octopamine remains the major missing evidence.<br /> The authors acknowledge that direct measurements of octopamine and tyramine would strengthen their conclusions but argue that technical limitations and cost prevented these analyses. While these practical considerations are understandable, they do not compensate for the absence of the critical mechanistic evidence.

      Overall, I appreciate the authors' revisions and agree that the manuscript provides interesting evidence that octopaminergic signaling is associated with virus-dependent changes in honeybee flight performance. However, I do not believe that the current data are sufficient to support the stronger mechanistic claims regarding regulation of the OA pathway or the specific involvement of the AmOARβ2 receptor.

      Unless direct measurements of endogenous OA (and ideally tyramine) can be provided, I recommend that the authors substantially moderate the mechanistic conclusions throughout the manuscript, including the Abstract, Results, and Discussion. The study should be presented primarily as evidence for a pharmacological association with octopaminergic signaling rather than as definitive proof of the proposed mechanistic model.

    3. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment:

      This important study investigates how distinct honey bee viruses differentially alter flight performance through interactions with octopamine signaling pathways. The combination of behavioral flight assays, pharmacological perturbation, and transcriptomic analyses provides solid evidence that virus-specific effects on flight are associated with octopamine signaling. However, some of the stronger mechanistic conclusions regarding direct regulation of octopamine signaling remain incomplete without more specific validation of receptor-level effects and direct quantification of octopamine levels or signaling activity.

      We revised some of text in the manuscript, since we agree that octopamine and tyramine quantification would strengthen the mechanistic interpretation of our findings. While we acknowledge that direct measurements of OA and tyramine would provide valuable complementary evidence, the current study relies on multiple independent lines of evidence—including gene expression analyses, OA supplementation experiments, and behavioral measurements—that collectively support a role for octopaminergic signaling in mediating the observed effects. The revised text better reflects the data included in this paper.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      Kaku and Flenniken investigate the mechanistic pathways through which specific viral infections alter the flight capabilities of honey bees. Building on their previous discovery that DWV impairs flight while SBV unexpectedly enhances it, the authors hypothesized that these behavioral shifts are driven by interactions with the insect's octopamine (OA) signaling pathway, which is responsible for the "fight-or-flight" neurohormonal stress response and energy mobilization. To test this, the authors experimentally infected adult honey bees with DWV or SBV and pharmacologically manipulated the OA pathway using either octopamine supplementation or epinastine (EP), an OA-receptor antagonist. They then evaluated the bees' flight performance (distance, duration, and speed) on custom flight mills and profiled their gene expression using qPCR and RNA sequencing.

      Strengths:

      A major strength of this study is the high prevalence of preexisting background DWV and SBV infections in the honey bee cohorts, which meant there were no completely "virus-free" control groups. However, the authors successfully mitigated this limitation by rigorously quantifying viral RNA copies for every individual bee via qPCR and utilizing these viral abundances as continuous variables in powerful linear mixed-effect models.

      Weaknesses:

      The primary weakness lies in the methodology used for targeted pharmacological manipulations, as well as the lack of OA quantification across different treatments. Thus, their claims are not sufficiently supported by the current data.

      We thank Reviewer #1 for these comments.

      (1) The authors utilize Epinastine to block octopamine signaling, describing it as a highly specific OA receptor antagonist. However, pharmacological inhibitors often lack absolute specificity. Epinastine might bind to other octopamine receptor subtypes present in honey bee neural and flight muscle tissues, or it could potentially cross-react with tyramine and dopamine receptors. Without further genetic validation (e.g., RNA interference targeting specific receptors), it is difficult to definitively conclude that the altered flight performance is solely due to the blockade of the specific Oβ−2R pathway.

      We thank the reviewer for this thoughtful comment and agree that pharmacological approaches have inherent limitations with respect to receptor specificity. However, among the available octopamine receptor antagonists, epinastine is considered one of the most selective compounds for insect octopamine receptors. Roeder et al. (1998) reported that epinastine exhibits affinities for octopamine receptors that are at least four orders of magnitude greater than those for other insect biogenic amine receptors, including dopamine, tyramine, histamine, and serotonin receptors. We updated the text to include this information.

      Honey bees encode four β-adrenergic-like receptors (AmOARβ1- AmOARβ4) and one αadrenergic-like receptor (AmOARα1). Our transcriptomic analyses indicated that expression of AmOARβ2 was substantially higher than that of other octopamine receptor genes. Specifically, AmOARβ4 transcripts were not detected in our RNA-seq datasets, while AmOARβ1 and AmOARβ3 were expressed at very low levels in most samples (Supplementary Table S9; Figure S5). Although AmOARα1 transcripts were detected in some samples, expression levels were consistently lower than those of AmOARβ2. These observations support the interpretation that the physiological effects observed following epinastine treatment are primarily mediated through disruption of AmOARβ2 signaling. We updated the text to include this information.

      We agree that receptor-specific genetic approaches would provide valuable complementary evidence. RNAi-mediated knockdown of AmOARβ2 is an attractive future direction; however, RNAi efficacy in honey bees is variable and influenced by factors including transcript turnover rates. In addition, dsRNA treatments can induce sequence-independent antiviral effects that could confound interpretation in studies involving viral infection (Flenniken and Andino PONE 2013; Brutscher, Daughenbaugh, and Flenniken Sci Reports 2017). We have revised the manuscript to more explicitly acknowledge these limitations and to clarify the basis for our interpretation of the epinastine experiments.

      (2) As a natural neurotransmitter, insects have evolved highly efficient "cleanup" mechanisms. OA is rapidly cleared from the synaptic cleft via reuptake transporters and quickly inactivated by enzymes such as N-acetyltransferase (NAT) or Monoamine Oxidase (MAO). Consequently, an injection of OA produces only a transient "pulse" of activity. It is often a poor "tool" for inducing prolonged physiological effects compared to synthetic formamidines like Amitraz.

      We thank the reviewer for this important point regarding the pharmacokinetics of octopamine. We agree that octopamine is rapidly metabolised and cleared under physiological conditions and that exogenous administration is unlikely to precisely mimic endogenous signaling dynamics. Our goal was not to induce a prolonged pharmacological activation of octopamine signaling comparable to that produced by synthetic agonists such as amitraz, but rather to determine whether increasing octopaminergic signaling could mitigate the flight impairments associated with DWV infection. Octopamine was administered either by injection or through feeding (Lines 86-89), both of which resulted in significant improvements in flight performance in DWV-infected bees (Figure 2). The observation that two independent delivery methods produced similar outcomes supports the conclusion that enhanced octopaminergic signaling can partially rescue the DWV-associated flight phenotype. We have revised the manuscript to clarify this distinction and to acknowledge that exogenous octopamine administration likely produces transient elevations in signaling rather than sustained receptor activation.

      (3) The study relies heavily on transcriptomics and quantitative PCR to measure the mRNA expression of key synthesizing enzymes, namely tyrosine decarboxylase (tdc) and tyramine βhydroxylase (tβh), to infer the activation or suppression of the octopamine pathway. However, changes in enzyme synthesis at the RNA level are often insufficient to accurately reflect the true physiological levels of biogenic amines. To robustly prove the authors' hypothesis of a "feedback loop that regulates intracellular OA concentrations", direct quantification of actual octopamine and tyramine titers in the bees (e.g., using high-performance liquid chromatography or mass spectrometry) is necessary.

      We thank the reviewer for this comment and agree that octopamine and tyramine quantification would strengthen the mechanistic interpretation of our findings. Previous studies have successfully quantified OA in honey bees using HPLC-based approaches, including KayaZee et al. (2022, eLife), who measured OA in honey bee muscle tissue (both naturally occurring levels and levels post-treatment with 10 mM OA), and Cook et al. (2017, J. Exp. Bio) who quantified OA in pooled honey bee brain samples.

      Prior to submission, we inquired with our institutional mass spectrometry facility regarding the feasibility of measuring OA in individual honey bee samples. The expected concentrations of OA in our samples was below their limit of detection, so we did not pursue these analyses at that time. During the review process, we explored the possibility of analyzing a subset of samples at external facilities that may have the sensitivity required to quantify OA and tyramine in honey bee tissues. Since such analyses would require substantial resources, with estimated costs of approximately $5,000–10,000 for 12–15 samples that have been stored in the -80C since the study, rather than flash-frozen in liquid nitrogen as described by Zee et al. 2022. While we acknowledge that direct measurements of OA and tyramine would provide valuable complementary evidence, the current study relies on multiple independent lines of evidence— including gene expression analyses, OA supplementation experiments, and behavioral measurements—that collectively support a role for octopaminergic signaling in mediating the observed effects. We thank the reviewer for this valuable suggestion. While these analyses are beyond the scope of this study, we will consider using this approach in future studies.

      Reviewer #2 (Public review):

      Summary:

      This highly original and well-designed study provides insight into how honeybee picorna-like viruses, Deformed wing virus (DWV) and Sacbrood virus (SBV), affect flight performance, and reveals the role of the octopamine (OA) pathway in virus-honeybee interactions. The authors used a flight mill to quantify the flight performance of bees with different levels of DWV and SBV. Bees were treated with OA and/or epinastine (EP) - an OA receptor antagonist; the study also quantified virus loads and expression of two key genes involved in OA biosynthesis.

      The results showed that reduced flight performance associated with high DWV levels could be alleviated by OA administration. In contrast, increased levels of SBV had the opposite effect, leading to enhanced flight performance. This suggests distinct physiological responses to DWV and SBV infections. Administration of EP had led to a reduction of flight performance in SBVinfected bees, indicating the involvement of the OA pathway.

      The authors also quantified levels of mRNAs of enzymes involved in OA synthesis, tyrosine decarboxylase (TDC) and tyramine beta-hydroxylase (TbH), and concluded that DWV induced expression of TbH, while SBV upregulated expression of TDC. Furthermore, the study identified upregulated and downregulated genes in response to SBV, DWV and DWV in combination with OA.

      Strengths:

      The study reported opposing effects of infections of related viruses, SBV and DWV, on honeybee flight performance, and identified the central role of the octopamine (OA) signaling pathway in the effect of viruses on honeybee flights.

      These findings were achieved by using a combination of approaches, including experimental measurement of flight distance, virus infections, and introduction of OA and EP. Experimental work with honeybees is technically challenging and requires specialized expertise, which makes the results produced in this study more valuable.

      DWV and SBV are among the most important honeybee pathogens affecting honeybee health and threatening the pollination service. Therefore, an understanding of the mechanisms underlying DWV and SBV pathogenesis has the potential to develop novel approaches to mitigate the negative impact of these viruses.

      Weaknesses:

      No weaknesses were identified by this reviewer.

      We thank Reviewer #2 for these comments

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      I have only minor suggestions for the manuscript.

      (1) L. 45-46

      Please note that not only high virus levels have a negative impact on honeybees. Low levels of DWV, typical of covert infections, can have long-term deleterious effects on honeybee foraging and survival. Please include citation (e.g., Benaets et al, 2017, Proc Biol Sci (2017) 284 (1848): 20162149. https://doi.org/10.1098/rspb.2016.2149).

      We thank the reviewer for these comments and edited the text accordingly, and apologize for our inadvertent omission of Benaets et al 2017, which is cited in our previous publication.

      (2) L. 113

      Clarify what is meant by "high DWV levels"

      "i.e., 10^8 copies / 2 ug RNA" -> "i.e., above 10^8 copies / 2 ug RNA"

      We thank the reviewer for this comment and corrected this in the text.

      (3) L.115

      "..mock infected bees.." /Figure 2A.

      Did these bees have low levels of DWV, below 10^8 / 2 mg RNA? What was the level of DWV in these bees?

      Mock-infected bees had an average of 3x10<sup>5</sup> DWV copies and 2x10<sup>3</sup> SBV copies per 2 µg RNA (reported in Lines 107-109 in revised manuscript, a few lines before in original manuscript).

      (4) Figure 2 / Legends to Figure 2

      Note that in Figure 2 legends, the grey areas show 95% confidence intervals for regression lines.

      We thank the reviewer for this comment and added this in the text.

      (5) Figure 2 / Legends to Figure 2

      Consider including correlation coefficients (R) and p-values for each of the regression lines in Figures 2A-F. (These could be included in the Figure 2 legends).

      We thank the reviewer for this suggestion and agree that providing sufficient statistical information is important for data interpretation. Because the analyses presented in Figure 2 are based on linear mixed-effects models that incorporate both fixed and random effects, the statistical outputs are more complex than those associated with simple linear regressions. For figure clarity, we chose not to include all model statistics within the figure panels or legends and the key statistical results, including p-values and model fit metrics (R<sup>2</sup> values), are reported in the main text (Lines 121+). In addition, complete model outputs, including all relevant coefficients, correlation estimates, and associated statistics, are provided in Supplemental Data Sheet S4. To address the Reviewer’s comments, we revised the figure caption to improve clarity and include key p-values.

      We believe this approach better balances accessibility in the main figures with comprehensive reporting of the statistical analyses and thank the Reviewer for this useful suggestion.

      (7) L.357-377 - virus-specific responses

      A previous honeybee transcriptome analysis study, which showed different responses to DWV and SBV, could be cited (Ryabov E. 2016. PeerJ 4:e1591 https://doi.org/10.7717/peerj.1591).

      We thank the reviewer for this point and included this citation in line 338 of original manuscript (line 349 in revised, tracked-changes manuscript).

      (7) L. 412

      "bees were collected 24 hours prior to eclosion" -> e.g. "bees were collected at pupal stage 24 hours prior to eclosion"?

      Specify if dark-eyed pupae were collected to make sure eclosion in 24 hr.

      We thank the reviewer for making this point, and we revised the methods and results text to improve clarity.

    1. eLife Assessment

      This manuscript reports an important study in which the authors apply smFRET imaging to probe HIV-1 Env conformational dynamics in the presence of antibodies. Previous implementations of smFRET imaging of HIV-1 Env, which focus on gp120 conformation, have yielded limited information on antibodies that target gp41. Through the cutting-edge application of smFRET imaging, the study provides convincing insights into the mechanisms of action of relevant antibodies.

    2. Reviewer #1 (Public review):

      The authors have considered a panel of antibodies that target epitopes at the gp120/gp41 interface (8ANC195 and PGT151), the fusion peptide in the gp41 domain (VRC34), and the MPER region of gp41 (DH511.2_K3 and VRC42). They also investigate 10E8.4/iMab, which is an engineered bispecific antibody that targets the MPER and the CD4 receptor. On a technical note, they have applied a double amber codon-readthrough strategy to incorporate the non-natural TCO*A amino acid, which gets labeled through click chemistry. This approach should result in less disruption of the native Env structure as compared to the peptide insertion previously used for smFRET imaging of Env. Furthermore, previous implementations of smFRET imaging of HIV-1 Env, which focus on gp120 conformation, have yielded limited information on antibodies that target gp41. Altogether, through the cutting-edge application of smFRET imaging, the study provides novel insights into the mechanisms of action of interesting and clinically relevant antibodies.

      Comments on revised version:

      The authors have nicely responded to all of my concerns. I have no further issues.

    3. Reviewer #2 (Public review):

      Summary:

      In this paper, Xu and co-workers unveil two distinct modes of neutralisation by gp41-targeted broadly neutralizing antibodies on HIV-1 Env. So far, it was unclear as to how the mechanism of neutralisation occurred for this subset of neutralising antibodies (that can target the fusion peptide or the membrane proximal external region of the gp41 subunit). Thanks to single-molecule FRET, the authors show that the majority of broadly neutralizing antibodies stabilize the closed Env conformation (named State 1 since the original work by Munro and colleagues PMID: 25298114). Interestingly, the bivalent 10E8.4/iMab stabilized in turn a CD4-bound open state of Env. The two modes of neutralization described for these antibodies show previously unknown allosteric mechanisms that stabilize closed and open Env conformation, stressing the importance of Env conformational dynamics and its efficiency during the process of fusion.

      Strengths:

      The article is well-written, and the figures fully depict the data in a convincing way. The authors have used smFRET, which is now established in the field as a good tool to assess Env dynamics.

      Comments on revised version:

      I am very happy with the comments, answers and the way the new manuscript is shaped after revision. I have no further questions or concerns.

    4. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment:

      This manuscript reports an important study in which the authors apply smFRET imaging to probe HIV-1 Env conformational dynamics in the presence of antibodies. Previous implementations of smFRET imaging of HIV-1 Env, which focus on gp120 conformation, have yielded limited information on antibodies that target gp41. Through the cutting-edge application of smFRET imaging, the study provides convincing insights into the mechanisms of action of relevant antibodies.

      We appreciate this positive assessment and thank the reviewers for their time and constructive comments. We have made the following changes in the revised manuscript to address all points raised by reviewers.

      (1) Clarify the distinction between suppression efficiency and functional cost.

      (2) Add controls: smFRET experiments in the presence of monovalent 10E8.4 and iMab individually.

      (3) All of the smFRET population contour plots have been removed, as suggested.

      (4) Repeat neutralization experiments of tagged viruses (carrying nc-AA-incorporated, amber-suppressed Env), add and compare infectivity profiles between before and after click-chemistry labeling of tagged viruses.

      (5) Add a section (Complementary views from smFRET and structural studies) to the Discussion on how these approaches complement each other.

      (6) Further clarify three prefusion conformational states identified by smFRET, the relation with previously identified States 1, 2, 3, and asymmetry, the heterogeneity of Env presentations and virion morphology, and the focus of this study.

      Please find below our point-by-point responses to the public reviews and recommendations for the authors.

      Public Reviews:

      Reviewer #1 (Public review):

      The authors have considered a panel of antibodies that target epitopes at the gp120/gp41 interface (8ANC195 and PGT151), the fusion peptide in the gp41 domain (VRC34), and the MPER region of gp41 (DH511.2_K3 and VRC42). They also investigate 10E8.4/iMab, which is an engineered bispecific antibody that targets the MPER and the CD4 receptor. On a technical note, they have applied a double amber codon-readthrough strategy to incorporate the non-natural TCO*A amino acid, which gets labeled through click chemistry. This approach should result in less disruption of the native Env structure as compared to the peptide insertion previously used for smFRET imaging of Env. Furthermore, previous implementations of smFRET imaging of HIV-1 Env, which focus on gp120 conformation, have yielded limited information on antibodies that target gp41. Altogether, through the cutting-edge application of smFRET imaging, the study provides novel insights into the mechanisms of action of interesting and clinically relevant antibodies.

      Thank you for the positive comments!

      In validating the functionality of the S401TAG/R542TAG Env, the authors performed infectivity assays and observed 20% infectivity as compared to wild-type (Figure S2A). However, the text equates this with "20% dual-amber suppression efficiency". This would benefit from some explanation. Why do the authors interpret infectivity as reporting on amber suppression efficiency, and not the functional cost of modifying Env, which is probably unavoidable? Or a combination of both? Is there data to suggest that 100% amber suppression would leave Env 100% functional? If so, this would be valuable to show. If not, the text should be clarified.

      We acknowledge this concern and have clarified the distinction between suppression efficiency and functional cost in this revised manuscript. The observed reduction in infectivity does not translate into functional loss; instead, it more reflects the efficiency of suppression (one of the critical limitations of applying genetic code expansion in mammalian cells). To support the preservation of Env functionality, we performed dose-response neutralization experiments of tag-free and 100% dual-ncAA-incorporated Env virions by two trimer-specific neutralizing antibodies, which exhibited similar dose-dependent neutralization sensitivity (Fig. 1D), providing stronger validation than infectivity assays. We also compared infectivity between labeled and unlabeled virions and observed no significant difference (Fig. S3B).

      We have previously discussed several limitations of amber suppression in mammalian cells when combined with smFRET viral systems (PMID: 38232732; PMID: 40716060) and, more recently, in our methodology chapters (PMID: 42349953; PMID: 42349954). In brief, orthogonal tRNA/aaRS pair–mediated amber suppression (reassigning/repurposing amber stop codons to non-canonical amino acids) of the introduced ambers in the target protein (Env in our case) must compete with the cellular translation system, particularly release factors that recognize amber codons and terminate translation. Readthrough of endogenous amber codons in virus-producing cells (in our case, HEK293T) can disrupt normal protein expression and virus production. Similarly, readthrough of pre-existing amber codons in HIV-1 ORFs other than the targeted ambers in Env can disrupt virus assembly, which we addressed by generating an amber-free provirus (PMID: 38232732). Introducing two amber codons into Env further reduces efficiency, as dual suppression requires two sequential successful suppression events within the same Env molecule.

      The authors state that the contour plots in Figure 2E reveal "dynamic sampling" of the observed FRET states. Strictly speaking, as presented, the contour plots (and FRET histograms) provide no information on dynamics per se. They indicate only the relative thermodynamic stabilities of the FRET states; transitions between states are a matter of interpretation. The TDPs, shown later in Figure 5A, nicely display the dynamics. More importantly, interpretation of the contour plots is challenging, as some seem to suggest an evolution toward lower FRET states. This is especially evident in Figures 2F and 3D, which suggest that the system evolves into a stable 0.1-FRET state (CO) after about 3 sec. Unless the authors want to conclude something from this, I would suggest that they consider removing the contour plots, since their interpretations are fully supported by the FRET histograms alone.

      We agree and have removed the contour plots, as they do not add meaningful information beyond what the histograms show.

      The data indicating that Env conformation is manipulated by 10E8.4/iMab is interesting. If I understand correctly, 10E8.4/iMab is an engineered antibody with one Fab targeting MPER and the second Fab targeting CD4. In the absence of CD4, could the difference between 10E8.4/iMab and the other MPER antibodies be due to 10E8.4/iMab being monovalent with respect to MPER binding?

      We appreciate this question. To address this, we have performed important controls: smFRET experiments in the presence of 10E8.4 and iMab individually in the absence of CD4. The results are shown in Fig. S9 in the revised manuscript, which indicates that 10E8.4 behaves similarly to other MPER-directed bNAbs we tested in this study, whereas iMab does not appear to affect the conformational populations of Env. The dual effect exerted by the bivalent 10E8.4/iMab is therefore very unexpected and thus interesting, as discussed in the Discussion section.

      Reviewer #2 (Public review):

      Summary:

      In this paper, Xu and co-workers unveil two distinct modes of neutralisation by gp41targeted broadly neutralizing antibodies on HIV-1 Env. So far, it was unclear as to how the mechanism of neutralisation occurred for this subset of neutralising antibodies (that can target the fusion peptide or the membrane proximal external region of the gp41 subunit). Thanks to single-molecule FRET, the authors show that the majority of broadly neutralizing antibodies stabilize the closed Env conformation (named State 1 since the original work by Munro and colleagues PMID: 25298114). Interestingly, the bivalent 10E8.4/iMab stabilized in turn a CD4-bound open state of Env. The two modes of neutralization described for these antibodies show previously unknown allosteric mechanisms that stabilize closed and open Env conformation, stressing the importance of Env conformational dynamics and its efficiency during the process of fusion.

      Strengths:

      The article is well-written, and the figures fully depict the data in a convincing way. The authors have used smFRET, which is now established in the field as a good tool to assess Env dynamics.

      We appreciate these positive comments!

      Weaknesses:

      (1) The limited controls on how click chemistry affects Env (as labelled Env HIV virions were not evaluated).

      We agree. Our previous validation focused on ncAA-incorporated Env HIV-1 virions, but not the fluorescently labeled virions. To address this, we have added infectivity results for labeled virions after click-chemistry labeling, compared with those before labeling. We did not observe any measurable difference in infectivity (Fig. S3B), indicating that the labeling procedure does not impair viral infectivity.

      We also attempted to perform dose-dependent neutralization after labeling. However, as anticipated in our provisional response, this remains technically challenging because the additional labeling and centrifugation steps substantially increase sample handling time, while the dual amber suppression system already limits virion production in cells. As a result, we were not able to obtain sufficiently robust datasets for this additional functional validation.

      Nevertheless, we have previously demonstrated real-time tracking of single click-labeled Env virions during internalization and intracellular trafficking in live cells (PMID: 38232732), providing independent evidence that click-chemistry-labeled Env retains functional competence.

      (2) Photobleaching of donor and acceptor molecules occurs right after 10sec exposure.

      We acknowledge this limitation and have included it in the revision.

      (3) Other limitations are well described in the corresponding section.

      We appreciate this comment.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      As a means of clarifying the mechanism of 10E8.4/iMab, the authors might consider performing separate smFRET experiments in the presence of the normal 10E8.4 antibody and the normal iMab antibody (a negative control). Alternatively, they could consider imaging in the presence of the DH511.2_K3 and VRC42 Fabs (as opposed to full-length Ig) to make a cleaner comparison, although this may be less informative given the high concentrations of antibodies used.

      We thank the reviewer for this excellent suggestion. To enable a direct comparison, we performed the most informative control by examining virus-associated Env in the presence of 10E8.4 alone and iMab alone. The corresponding smFRET results are presented in Fig. S9. We found that 10E8.4 behaves similarly to other MPER-directed antibodies, whereas iMab alone does not appear to have a notable effect on the conformational propensity of Env. For transparency and to facilitate future antibody design, we have also included the Fab region sequences of the antibodies in Table S2.

      Reviewer #2 (Recommendations for the authors):

      The article is well-written, the findings are of high interest for the community. The article should be shared once the points stated below are clarified and revised by the authors.

      We appreciate this comment and the points raised by the reviewer and have revised the manuscript accordingly.

      (1) In Figure 1C, the tomographic slices showing HIV-1 WT as compared to HIV-1 decorated with EnvBG505 S401ncAA R542nCAA are quite different morphologically. The micrographs chosen show a big particle with two capsids close to a smaller one without a capsid and at least in this plane bold (no Env incorporation) for the WT; whilst for the HIV-1 decorated with EnvBG505 S401ncAA R542nCAA no capsid is apparent in both particles, one (the right one) is very small and the right one does present a number of Envs but no apparent capsid is visible here. Please comment - perhaps it would be important to average the morphological traits of both and look at average diameter, average Env incorporation, morphology of the capsid, percentage of immature particles, capsid abnormalities (as the one shown in the upper micrograph).

      We thank the reviewer for this thoughtful comment.

      The tomograms in Fig. 1C were included to demonstrate the overall size and shape of the viral particles rather than to provide a quantitative structural comparison. HIV-1 viral particles are inherently heterogeneous, and the original slides were selected as representative examples. Following the reviewer's suggestion, we replaced the representative wild-type (Fig. 1C, top panel) and tagged virus (Fig. 1C, bottom panel) tomographic slides with those that better reflect the overall quality of each sample. To further address this concern, we refer the reviewer to the nanoparticle tracking analysis (NTA) shown in Fig. S3, which shows no significant difference in particle diameter between the wild-type and tagged viruses. In the revised manuscript, we now replace "morphology" with "shape" or "size," as these terms better reflect what our results can say.

      We agree that a quantitative analysis of capsid morphology, Env spike incorporation, and the proportion of immature particles would be informative. However, such analyses would require a substantially larger cryoET dataset, which is beyond the scope of the present study; nevertheless, it is certainly in our interest to pursue a cryoET-focused study of EnvCA interactions, with Env complexed with 10E8.4/iMab. Our primary objective is to study Env conformational dynamics by smFRET rather than viral morphogenesis or capsid maturation, whose relationship to Env dynamics remains largely unexplored. It is also worth noting that the optimal particle populations for smFRET and cryoET differ. smFRET measures the conformational dynamics of individual Env trimers and therefore selectively analyzes virions containing a single dually labeled Env trimer, whereas cryoET structural analyses typically benefit from particles with higher Env spike densities. Therefore, the particle populations favored for the two techniques are not identical.

      (2) In Figure 1D, there is a difference in neutralisation with PGT151 - how different are these two curves - how does the labelling affect neutralisation for bNAbs targeting gp41? Would it be possible to assess also the infectivity, fusion and neutralisation profiles of particles where the flurophores are included? This would be without diluting the Env for single particle analysis, but just to understand how harsh the organic reaction is and how it affects Env function (as all experiments and conclusions in the manuscript are based on labelled Env).

      Again, we sincerely appreciate these questions, which have helped us improve the manuscript. Neutralization assays for the tagged viruses were performed using the ncAA-incorporated, amber-suppressed viruses, whereas the engineered wild-type is amber-free. The differences between these two dose-response curves in the original Fig. 1D are small and within the experimental variation routinely observed under even identical conditions (same virus and same bNAb). We have repeated these experiments, and the new results are shown in the revised Fig. 1D. Although minor variations remain, the overall neutralization profiles and IC50 values are highly consistent.

      To assess whether the fluorophore labeling reaction affects Env functionality, as noted above, we have included infectivity results for labeled virions after click-chemistry labeling, compared with those before labeling. We did not observe any measurable difference in infectivity (Fig. S3B), indicating that the labeling procedure does not impair viral infectivity. We also attempted to perform dose-dependent neutralization after labeling. However, as anticipated in our provisional response, this remains technically challenging because the additional labeling and centrifugation steps substantially increase sample handling time, while the dual amber suppression system already limits virion production in cells. As a result, we were not able to obtain sufficiently robust datasets for this additional functional validation. Nevertheless, we have previously demonstrated real-time tracking of click-labelled Env virions during internalization and intracellular trafficking in live cells (PMID: 38232732), providing independent evidence that click-chemistry-labelled Env retains functional competence.

      We believe that the unchanged infectivity of labeled viruses relative to their unlabeled counterparts, together with our previously observed real-time trajectories of click-labeled virions in live cells, provides strong evidence that our labeling strategy does not measurably impair Env function.

      (3) In Figure 2E and 2G, the authors employ a three Gaussian fit approach to recover the three populations (pre-triggered - pre-fusion closed - CD4 bound open). Can you please relate these with State 1, 2 and 3 from the original article (PMID: 25298114). Comment on the possibility that more than three populations could be fitted and what this could mean - pre-triggered and partially open (one gp120 asymmetrically open) could occur? Could this labelling approach account for this asymmetry?

      Thanks for this suggestion. In this study, we compared our results obtained using the gp120-gp41 structural axis with those obtained using the referenced gp120 V1-V4 structural axis to confidently assign the FRET-identified states to the previously reported three primary populations. The referenced axis is comparable to those used in the original article (PMID: 25298114) and later confirmed using the amber-click strategy (PMID: 38232732). We observe the same structural changes from these two distinct structural angles, as probed under ligand-free conditions (Fig. 2E and 2G) and CD4-triggered open conditions (Fig. 2F and 2H).

      The pre-triggered state corresponds to State 1; the pre-fusion closed state corresponds to the symmetric State 2 (which the SOSIP-based soluble Env primarily adopts; PMID: 30971821); and the CD4-bound open state corresponds to the fully open State 3. The assignment of the FRET states observed from the gp120 V1-V4 structural axis to States 1, 2, and 3 was originally reported in two studies (PMIDs: 27795397 and 29561264). In the asymmetric trimer configuration, the State 2 FRET signal originates from the free protomer, while the other one or two protomers bind CD4 and adopt the open conformation (PMID: 29561264). The asymmetric intermediate (PMID: 29561264) was identified using a heterotrimer experimental design consisting of a mixture of wild-type and CD4-binding-incompetent D368R protomers, which was not used in the present study. Therefore, our labeling approach cannot unambiguously resolve this asymmetry.

      Regarding the possibility of more than three populations, evidence from current and previous studies (PMIDs: 25298114, 27795397, 29561264, 38232732, 30971821) strongly supports the presence of three primary states of virus-associated Env, with additional substates that can be resolved under specific triggering conditions (PMIDs: 30974085, 41326374, 39640534). The assignment of such substates requires well-controlled experimental designs (PMIDs: 30974085, 41326374, 39640534).

      We have related PT, PC, and CO to States 1, 2, and 3, and added comments on multiple states and asymmetry in the revised manuscript.

      (4) When comparing smFRET with CryoET or structure, one can see that in smFRET there are always many potential conformations for big sub-populations of Env. Indeed, there is a trend, and the addition of bNAbs (Figure 4) clearly has an impact on increasing and stabilizing a particular state as defined by the authors (e.g. PT at 45% upon addition of 8ANC195, but also 32% PC and 23% CO). I assume that when analysing single particle CryoET or single virus CryoET, one needs to discard after template matching different scenarios that do not necessarily contribute to the highest resolution and this information is not always discussed. It would be interesting to address this in the discussion as the effect on Env dynamics of adding ligands (including CD4 and 17b) is not inducing in all Envs a drastic conformational change - this could be derived from the Ka of the ligands, but also from the intrinsic Env heterogeneity in both dynamics and architecture - I think that addressing these matters in the discussion could be of interest for the community. In this regard, the transition density plots are very helpful.

      We completely agree and appreciate this insightful suggestion. We have expanded the Discussion to better address the complementary insights provided by smFRET and structural approaches. In single-particle cryoEM, we do not observe the full spectrum of Env conformations for technical reasons, not because particles are intentionally discarded to obtain only the highest-resolution structures. One reason is that open Env conformations are much more sensitive to radiation damage than closed Env. Likewise, ligand-free closed Env is more sensitive to radiation damage than a bNAb-stabilized closed Env. Thus, the outcome of an SPA cryo-EM study depends strongly on the biological question being addressed and the conformational state that is preferentially preserved under the experimental conditions. Although one could hypothetically collect much larger datasets to recover lower-abundance conformations, this would be both cost-prohibitive and unlikely to faithfully represent the relative conformational populations due to differential, conformation-dependent radiation damage.

      We agree that the smFRET data highlight an important aspect of Env dynamics. Ligands, including bNAbs, CD4, and 17b, generally shift the conformational equilibrium toward particular states rather than driving all Env trimers into a single conformation. This likely reflects both differences in ligand binding properties and the intrinsic conformational heterogeneity of Env. We therefore believe that structural studies and smFRET provide complementary information. Structural methods resolve the molecular architecture of individual conformational states at atomic (by cryoEM) and near-atomic (by cryo-ET) levels, whereas smFRET quantifies their relative populations and dynamic interconversion. As the reviewer pointed out, the transition density plots are particularly valuable in illustrating these dynamic changes.

      We have incorporated these points into the revised Discussion (Subtitle: complementary views from smFRET and structural studies in general).

      (5) One of the very interesting findings of the paper is that the effect of bNabs (at least the ones tested) have an impact on the Env dynamics and how this shift can alter entry - therefore the structural view is perhaps less important - In spite of this, we still employ a structural jargon to refer to "Env open conformation stabilisation" for instance - even if the data shows that upon ligand exposure dynamics are still important but shifted. Please comment.

      This is a great point. Our smFRET data show that bNAb binding generally shifts the conformational distribution toward and stabilizes particular Env states by lowering their free energy and increasing their occupancy. We have clarified that ligand-induced stabilization reflects a redistribution of the conformational ensemble while preserving the intrinsic dynamic nature of Env.

    1. eLife Assessment

      This important study investigates the role of Nav1.7 voltage-gated sodium channels in regulating excitability of human dorsal root ganglion (hDRG) neurons. The authors characterize a previously identified Nav1.7 channel inhibitor using recombinant channels and human neurons. The study convincingly shows that inhibition of Nav1.7 channels with AM-2099 causes a modest decrease in neuronal excitability and prolongs the refractory period. This work offers new insights into the mechanism of clinically relevant pharmacological targets for pain relief.

    2. Reviewer #1 (Public review):

      Summary:

      Fujita and colleagues investigated two selective peripheral nerve voltage-gated sodium channel inhibitors targeting either Nav1.7 or Nav1.8 on excitability of human dorsal root ganglion neurons. The authors discovered that Nav1.8 inhibition is more effective at suppressing repetitive firing of DRG neurons and this may explain the greater clinical efficacy observed for suzetrigine.

      Strengths:

      The study is interesting and the findings are conceptually satisfying in that they may explain one aspect of Nav1.7 vs Nav1.8 targeting success.

      Weaknesses:

      (1) The use of postmortem human DRG neurons provides translational relevance, but the use of these cells is also a liability given their high degree of variability. Of note are the 10 to 20-fold differences in baseline properties among cells, which dwarfs the effects of the test compounds. The experiments may suffer from under sampling.

      Comments on revised version.

      The revised manuscript addresses my prior concern with reasonable effort given the limitations of human postmortem DRGs.

    3. Reviewer #3 (Public review):

      Summary:

      In this manuscript, Fujita/Jo/Stewart/Osorno et al., investigate the contribution of Nav1.7 in regulating the excitability and firing properties of human dorsal root ganglion (hDRG) neurons in vitro. The authors characterize the effects of a previously reported Nav1.7-selective blocker AM-2099 in recombinant human Nav1.7 channels and in cultured hDRG neurons from postmortem organ donors. The authors observed modest changes in many of the properties expected by inhibiting Nav channels, including decreased action potential upstroke rate and amplitude, while increasing the voltage and current thresholds for spike generation. However, AM-2099 did not change the maximum number of APs in response to suprathreshold stimulation, leading the authors to conclude that Nav1.7 inhibition alone has limited efficacy in reducing the firing properties of hDRG neurons at the soma, and discuss that the effects of Nav inhibition may be different at distal axons.

      Strengths:

      Experiments are well-designed and executed, and the results presented are convincing. The focus on voltage-gated sodium channels in native human DRG neurons is highly relevant to recent efforts to develop safer analgesic options for chronic pain in people.

      Comments on revised version.

      The authors have done an excellent job addressing my prior critiques.

    4. Author response:

      The following is the authors’ response to the previous reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      Fujita and colleagues investigated two selective peripheral nerve voltage-gated sodium channel inhibitors targeting either Nav1.7 or Nav1.8 on the excitability of human dorsal root ganglion neurons. The authors discovered that Nav1.8 inhibition is more effective at suppressing repetitive firing of DRG neurons, and this may explain the greater clinical efficacy observed for suzetrigine.

      Strengths:

      The study is interesting, and the findings are conceptually satisfying in that they may explain one aspect of Nav1.7 vs Nav1.8 targeting success.

      Weaknesses:

      (1) The use of postmortem human DRG neurons provides translational relevance, but the use of these cells is also a liability, given their high degree of variability. Of note are the 10 to 20-fold differences in baseline properties among cells, which dwarf the effects of the test compounds. The experiments may suffer from undersampling.

      We have added data from an additional 3 donors for the key results on increase in threshold and reduction of action potential upstroke, more than doubling the number of neurons for this data. We have also added a Supplementary Figure (Figure S2) that breaks out the effects on these parameters for each donor. This illustrates that there is a high degree of neuron-to-neuron variability in the effect of inhibiting Nav1.7 channels even within a single donor, even though we confined data to neurons that were verified to be capsaicin-sensitive. We also now note that there is a similar high degree of cell-to-cell variability in relative functional expression of Nav1.7 and Nav1.8 channels in capsaicin-sensitive mouse DRG neurons.

      (2) A potential confounder when using post-mortem human DRG neurons is heterogeneity of cell types. The methods clearly state that the cells selected for recording were of 'generally' small size, but specific criteria for what constitutes 'small' or other unstated selection criteria were not provided. A table of individual cell capacitance and input resistance values, along with information about individual donors (age, sex, ethnicity), is important to include. Additionally, some discussion of how DRG neuron heterogeneity impacts the findings. This relates to concern #1 about sample size determination and how cell heterogeneity factored into this calculation.

      We have added a figure (Figure S1) showing histograms and box plots of individual cell capacitance, input resistance, resting potentials, and maximum upstroke. We have also added a table with the information about donors (Table S1). As noted, we have also added a figure (Figure S2) that breaks out the effects on these parameters for each donor, illustrating that there is a high degree of neuron-to-neuron variability in the effects even within a single donor. We have added several sentences to the Discussion concerning the neuron-to-neuron variability in the effects of Nav1.7 inhibition, including the possibility that this may reflect heterogeneity of cell function.

      Reviewer #2 (Public review):

      Summary:

      The authors examine the functional role of Nav1.7 voltage-gated sodium channels in human sensory neuron electrogenesis using a Nav1.7 selective inhibitor and human dorsal root ganglion neurons obtained from organ donors. Patch-clamp electrophysiology is used at physiological temperature to measure the impact of Nav1.7 inhibition on sensory neurons' action potential firing. This is an important topic as Nav1.7 and Nav1.8 have been identified as therapeutic targets for the treatment of pain, but there has been mixed success with isoform-specific inhibitors in clinical trials. The data suggest that Nav1.7 and Nav1.8 have overlapping yet complementary functions in nociceptor neurons and that targeting both may be most effective for reducing nociception.

      Strengths:

      The data are of high quality. Action potential properties are measured at 37 degrees Celsius. Threshold is measured using brief pulses. The Nav1.7 inhibitor has been reported to be highly selective for Nav1.7 over Nav1.8 and moderately selective for Nav1.7 over Nav1.1 and Nav1.6. Data are collected using identical conditions and protocols to a previous study on the role of Nav1.8 in similar neurons.

      Weaknesses:

      The study relies on a single Nav1.7 inhibitor that has not been extensively characterized. One prior study indicates that the IC50 is around 140 nM, thus the 600 nM concentration used in this study could be predicted to reduce Nav1.7 currents by 80%. However, there is no voltage-clamp data in the current study to confirm this, and therefore, it is unclear if the batch of AM-2099 is as potent as reported in the paper that initially described its selectivity. The impact of Nav1.7 inhibition is compared to data from a previous study by this lab, and this is a minor concern. It would have been interesting to see if the combined inhibition of Nav1.7 and Nav1.8 completely blocked action potential generation in the human DRG neurons.

      We have done experiments to directly characterize the potency of the AM-2099 sample we used on both cloned human Nav1.7 channels and on native currents in the DRG neurons. Using a stable cell line expressing human Nav1.7 channels, we determined dose-response curves at both 22°C and 37°C, using an automated patch clamp instrument. These results are shown in a new Figure 1. Interestingly, we found that the IC<sub>50</sub> is substantially higher at 37°C than at room temperature. We also did experiments quantifying the effect of 100 nM and 600 nM AM-2099 on native sodium currents in the human DRG neurons, which align well with the results on the cloned Nav1.7 channels in suggesting that at 37°C, 600 nM AM-2099 inhibits Nav1.7 channels by about 85%.

      We have also added a new figure (Figure S3) showing the effects of a different Nav1.7 inhibitor, PF04856264. The effects of this inhibitor were qualitatively identical but quantitatively smaller than those of AM-2099. When we realized this, we did voltage clamp experiments on cloned Nav1.7 channels and discovered that the potency of PF-04856264 at 37°C was weaker than expected from the published IC<sub>50</sub>, which was determined at room temperature.

      We are currently doing experiments testing combined inhibition of Nav1.7 and Nav1.8 channels whenever we can obtain human neurons. Because it is of interest to examine effects of partial as well as full inhibition of each channel type, there are multiple permutations of inhibitors combined and alone that are of interest to characterize, and these studies are still in progress. We think the results in the present manuscript stand on their own and together with previous data on effects of Nav1.8 inhibitors alone provide a foundation for on-going and future studies on combinations of inhibitors by ourselves and others.

      Reviewer #3 (Public review):

      Summary:

      In this manuscript, Fujita/Jo/Stewart/Osorno et al. investigate the contribution of Nav1.7 in regulating the excitability and firing properties of human dorsal root ganglion (hDRG) neurons in vitro. The authors characterize the effects of a previously reported Nav1.7-selective blocker AM-2099 in cultured hDRG neurons from postmortem organ donors. The authors observed modest changes in many of the properties expected by inhibiting Nav channels, including decreased action potential upstroke rate and amplitude, while increasing the voltage and current thresholds for spike generation. However, AM-2099 did not change the maximum number of APs in response to suprathreshold stimulation, leading the authors to conclude that Nav1.7 inhibition alone has limited efficacy in reducing the firing properties of hDRG neurons and that Nav1.7 blockers may have limited efficacy as analgesics. This is surprising, given that patients with loss-of-function mutations in Nav1.7 suffer from congenital insensitivity to pain. While it may indeed be true that pharmacological inhibition of Nav1.7 is unlikely to produce analgesia, the present study was limited to a single concentration of AM-2099. The manuscript would be significantly strengthened by a more careful and thorough pharmacological characterization of this compound, which has not been widely used or validated in native human DRG neurons.

      Strengths:

      Experiments are well-designed and executed, and the results presented are convincing. The focus on voltage-gated sodium channels in native human DRG neurons is highly relevant to recent efforts to develop safer analgesic options for chronic pain in people.

      Weaknesses:

      Only a single concentration of AM-2099 was used for all experiments. This compound was reported to be selective for cloned human Nav1.7 channels in heterologous systems, but has not been validated in other studies after the original publication in 2016. Since the original study reported a substantial statedependent block of recombinant Nav1.7 channels, more detailed pharmacological characterization of AM-2099 is needed in human DRG neurons to fully support these claims. This study would be significantly strengthened by the inclusion of dose-response curves to assess how much of the sodium current is inhibited at this concentration, confirming selectivity in hDRG, and whether maximal inhibition of Nav1.7 still has limited efficacy in reducing the firing of native human sensory neurons.

      We have added results from experiments to directly quantify the potency of AM-2099 on both cloned human Nav1.7 channels (new Figure 1) and on native currents in the DRG neurons (new Figure 2). These show that 600 nM AM-2099 produces about 85% inhibition of Nav1.7 channels at 37°C. We have added a paragraph to the Discussion explaining that we chose this concentration of AM-2099 to produce reasonably complete inhibition of Nav1.7 channels while minimizing potential inhibition of a component of non-Nav1.7 TTX-sensitive current.

      With regard to the broader point about reconciling the variable and sometimes relatively modest effects of Nav1.7 inhibition with the complete loss of pain sensation in humans with loss-of-function mutations, we have modified the Introduction and Discussion to eliminate any implication that the results in the manuscript suggest that pharmacological inhibition of Nav1.7 is unlikely to produce analgesia. Our experiments are only on action potential firing in the cell body, and it is perfectly possible that inhibiting Nav1.7 channels in the axon could disrupt generation or propagation of action potentials, either in the main axon or in the fine axon terminals in the spinal cord. We have modified the Discussion to explicitly point this out, which would reconcile the loss of pain sensation in humans with loss-offunction mutations with the incomplete effects of Nav1.7 inhibitors on excitability of cell bodies.

      Recommendations for the authors:

      Reviewing Editor comments:

      In addition to the points noted in the eLife assessment summary above, the study has several important strengths, including use of human primary neurons and recordings performed under physiologically relevant conditions (at 37 {degree sign}C using brief current injections). However, reviewers also identified several key weaknesses that must be addressed to support the central conclusions. In particular, multiple reviewers raised concerns regarding the lack of voltage-clamp data evaluating the efficacy and specificity of AM-2099 inhibition of Nav1.7 currents. A single dose of 600 nM was used based on the report of Marx (2016) in recombinant systems (Marx, 2016). Since no other studies other than the single Amgen report exist on this compound, it is important to validate its effects directly in the human DRGs used here. Additional concerns include the lack of dose-response analysis, as well as the large variability in baseline properties, which complicates the interpretation of the results. To assist the revision of the study, we outline below the key issues that should be addressed.

      Recommendations for authors:

      (1) Add voltage-clamp experiments to directly measure Nav1.7 current inhibition by AM-2099 in hDRG neurons. Given the limited previous characterization of this compound, it is important to confirm that the concentration used here (600 nM) effectively blocks Nav1.7 currents in the native system used here.

      (2) Related to point 1 above, perform a dose-response of AM-2099 on hDRGs on Nav1.7 currents in human DRGs. Since this study, at least in part, is framed as a comparative analysis of Nav1.7 vs Nav1.8 channel subtypes in DRGs, it seems important to establish pharmacological equivalence to ensure that the comparisons are made at functionally comparable levels of channel block.

      We have added results from experiments to directly characterize the potency of AM-2099 on both cloned human Nav1.7 channels and on native currents in the DRG neurons. Using a stable cell line expressing human Nav1.7 channels, we determined dose-response curves at both 22°C and 37°C, using an automated patch clamp instrument. These results are shown in a new Figure 1. Interestingly, we found that the IC<sup>50</sup> is substantially higher at 37°C than at room temperature. We also did experiments quantifying the effect of 100 nM and 600 nM AM-2099 on native sodium currents in the human DRG neurons, which align well with the results on the cloned Nav1.7 channels in suggesting that at 37°C, 600 nM AM-2099 inhibits Nav1.7 channels by about 85%.

      (3) Reviewer 1 notes that there seem to be 10-20-fold differences in baseline firing properties, which would exceed the effects of the test compound. This raises concerns about undersampling. Additional analysis or experiments would strengthen the conclusions.

      We have added data from an additional 3 donors for the key results on increase in threshold and reduction of action potential upstroke, more than doubling the number of neurons for this data. We have also added a Supplementary Figure that breaks out the effects on these parameters for each donor. This illustrates that there is a high degree of cell-to-cell variability in the effects even within a single donor, even though we confined data to neurons that were verified to be capsaicin-sensitive. Reviewer 1 made the excellent suggestion that because of the neuron-to-neuron variability in baseline properties, the effects of compounds could be better illustrated by displaying changes from baseline. Following this suggestion, we have added Tukey-style box plots displaying the data in this way. Together with the donor-to-donor breakout of data in the new Figure S2, these plots make it clear that the neuron-to-neuron variability reveals genuine differences in the channel make-up of each neuron and not experimental error.

      (4) Reviewer 2 notes an interesting experiment: does a combined block of Nav1.7 with the AM compound and Nav1.8 block action potential generation? If Nav1.7 controls threshold and Nav1.8 controls firing, then the combined inhibition should be highly effective in blocking nociceptive output, which could have therapeutic relevance.

      We are currently doing experiments testing combined inhibition of Nav1.7 and Nav1.8 channels whenever we can obtain human neurons. Because it is of interest to examine effects of partial as well as full inhibition of each channel type, there are multiple permutations of inhibitors combined and alone that are of interest to characterize, and these studies are still in progress. We think the results in the present manuscript stand on their own and together with previous data on effects of Nav1.8 inhibitors alone provide a foundation for ongoing and future studies on combinations of inhibitors by ourselves and others.

      Reviewer #1 (Recommendations for the authors):

      Concerns in addition to those in the Public Review:

      Major:

      (1) As per point 1 of the weaknesses in the Public Review, I'm concerned that the experiments suffer from undersampling. This requires a discussion of how the sample size was determined.

      We have added experiments from an additional 3 donors to the key results in Figures 3-5, more than doubling the number of neurons for these measurements.

      (3) The effect of compounds could be better displayed as a change from baseline in Figure 1C-E. Also, are the AP traces and phase plots shown in Figures 1AB and 2AB averages or representative?

      Thanks for this excellent suggestion. We have added box-plots that show changes from baseline for the various parameters. We have also clarified that the action potential traces and phase plots are from application of AM-2099 in a single representative neuron.

      Minor:

      (1) Provide source of VX-548 and report the purity of both compounds.

      We have provided the information for VX-548 and added the information on the purity of both compounds

      (2) Clinical failures of Nav1.7 blockers may not be solely due to pharmacodynamic limitations as implied by this study. Pharmacokinetic differences and toxicity (e.g., effects on the autonomic nervous system) may also have contributed.

      Thanks for raising this important point. We have added this point to the Introduction.

      Reviewer #2 (Recommendations for the authors):

      It is an interesting study, and the conclusions are reasonable. However, it would have been good to see validation of the potency of AM-2099 on native DRG sodium currents and/or recombinant human Nav1.7 channels expressed in a heterologous expression system.

      We have added results from experiments to directly quantify the potency of AM-2099 on both cloned human Nav1.7 channels (new Figure 1) and on native currents in the DRG neurons (new Figure 2).

      Minor comments:

      (1) Page 3, middle paragraph - there is a "(" missing before Renganathan.

      Thanks, corrected.

      (2) Page 4: Is anything known about AM-2099 in terms of state-dependence? It seems like Marx 2016 is the only previously published study using it, so additional information on the inhibitor would be helpful.

      We have not characterized the state-dependence of AM-2099, but we characterized its potency in voltage clamp using holding voltages similar to the average resting potentials of the cells in current clamp conditions.

      (3) Page 6 discusses that there might be differences between human and rodent DRG neurons in terms of Nav1.7 and Nav1.8. It would be nice if this were directly tested with these same Nav1.7 and Nav1.8 inhibitors.

      We have recently done such a study on mouse DRG neurons which has just been published (J Physiol. 604:6104-6127, doi: 10.1113/JP290574).

      (4) Figure 2A, right panel: I could not figure out the difference between the red and green traces. Perhaps this could be explained in the figure legend?

      Thank you for pointing out that this was confusing. These two traces showed two different subthreshold responses, one of which was slightly regenerative without generating a full-blown spike. We have simplified the figure by now showing only a single subthreshold response.

      Reviewer #3 (Recommendations for the authors):

      (1) The conclusion that Nav1.7 inhibition has limited efficacy for inhibiting the firing of human DRG neurons is not fully supported by the data. This may be true, but it cannot be concluded without a more thorough pharmacological characterization of this compound. Dose-response curves and experimental confirmation of Nav1.7 selectivity (maybe just total Nav current, TTX-sensitive and TTX-resistant components) are needed.

      We agree and have now added two new figures with this data.

      (2) How was the 600 nM concentration chosen? Given that AM-2099 was reported to exhibit state dependent block, how much of the Na current is inhibited by this concentration at the initial voltage used in current clamp experiments (~-80 mV)?

      We have added a paragraph to the Discussion recognizing the limitation that 600 nM AM-2099 produces ~85% rather than complete inhibition of Nav1.7 current and explaining that we chose this concentration of AM-2099 to produce reasonably complete inhibition of Nav1.7 channels while minimizing potential inhibition of a component of non-Nav1.7 TTX-sensitive current.

      (3) It appears that the effects of AM-2099 on the refractory period are bimodally distributed, where neurons that recovered more slowly at baseline were preferentially affected by AM-2099 (Figure 4). Do these reflect different neuronal populations (e.g. smaller or larger diameter DRG) or different resting voltages in these experiments?

      We agree that there seem to be two groups based on initial refractory period. Examining the parameters for the cells, there is no clear correlation between the effects of AM-2099 on the refractory period with resting potential or cell diameter. At this time, it is not obvious what determines the differences in refractory period. We speculate that neuron-to-neuron differences in the potassium conductances that generate the after hyperpolarization may be different in these cells but it will take further work to explore this.

      (4) How much of the sodium current is mediated by Nav1.7 in hDRG neurons? How does inhibition of both Nav1.7 and Nav1.8 affect hDRG excitability?

      The new Figure 2 shows data quantifying the AM-2099-sensitive current in the DRG neurons. With regard to combined Nav1.7 and Nav1.8 inhibition, we are currently doing experiments examining inhibition of excitability by combined Nav1.7 and Nav1.8 inhibition, which we agree is the logical next step in exploring how the two components of current control excitability. These are still in progress. Because designing and interpreting these experiments is facilitated by the current experiments with Nav1.7 inhibition alone, we believe that reporting the current results now will serve the scientific community better than waiting to obtain and interpret a body of data on dual inhibition in a sufficient number of donors, which we obtain only sporadically.

      (5) Donor information and soma diameters should be included. Capsaicin sensitivity testing was mentioned in the methods, but I was unable to find any inclusion of these data in the results. These may be useful to potentially infer effects in different cell types.

      We have added a figure (Figure S1) showing histograms and box plots of individual cell capacitance, input resistance, resting potentials, and maximum upstroke. We have also added a table with the information about donors (Table S1). We have now clarified that data were confined to cells verified to be capsaicin-sensitive and that ~95% of all cells tested were capsaicin-sensitive.

      (6) Please check statistical tests and reporting. Several graphs do not appear to have paired responses (e.g. Figure 1E, Figure 4B). As a result, two-tailed Wilcoxon tests would not be appropriate. Also, check reported p-values (e.g. p=.0002), which are identical for multiple panels in the Results section.

      We have clarified that the symbols of action potential width in control without a corresponding value after AM-2099 represent neurons in which the action potential in AM-2099 had a peak < 0 mV. These cells were not included in the data set of paired parameters used for the Wilcoxon test. We have also checked and verified all statistical tests.

    1. eLife Assessment

      A computational model potentially provides important new insights into the circuit mechanisms underlying navigational control in insects. The authors compare high speed video recordings of ants with detailed predictions from a new computational model. The similarities between model and behavioral data are striking and convincingly suggest how complex behavioral motifs can emerge from a simple neural circuit.

    2. Reviewer #2 (Public review):

      The paper by Freas and Wystrach is an interesting computational study, exploring the detailed mechanisms of how simple neural circuits could explain complex behavioral patterns observed in navigating ants. The authors compare detailed, high speed video recordings of Australian desert ants (Melophorus bagoti) with predictions made by their new computational model and find convincing similarities between the model and the behavioral data, at a level of detail not previously studied. Particularly interesting are emerging properties of the model, yielding behavioral motifs it was not designed to reproduce, but which occur in natural ant behavior.

      A strength of the study is that the model is based on previous models, without making major novel assumptions. It combines existing models of the insect central complex with a model of the lateral accessory lobe and adds a stochastic inhibition of forward velocity to the interaction of central complex and lateral accessory lobes. In essence, the central complex provides corrective steering signals when the goal direction and the current heading of the insect are not aligned, while the lateral accessory lobes provide an intrinsic oscillator underlying the behavioral oscillations shown by walking ants at all times. These background oscillations are modulated by the steering signals from the central complex. Depending on which phase of the intrinsic oscillations coincides with the corrective signals, and how fast the ant is moving forward during this time, a complex set of behaviors emerges.

      Most prominently, scanning behaviors, which are regularly carried out by the ants, are recapitulated in great detail by the model. Additionally, other behaviors, such as full loops, emerge naturally from the model. While computational models are not to be seen as definite evidence for any biological reality, they can provide strong support for particular neural implementations. The current study is an excellent example in that it provides evidence for a serial arrangement of central complex circuits upstream of the lateral accessory lobe circuits, modulated by speed regulating input. While the latter is hypothetical, it yields a clear hypothesis that can be validated by connectomics studies and functional work in the future.

      The computational model is explained in detail and information about all model parameters is provided in an accessible way. The approach is thus transparent and reproducible, leaving it to the readers to assess the assumptions made in the model and how the studied complex behaviors emerge. This also provides the possibility to combine this new model with existing models to expand the scope and to more comprehensively capture the behavioral repertoire of ants, and insects in general.

      Importantly, the study shows that even complex behavioral motifs do not require dedicated neural modules, but can rather emerge from the interplay of already known circuits - highlighting the efficiency of insect brains and possibly providing the path towards embodied hardware solutions of such circuits in autonomous agents.

    3. Author response:

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #1 (Public review):

      Freas and Wystrach present a computational and experimental study of ant navigation. The main innovation of the computational model is the insertion of an oscillatory element between the steering signal and the motor control that results in a trajectory whose heading oscillates around a goal direction. Additionally, the model imposes periodic cessations of forward movement and inversely couples rotational speed to forward velocity. As a result the model periodically makes larger reorientations reminiscent of those seen in behaving ants.

      The behavioral data consists of two experimental sets: experienced Melophorus bagoti foragers, recorded in 2010 and inexperienced M. bagoti foragers, recorded in 2023-2024 at the same site. The behavioral data is qualitatively compared to the model in Figures 3 through 6. In figures 3-5, all ant sets are grouped together while in Figure 6 they are separated. In Figure 6, the authors should do a careful job of making sure the reader is aware that comparisons are being made between behavioral data sets captured more than a decade apart and of justifying the validity of a quantitative comparison between these sets.

      We now make explicit in the methods and figure that the two datasets were recorded at different times: experienced foragers from 2010 (Deeti et al., 2023) and inexperienced foragers from 2023. Their comparison is used to test a qualitative difference predicted by the model.

      The manuscript also describes Myrmecia ants and makes comparisons between modeled Myrmecia ants and supplemental videos of these ants (Videos 3,4). These videos are not described in the methods. While the captions describe these as ants "homing in an unfamiliar environment," the videos show tethered ants walking on a ball. Without more information and absent any analysis, it is difficult for me to understand how these videos support granular points in the text about coupling between rotation and forward velocities.

      We have added a description of Videos 3 and 4 to their respective captions and now state explicitly that these clips were recorded from ants on a tethered trackball apparatus and that they provided as qualitative examples of behaviour discussed in the text.

      Strengths:

      The manuscript's main thesis, that an oscillatory element interspersed between the control signal and the motor unit can reproduce aspects of ant navigation, appears supportable.

      Weaknesses:

      Qualitative agreement between aspects of a model and aspects of a behavioral measurement do not prove the correctness of a model. In the section (802), "An ancestral design? Striking parallels with crawling Drosophila larvae," the authors argue that behavioral data in larvae support their model, despite the larva's lack of a (known) central complex. C. elegans navigation can also be segmented into longer runs and shorter exploratory behaviors (Chen 2025), comparable to the runs and scans described here. C elegans definitively does not have a central complex. In general, multiple internal mechanisms are capable of producing the same macroscopic behavioral outcome. This fact limits the ability of behavioral data to confirm the details of a particular model; it does not imply that observation of similar behaviors in multiple species shows that a particular model is correct or generalizable.

      Here the ability of the behavioral data to confirm or constrain the model is further limited by the qualitative nature of the comparisons. Some of the comparisons are trivial (e.g. Figure 5E-F: any first order process will produce a Poisson distribution, and in the model a Poisson process was explicitly coded in with parameters chosen (1070) to match the behavioral data). Finally, the number of adjustable parameters (13) is comparable to the number of comparisons made; it is unclear that the model could not be adjusted to fit any set of behavioral measurements.

      Our model is a minimal neuro-mechanical model. It is not a mathematical model where each parameter can be optimised to a final output.

      From our 13 parameters, 10 were either taken directly from independent prior studies. 5 concern the oscillator, and have been arbitrarily chosen and simply need to produce regular oscillation (as explained in supplemental material). 4 are the necessary motor gain and noise, which scale neural activations values into movement units, note that this conversion is backed up by previous evidence in drosophila and present in previous ant models). 1 parameter specifies the width of the bump of activity in the CX, and is roughly matched to neural imaging data in flies. None of these parameters have been introduced or adjusted to back up our claim.

      Only 3 parameters have been added to produce scannings. Two of them were tuned to match local scanning-specific data (the probabilistic trigger to stop (p_stop) and the threshold for triggering a saccade (θ_CPG) enabling us to tune fixation duration). This level of parametrisation enables the model to reproduce realistic scans, but does not influence the qualitative predictions of this article. For instance, we agree that the probabilistic trigger producing the Poisson distribution of scan duration (Figure 5’s E) is used to parametrise the model to scanning data, which does not constitute an emerging prediction of the model. We do not include it as evidence (see ~490). Finally, the CX_output_gain, is a new parameter we invoke to implement our hypothesis that CX steering modulates the oscillator. From these three added parameters emerge the large array of behavioural signatures and predictions. These are emerging consequences of the model's architecture rather than curve-fits.

      While the introduction is improved, there is still room to eliminate confusion as to what aspects of the model reflect hypothesized rather than measured neural circuits. For instance, if there is data showing LAL oscillations in insects, the authors should cite it and call it out clearly.

      Alternatively they should say that the oscillator is hypothesized based on measured bistability. They should also clarify whether they are discussing neural oscillations or motor oscillations and whether these oscillations are measured, modeled, or hypothesized.

      As one example: Lines 283-284 "This oscillator [referring to the model's intrinsic oscillator described in the previous paragraph], which is widespread in insects (Cheng, 2024; Kanzaki, 2005; Kanzaki and Mishima, 1996), resides in the lateral accessory lobes (LAL)" reads as though it is known that a neural oscillator occupies the LAL. Cheng 2024 is a brief review of behavioral oscillation. Kanzaki et al. 2005 describes numerical modeling and simulation with a physical robot. Kanzaki and Mishima, 1996 demonstrates bistability (flip-flopping) in moth descending neurons. None of these show neural oscillations and none of them describe the LAL. The authors should review the paper and be scrupulously careful that the claims made in the text are supported in the cited references. These difficulties were pointed out in a previous round of review; hopefully they can be fully corrected this time.

      Kevin S. Chen, Jonathan W. Pillow*, Andrew M. Leifer*, "State-switching navigation strategies in C. elegans are beneficial for chemotaxis," arXiv:2508.00191 31 July 2025.

      We have softened the text to be more explicit about what is modelled versus what is neurally shown (labelled each as behavioural, electrophysiological, or modelled)

      Reviewer #2 (Public review):

      The paper by Freas and Wystrach is an interesting computational study, exploring the detailed mechanisms of how simple neural circuits could explain complex behavioral patterns observed in navigating ants. The authors compare detailed, high speed video recordings of Australian desert ants (Melophorus bagoti) with predictions made by their new computational model and find convincing similarities between the model and the behavioral data, at a level of detail not previously studied. Particularly interesting are emerging properties of the model, yielding behavioral motifs it was not designed to reproduce, but which occur in natural ant behavior.

      A strength of the study is that the model is based on previous models, without making major novel assumptions. It combines existing models of the insect central complex with a model of the lateral accessory lobe and adds a stochastic inhibition of forward velocity to the interaction of central complex and lateral accessory lobes. In essence, the central complex provides corrective steering signals when the goal direction and the current heading of the insect are not aligned, while the lateral accessory lobes provide an intrinsic oscillator underlying the behavioral oscillations shown by walking ants at all times. These background oscillations are modulated by the steering signals from the central complex. Depending on which phase of the intrinsic oscillations coincides with the corrective signals, and how fast the ant is moving forward during this time, a complex set of behaviors emerges.

      Most prominently, scanning behaviors, which are regularly carried out by the ants, are recapitulated in great detail by the model. Additionally, other behaviors, such as full loops, emerge naturally from the model. While computational models are not to be seen as definite evidence for any biological reality, they can provide strong support for particular neural implementations. The current study is an excellent example in that it provides evidence for a serial arrangement of central complex circuits upstream of the lateral accessory lobe circuits, modulated by speed regulating input. While the latter is hypothetical, it yields a clear hypothesis that can be validated by connectomics studies and functional work in the future.

      The computational model is explained in detail and information about all model parameters is provided in an accessible way. The approach is thus transparent and reproducible, leaving it to the readers to assess the assumptions made in the model and how the studied complex behaviors emerge. This also provides the possibility to combine this new model with existing models to expand the scope and to more comprehensively capture the behavioral repertoire of ants, and insects in general.

      Importantly, the study shows that even complex behavioral motifs do not require dedicated neural modules, but can rather emerge from the interplay of already known circuits - highlighting the efficiency of insect brains and possibly providing the path towards embodied hardware solutions of such circuits in autonomous agents.

      We thank Reviewer 2 for this assessment.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      The paper would benefit if the authors would take a more traditional/formal approach to the presentation and interpretation of results. They should avoid drawing conclusions or presenting interpretations in the figure captions (e.g. caption to figure 6) and avoid unnecessary modifiers (e.g. just use "supports" instead of "strongly supports"). This might help correct the tendency of the manuscript to overstate or over-interpret the correspondence between the model and the data.

      Figure captions are now revised to remove interpretive discussion. We also removed unnecessary modifiers throughout the manuscript.

      We also added clarifying text to the “An ancestral design?” section to make clear that we are not implying homologous neural implementations across taxa.

      Reviewer #2 (Recommendations for the authors):

      The authors have addressed my comments fully and I only have a few minor, mostly editorial points:

      line 142: it appears that the references should refer to goal encoding, but both references are head direction papers (one review, one research paper). The only paper showing goal encoding in the CX is Mussels-Pires et al 2024. This should be fixed to ensure the citations are not misleading.

      Changed citations.

      line 165: maybe remove 'intrinsic' to not suggest that this reflects what it known from Biology? It is clear that in the model it is an intrinsic oscillator, but it should not leave the impression that this is an established fact for the LAL

      Removed when not discussing the model.

      line 168: remove either 'a diversity' or 'key qualitative'

      Changed to reproduce multiple key qualitative…. (~Line 170)

      section: 'Neural substrate of insect navigation':

      - '... compares to output steering commands.' grammar is misleading as 'output' might be read as an adjective rather than a verb, replace with 'generate'?

      Changed.

      - as above, the only paper showing goal directions is Mussels-Pires et al 2024. Any other paper either assumes this in models (such as Stone et al and the Honkanen review) or deal with head direction encoding. Pfeiffer and Homberg, 2014 is a general review. Please ensure that references are used more accurately.

      We revised the text to clarify the specific evidence provided by each cited reference.

      - The goal direction in the CX can be updated by various pathways....' After this, behavioral and modeling papers are cited, which is misleading. None of these papers deal with the CX or the neural representation of goals. Same with the rest of the sentence, referring to MB output and PI, which is only shown in modeling, not data.

      We now explicitly distinguish between behavioural evidence and modelling evidence.

      line 210: use CX, not central complex

      Changed.

      Figure 2: I suppose all data shown are modeling data? This should be more explicit in the figure caption (a bit unclear what 'using the neural circuit model.' in the caption heading means. Maybe rephrase to: 'Schematic of neural circuit model and its outputs across navigation relevant brain regions.' (or something like that, putting model first, not brain regions)

      Changed.

      Figure 7: Axis labels in the graphs are still much too small to be read on a printed version (ensure at least 5pt font size in the actual figure on the printed page)

      Enlarged axis labels.

      Line 654: mirror, not mirrors

      Changed.

      line 816: insert 'fly' before larva, as otherwise one might assume this refers to ant larva

      Added (~line 822).

      line 819: The CX does (as we currently know) not exist in fly larvae. At least not as a brain structure, or a set of homologous neurons. There might be equivalent circuits for action selection, but they have not yet been convincingly described. I suggest to rephrase to: 'Although no present as neuropils in fly larvae, the CX and LAL....'

      Changed as suggested (~Line 830).

    1. eLife Assessment

      This valuable study advances our understanding of habituation and related potentiation behavior in a single-celled organism, Stentor. The authors provide convincing evidence for their claims in the form of extensive data on habituation behavior. The work will be of interest to cell biologists, neuroscientists, and also scientists broadly interested in behavior in non-neural organisms.

    2. Reviewer #1 (Public review):

      This interesting paper addresses the phenomenon of potentiation in single-cell habituation in Stentor coeruleus. This is an important "hallmark" of habituation that helps to establish single-cell learning as being similar to habituation in animals. Prior studies from Wood, as well as our own results, have shown that potentiation occurs in Stentor, but I have always remained a little bit skeptical that this effect was possibly just due to incomplete recovery after the first trial. When I first read this paper and saw the habituation curves for the first and second trials, such as in Figure 5, I thought, yes, that is definitely what is happening, and so is this really potentiation?

      The authors were also clearly aware of this issue and, notably, they embraced it head-on by developing an analysis that allows potentiation effects to be detected even despite failure of the cell to fully recover after the first trial. The key is their "phase portrait" that allows the learning process to be depicted as a curve capturing how learning rates and response probability evolve over time, thus allowing the curves to be compared between trials. If my interpretation was correct that so-called potentiation was just incomplete recovery, the prediction would be that the curves for two successive trials would overlap, with the first trial curve extending beyond the second one towards higher response probabilities, which would be lost in the second trial due to failure to recover fully. But the data clearly are not consistent with that idea. I think that this result is very strong and important.

      Especially nice is the approach of Figure 7C, which uses a vertical shift in the phase portrait as an indicator of potentiation. I did, however, find Figure 6 a little hard to digest at first, and I have a few suggestions about that. First, I think it would be a good idea to explicitly say which curve is the first trial and which is the second. Second, I think it would help readers if the authors could start with a cartoon that explains visually what the curves mean. For example, show a habituation curve, indicate how the slope is calculated at different parts of the curve, and then show how the slope versus response are plotted to make the phase portrait. It is all spelled out in the text, but it would help a lot of readers to see it visually, I think.

      One question I have about Figure 6 is that it looks like the specific case of ITI 1 hour ISI 2 min has some kind of pathological behavior in the second trial, despite not seeing any indication of any 'weirdness' in Figure 5. I gather that this is meant to be due at least in part to the incomplete recovery seen after the first trial, but then I don't see why this would not also be an issue for ITI 1 hour ISI 3 min. I would not require the authors to explain every anomaly, but this one stands out, and I feel it could be telling us something interesting.

    3. Reviewer #2 (Public review):

      Summary:

      The authors address habituation and potentiation in the single-celled organism Stentor in a large data set by systematically varying stimulus frequency and recovery duration. They analyze habituation dynamics on the level of single cells within a Bayesian inference framework to map out how the response probability of individual cells decays during training. Mapping out the progression of habituation quantified by learning rate versus decaying response probability, they observe different dynamics for different stimulus frequencies and recovery durations, which they reconcile with multiple time-scales governing the memory of prior training.

      Strengths:

      The authors accumulate a systematic, broad data set of Stentor habituation and potentiation, which, in combination with the Bayesian framework they developed, unfolds its power to probe underlying habituation dynamics and challenge theoretical frameworks.

      Weaknesses:

      The interlacing of theoretical framework, existing concepts and expectation, and experimental data in their narrative may challenge readers. The Bayesian inference of habituations is very successful in concluding that their variation with stimulus frequency and recovery duration points to multiple time scales of memory are involved. However, the authors' comprehensive analysis of potentiation may need more guidance to follow the authors' conclusions.

      The combination of a dynamical systems-driven hypothesis, experimental data, and statistical analysis, as put forward in this work, is immensely powerful for uncovering the mechanisms that facilitate learning, such as habituation and potentiation, in single-celled organisms.

    4. Author response:

      We thank the reviewers and editor for their thoughtful and constructive comments. Our goal was to connect theoretical work on habituation with empirical findings on intracellular habituation in Stentor coeruleus. We developed the phase-portrait analysis to provide a more formal way to examine habituation dynamics and to address the concern that apparent potentiation might simply reflect incomplete recovery. We are glad that the reviewers found this approach useful, and we hope to build on it in future work through mechanistic modeling.

      We will submit a revised version with the following changes:

      (1) We will include a supplementary figure that provides visual intuition for the habituation curves and phase portraits.

      (2) We agree that the anomalous behavior of the 2 min ISI / 1 hr ITI condition is noteworthy. This behavior arises from a subtle difference in the fitted shape of the trial 2 habituation curve: its Hill coefficient is less than 1, so the curve has no inflection point and its initial slope has nonzero magnitude. As a result, the corresponding phase portrait begins away from the x-axis, unlike the other conditions, whose Hill coefficients are greater than 1 and whose phase portraits are U-shaped. We will discuss this explicitly in the revision.

      (3) We will reconsider the layout to make the background and results easier to follow. In particular, we will consolidate the repeated material while preserving the context needed to interpret the theoretical consequences of the empirical findings.

      (4) We will revise the conclusions and discussion to leave claims about the decay of potentiation more open-ended.

    1. eLife Assessment

      This valuable work discusses the phylogenetic conservation of the hippocampal region and primary sensory cortical regions in mammalian species. The authors propose that species-specific differences in behavior and mnemonic functions may be due to differences in cortico-hippocampal connectivity patterns. However, the manuscript, in its present form, is speculative, and the strength of evidence for this proposition is incomplete.

      [Editors’ note: the final version of this work has been published in the journal Hippocampus (https://doi.org/10.1002/hipo.70119)]

    2. Reviewer #1 (Public Review):

      The paper itself has a reasonable aim, to compare the inputs to the hippocampus from cortical regions across mammals. But for some reason, the conclusions that are reached are very limited. We know for example that the main laboratory rodents investigated, rats and mice, are nocturnal, live in underground tunnels, and have a very wide field of view with no fovea. In contrast, primates have a highly developed cortical system for vision and a fovea, and so have very different capabilities to rodents, as they have an ability to identify people or objects at a distance, and to remember where they have been seen. Despite this major difference in the visual cortical processing in these different mammals, somehow important points are missed in this paper about how the cortical processing is organised in these different mammals, and how this is reflected in the anatomy.

    3. Reviewer #2 (Public Review):

      Summary:

      The manuscript emphasizes a phylogenetic conservation of the hippocampal region and primary sensory cortical regions in mammalian species. The authors then propose that the evident species-specific differences in behavior and memory-related functions may be due to differences in type and amount of cortico-hippocampal connectivity.

      Strengths:

      The authors are well-established researchers with a long history of excellent results and publications. The question (co-influence of cortical and hippocampal connections) is potentially interesting.

      Weaknesses:

      The treatment is very broad and macro scale, ignoring the likelihood that hippocampal-cortical connectivity and behavioral outcomes result from multiple differences at a more micro-scale. The designated "mammalian" sample is also broad. Thus, it can appear incomplete as a sample, and incompletely discussed.

    1. eLife Assessment

      In this study, electroencephalography was recorded during a binocular-rivalry paradigm in which participants viewed a target and a distractor stimulus. The stimuli were independently frequency-tagged, allowing the researchers to track target and distractor processing over time. The findings suggest that parietal alpha oscillations initially segregate competing inputs, followed by frontal theta activity associated with suppression of distractor representations. The significance of these findings is valuable, although the strength of the evidence remains incomplete, because of specific experiment design choices and methodological limitations.

    2. Reviewer #1 (Public review):

      Summary:

      These authors used a binocular rivalry task with flickering stimuli in which subjects had to report the color of the target grating at the end of each trial. Target or distractor cues provided information about the orientation of the respective stimulus prior to each trial. The stated goals of this project include testing the neural mechanisms underlying strategic target and distractor processing. Behavioral enhancement was observed for target cueing, while no cost was noted for distractor cueing. These authors present evidence for reactive suppression, characterized by pronounced frontal theta activity that reduced the sensory gain (SSVEP) of the distractor. Distractor cues also increased alpha activity over parietal areas, which these authors link to attentional gating while pointing out no relationship with sensory gain.

      Strengths:

      This manuscript clearly reflects thoughtful analysis of the available data. Alongside a simple and effective task design, sophisticated methods provide good support for most of the claims made by these authors.

      Weaknesses:

      Lack of temporal precision for SSVEP effects. I would like to see how sensory gain is/isn't dynamically modulated in the moments after the initial ERP to see if there could be differences compared to the broader window used presently (1.3 to 3.1 seconds).

      These authors indicate that persistence of the neural representation of cued distractor orientations into the rivalry period is evidence against a "search-and-destroy" type mechanism where distractors are enhanced to then be suppressed reactively. This claim relies on an indirect link between the maintenance of information about distractor orientation (i.e., successful orientation decoding) and the processing of sensory representations. This claim would be backed up more substantially if the SSVEP (a measure of sensory processing) could reveal temporal dynamics on a finer scale.

    3. Reviewer #2 (Public review):

      Summary:

      The findings are conceptually useful - a sequential alpha-then-theta architecture for proactive gating and reactive distractor suppression would be a compelling contribution to the attention control literature - but the evidence is incomplete at best. The central dissociation rests on an inadequate proxy for perceptual dominance, the key alpha-behavior effect is small (d = 0.199) and confined to a single unprotected data quadrant, and the GLMM uses an inadequate random effects structure that inflates false-positive risk.

      Strengths:

      The SSVEP frequency-tagging + binocular rivalry combination is genuinely inventive for isolating sensory gain signals from the two competing stimuli simultaneously. The finding that distractor cueing enhances sensory processing of the distractor yet fails to impair behavior is a clean result that directly addresses a behavioral paradox in the attentional suppression literature. The non-phase-locked TF analysis and the use of RESS for SSVER extraction are methodologically sound.

      Weaknesses:

      The most consequential flaw in the paper is the operationalization of "perceptual dominance." The authors explicitly acknowledge in a footnote that trial categorization as "target-dominant" or "distractor-dominant" is based on which eye received the stimulus, not on participants' actual perceptual reports. Because participants were never asked to report which stimulus was dominant (only to reproduce the target's color), the assignment is an anatomical proxy, not a perceptual measure. This matters enormously for the paper's central claims. Specifically: (a) The entire two-mechanism dissociation (theta for target-dominant trials, alpha for distractor-dominant trials) is built on a trial-type categorization that may not reflect subjective perceptual experience on a given trial, and (b) Dominant-eye stimuli do typically win initial rivalry dominance, but dominance alternates, and in a 2-second window (the stimulus duration used), perceptual states likely fluctuate in many trials. The lack of button-press perceptual tracking (e.g., continuous dominance reports) means the authors cannot verify that their neural effects actually correspond to the perceptual states they claim. This is a major structural limitation of the design that can't be retroactively corrected, and it significantly weakens the consciousness/awareness framing of the findings.

      Another significant issue is that the parietal alpha effect on behavior is confined to a very specific quadrant of the data: distractor-dominant trials where both target and distractor SSVERs are weak simultaneously. The authors present this as an elegant result - "alpha helps most under high perceptual uncertainty" - but it could equally reflect insufficient statistical power for effects in the other three SSVER-strength cells (target strong/distractor weak; target weak/distractor strong; both strong). The Cohen's d for the alpha effect on target reporting probability is only d = 0.199, which is a very small effect. With N=36 and no correction for the multiple SSVER-strength subgroupings tested, there is a real risk that this specific cell-finding is a false positive, while the null in adjacent cells reflects inadequate power rather than a genuine boundary condition.

      A third major limitation is that with a design that includes 6 fixed effects and all their interactions, the random effects structure should include random slopes for at least the key predictors (cueing condition, dominance). Fitting maximal random effects models or justified reduced structures (Barr et al., 2013) is standard in within-subjects EEG research. Using only random intercepts risks inflating Type I error rates for the interaction terms that form the core of the paper's claims. The authors provide a supplementary table (Table S1) but do not describe whether model convergence was verified or alternative random effects structures were tested.

      Fourth, the paper's title and central claim are that alpha and theta dynamics operate sequentially. However, the temporal ordering (preparatory alpha -> rivalry-phase theta) is primarily shown by examining each oscillation in its respective analysis window, not by a single analysis testing whether the sequence itself predicts behavior better than either mechanism alone. A path analysis or cross-lagged model linking trial-level alpha to subsequent theta, and both to behavior, would directly substantiate the "relay" framing. Without this, the sequential architecture is more of an interpretation than a demonstrated property.

      Finally, the frontal theta cluster identified by permutation testing spans 3 to 16 Hz - a range that extends well into the alpha band. Calling this a "theta" effect while simultaneously discussing alpha as a separate mechanism is difficult to reconcile. At minimum, this frequency boundary issue warrants explicit discussion.

    4. Reviewer #3 (Public review):

      Summary:

      Interest was especially focused on how foreknowledge of the orientation of either the target or the distractor could be used to resolve the competition between these stimuli and properly report the target color. The target or distractor was pre-cued by a solid or dashed orientation cue. They were displayed with slightly different presentation frequencies, which allowed for examining their sensory processing with steady-state visual evoked responses (SSVERs). Furthermore, orientation decoding was performed, which revealed that orientation cues selectively affected processing after stimulus onset related to the dominant but not the non-dominant eye. EEG analyses additionally focused on parietal alpha activity and frontal theta, both during the anticipatory phase and the stimulus-processing phase. Cueing the distractor vs. the target induced increased right parietal alpha power during the anticipatory phase, but this did not result in direct inhibition of distractor features. During the stimulus-processing phase, cueing the distractor resulted in increased theta activity. Finally, a generalized linear model was employed wherein trial-by-trial behavior (precision in target color report) was predicted by target and distractor SSVERs, type of pre-cued stimulus (target/distractor), preparatory parietal alpha power, stimulus processing-related frontal theta power, eye dominance, and all their interactions. Performance in the case of reduced sensory processing of the target (based on SSVER) showed more deviations when sensory processing of the distractor was high, but no such effect was observed when sensory processing of the target was high. The latter effects were modulated by eye dominance and cue. Increased theta reduced distractor sensory processing but not target sensory processing. Increased alpha was only beneficial when sensory evidence for both target and distractor was low. Results were interpreted as favoring sensory gating before stimulus onset, reflected by increased parietal alpha (i.e., pro-active control), while theta activity especially seemed relevant to suppress distractor activity (i.e., reactive control) thereby favoring target-related performance.

      Strengths:

      The authors convincingly show that EEG can provide crucial information about how the human brain deals with the conflict between a target and distractor in a binocular rivalry paradigm with pre-cues signaling either the target or the distractor orientation. An important aspect of the study is the focus on precision of target color report, in combination with the possibility to assess SSVERs to the target and distractor. The strength of this study may actually also be its weakness; the question is whether the presented ideas on proactive and reactive mechanisms can be generalized to paradigms that do not employ binocular rivalry. Separation of target and distractor processing by selectively presenting them to the left/right eye increases the conflict when the target is presented at the non-dominant eye, but what happens in the absence of binocular rivalry concerning the target-distractor conflict?

      Weaknesses:

      An important aspect of the study relates to the cue manipulation. In many studies, cues are often informative but not mandatory. Couldn't one argue that in this study task performance crucially depends on cue processing, as without the cue, it becomes difficult to tell apart the target from the distractor. It could be argued that participants are able to do this based on the slight difference in flickering frequency, but I doubt whether this is possible at all. However, if this were the case, then they might use this as an alternative cue and ignore the orientation cue. What do participants experience while performing this task? As the cue can be considered to be mandatory, the question may be raised what strategy the participants actually employed. If the target was cued, they simply may have prepared for this orienting and could ignore the distractor. However, if the distractor was cued, they could use two strategies: search for the stimulus without the cued orientation, or first detect the distractor, and then orient towards the other stimulus. The ideas and results on parietal alpha and frontal theta in combination with the other findings are certainly very interesting, but recently, it has also been argued that frontal theta may be more related to action control (e.g., see Panek et al., https://doi.org/10.1093/cercor/bhaf276) and also pro-active control (Cooper et al., 2017). So, it might be that increased theta reflects suppression of the response related to the distractor, which feeds back on its sensory processing. This raises the question whether there is possibly also some evidence on functional connectivity between frontal and posterior regions that varies depending on the precise condition. Are the results also shining a new light on the relation between attentional orienting and eye dominance (e.g., see Schintu et al., 2020)?

    1. eLife Assessment

      This important study introduces a method, based on active control, for determining the functional connectivity between neurons or groups of neurons. While the theory is derived for a linear model with a perfectly observed control signal, the authors present compelling evidence that their method will work even when the system is nonlinear and the control signal is not fully observed. Furthermore, they provide a practical recipe for doing so.

    2. Joint Public Review:

      Summary:

      Inferring so-called "functional connectivity" between neurons or groups of neurons is important both for validating models and for inferring brain state, including in human patients. This study aims to enhance this inference process by using closed-loop perturbation-based approaches. To this end, the authors develop a framework based on linear dynamical models that minimizes the estimation error. Based on this framework, the authors provide a practical guide for applying it in realistic experiments. Modalities include non-invasive ones, such as fMRI, iEEG, and invasive ones, such as optogenetic perturbations combined with neuropixel probes or calcium imaging.

      Strengths:

      A main strength of this paper is the application and adaptation of an explicit error expression to system dynamics estimation from evoked neural responses, bringing a useful theoretical tool into computational neuroscience for, as far as we know, the first time. Importantly, while the analytical derivation assumes the neural dynamics is linear and the control signal is known, these assumptions do not appear to be essential: their method outperforms passive observation even when the true dynamics is nonlinear or the control input is not known perfectly. Moreover, the relative simplicity of the method makes its practical applications straightforward, as the authors illustrate in the context of brain state classification and neural control.

      Besides being of practical importance, simply pointing out that passive observation can lead to large mis-estimation of functional connectivity should serve as a wakeup call to anybody engaged in this endeavor.

      Weaknesses:

      None.

    3. Author response:

      The following is the authors’ response to the original reviews.

      Joint Public Review:

      Weaknesses:

      (1) The derivation of the main error term misses some important steps, which complicates peer review at this stage. In particular, factorisation of the covariance into noise and the inverse of the observation covariance matrix needs a more thorough justification. The cited sources do not contain the derivation for a noise term with full covariance, which is essential for deriving this error term.

      The derivation of the main error term misses some important steps, which complicates peer review at this stage

      We thank the reviewers for this careful observation. This concern is associated with the error term. Thus, we first clarified the noise assumption explicitly. We assume that ξ(t) is i.i.d. over time with zero mean and an arbitrary (not necessarily diagonal) positive definite covariance matrix .

      In particular, factorisation of the covariance into noise and the inverse of the observation covariance matrix needs a more thorough justification. The cited sources do not contain the derivation for a noise term with full covariance, which is essential for deriving this error term.

      The cited sources do contain the derivation for a noise term with full covariance. We have updated the citation that directly supports Eq. (S.2): Proposition 11.1 of Hamilton (1994, TimeSeries Analysis), which establishes the asymptotic distribution The proof, given in Appendix 11.A of Hamilton (1994), proceeds via a CLT for martingale difference sequences. See Theoretical Details of the Supplementary Materials.

      (2) The practical recommendation at the end of the paper also requires clearer guidance on how the design perturbations are constructed, and how many times and for how long the system is stimulated in each iteration of the experiment.

      Thank you for this helpful suggestion. We agree that the practical implementation of the experimental design should be explained more clearly. We have addressed this concern in two ways. First, we have revised the manuscript to explicitly describe the parameter design procedure. Second, we have revised the manuscript to clearly provide a reference to the detailed experimental condition table in the supplementary material. See Results - Main Manuscript.

      (3) Finally, there is no analysis of model mis-specification. In particular, the true dynamics are unlikely to be linear; the noise is unlikely to be either Gaussian or uncorrelated across time; and the B matrix is unlikely to be known perfectly. We’re not suggesting that the authors consider a more complex model, but it’s important to know how sensitive their method is to model mismatch. If nothing can be done analytically, then simulations would at least provide some kind of guide.

      We thank the reviewer for raising this important point regarding model mis-specification. We agree that it is important to run simulations to assess the impact of these model mismatches, therefore we conducted preliminary simulations to assess the sensitivity when two primary assumptions are violated: linear state dynamics and a perfectly known input matrix B. We added these preliminary results in the Supplementary Material, and revised the main manuscript to include a mention of these results. In summary, the simulations showed:

      The model estimation error increases with the strength of the nonlinearity; however, perturbation can increase the information and lead to accurate estimation of the hidden mode even under nonlinearity.

      The matrix A can be estimated roughly even if the assumed input matrix B differs from the true matrix in some cases.

      The matrices A and B can be estimated jointly without bias, provided the stimulation pattern excites the full state space.

      In such joint estimation, preferentially exciting the hidden modes directly leads to more accurate estimation of A than stimulating other modes. See Background and Results - Main Manuscript, Experimental Conditions and Results - Supplementary Material.

      Recommendations for the authors:

      (1) Please tell us what tACS, tDCS, and TMS are, and how much control the experimenter has over them. That’s important, because they are going to be used as control signals, so we need to know how accurately u(t) can be specified, and what its range is.

      Please tell us what tACS, tDCS, and TMS are, and how much control the experimenter has over them.

      We appreciate the reviewer’s helpful comment. We agree that it is important to describe these stimulation methods with appropriate references. We have added a new section titled “Neural Stimulation as Control Inputs” to the background, which connects our theoretical framework to practical experimental settings.

      We need to know how accurately u(t) can be specified, and what its range is.

      The accuracy and range of control inputs vary substantially depending on the specific stimulation technique and experimental setup, and a thorough discussion would require a dedicated review beyond the scope of this manuscript. Instead, we have added a sentence acknowledging the gap between practical experimental implementations and the theoretical formulation, and cited relevant references for readers interested in further details.

      Background - Main Manuscript

      “Neural Stimulation as Control Inputs

      This section describes how commonly used neural stimulation techniques can be related to input signals in control theory. Their adjustable parameters vary depending on how the stimulation inputs are modulated.

      Three non-invasive electrical stimulation methods illustrate how stimulation paradigms map onto basic control inputs. Transcranial magnetic stimulation (TMS) induces brief and transient perturbations via electromagnetic pulses [19], which are naturally represented as a sequence of impulse-like inputs, where the timing and intensity of each pulse are the primary controllable parameters. Transcranial direct current stimulation (tDCS) primarily modulates neural activity through approximately constant inputs [26], which can be viewed as a step-like signal whose main controllable parameter is the amplitude of the applied current. Transcranial alternating current stimulation (tACS) delivers oscillatory inputs [7], corresponding to sinusoidal signals characterized by amplitude, frequency, and phase. In control theory, impulse, step, and sinusoidal inputs are the basic components used to characterize system responses and dynamics [22, 21].

      The control input framework extends beyond non-invasive techniques to invasive and optogenetic stimulation. Invasive electrical stimulation, including intracranial microstimulation and deep brain stimulation (DBS), enables direct delivery of electrical inputs to neural tissue [17], providing flexible control over amplitude and timing through pulse trains or temporally structured waveforms. Optogenetic stimulation allows genetically targeted activation or inhibition of specific neurons using light [5], providing fine-grained control over multiple input dimensions, including amplitude (light intensity), temporal pattern, and cell-type specificity. In particular, recent developments enable stimulation at the level of individual neurons with high temporal precision [27, 16], allowing flexible construction of spatiotemporal input patterns.

      These stimulation examples demonstrate that the theoretical framework developed in this paper connects to practical experimental settings. While a substantial gap remains between idealized control inputs in theory and experimentally realizable stimulation, the core principles established in the following sections provide a foundation that naturally extends to these practical stimulation paradigms.”

      (2) Is the solid curve in the last panel of Figure 4b the prediction? If not, are the data points consistent with the prediction in Equation 8? This should be clear.

      Thank you for this question. The solid curve shows the empirical eigenvalues of the state covariance matrix, not the eigenvalues of matrix A. As shown in Equation 9, the estimation error is proportional to the inverse of the sum of the eigenvalues of the state covariance matrix. We have clarified these points by revising the main text and the figure captions. See Results - Main Manuscript.

      (3) Page 10, "Nodes 7 and 8 have only outgoing edges". It looks like node 7 has an incoming edge from node 8 (Fig. 6b). Or are we misinterpreting something?

      Thank you for pointing out this inconsistency. You are correct. Node 7 did have an incoming edge from Node 8 and contradicted the statement in the text. Moreover, we now think that two hub nodes (Node 7 and 8) are not necessary for demonstrating our primary theory. To resolve these issues, we have revised the simulation so that the network now contains a single hub node with only outgoing edges. Please refer to Fig. 6.

      (4) Figure 6f, g: why isn't the estimation error proportional to tr[Sig_x^{-1}]?

      We thank the reviewer for this observation. In the original manuscript, the estimation error in Figure 6f, g was plotted on a logarithmic scale, which obscured the proportional relationship with . The underlying values are indeed proportional, consistent with our theoretical prediction.

      In the revised manuscript, we have substantially reorganized Figure 6 to make the theoretical reasoning more transparent by using 1/µ instead of following Eq. 9. The sum across the column in panel (g) is proportional to the estimation error shown in panel (h).

      (5) Page 11, "The simulation was conducted with a single node receiving an impulse-shaped perturbation input." What’s an "impulse-shaped perturbation input"? A delta function? Please make this clear.

      Thank you for this clarifying question. We clarified the explanation as follows.

      Results - Main Manuscript

      “Each node was individually perturbed by an impulse input. Here, an impulse input is defined as a Kronecker delta at t = 0 with fixed amplitude α = 10, with no external input applied at any subsequent time step.”

      (6) Page 11, "Fig. 6d shows that the system possesses modes with small absolute eigenvalues." According to Figure 6d, all the absolute eigenvalues are between 9.9 and 9.98. So this statement appears not to be correct. Are we missing something?

      Thank you for pointing this out. You are correct. The previous statement was inconsistent with the figure. In the revised simulation, we have redesigned the network so that it clearly contains modes with distinct damping characteristics: heavily damped modes with absolute eigenvalues below 0.8 and lightly damped modes with absolute eigenvalues close to 1.0 (see revised Fig. 6c). This makes the relationship between mode damping and perturbation effectiveness much more transparent.

      Results - Main Manuscript

      “Figs. 6c shows the damping rates |λA| of the eigenvalues of A: the mode formed by Nodes 1 and 2 (15Hz) is heavily damped, while those formed by Nodes 3–6 are moderately damped. Node 7 serves as a hub with only outgoing edges.”

      (7) Page 11, "Crucially, the eigenvectors corresponding to these rapidly decaying modes (e.g., evec 7, evec 8) have their largest components concentrated at Nodes 7 and 8." If "nodes" are the same as "eigenvalue index", then the components on node 8 are zero (Figure 6e, bottom). In any case, it should be clear what you mean.

      Thank you for this comment. We agree that the previous description regarding eigenvectors was unclear and inconsistent with the figure. We now think that explaining with eigenvectors is not necessary for demonstrating our primary theory. We have therefore replaced the plots of eigenvectors with the reciprocals of the eigenvalues of Σ<sub>X</sub>, which are directly and rigorously explained by Equation (9). These reciprocals clearly show that the modes associated with Nodes 1 and 2 are heavily damped and applying perturbations to these nodes contribute to the estimation error and the perturbations to hub node (Node 7) broadly increases all eigenvalues of Σ<sub>X</sub>, thereby reducing the estimation error across all modes. This change ensures that all simulation results are grounded in the theory presented in the paper. Please refer to the revised Fig. 6d–g for details.

      (8) It’s not clear to us what’s plotted in Figure 6e. The real part of the eigenvectors? Which would explain why some of the eigenvectors are the same (e.g., 1 and 2). But that does not seem like a good idea, since the eigenvectors can be rotated by an arbitrary complex phase. Also, eigenvectors 7 and 8 seem totally opposite, and there’s no weight at all on index 6 and index 8. Could there be a mistake in the figure? In addition, Figure 6e is explained and interpreted in two different paragraphs, discussing the same observation. It would be easier to understand if they were moved to the same paragraph. Also, in the second paragraph where plot 6e is referenced (page 11, line 19), what does ’concentrated’ mean?

      Thank you for this comment. As described in our response to Comment (7), we have removed the eigenvector-based analysis from the revised simulation to prevent confusing readers. The revised results focus on quantities that are directly explained by Equation (9). Please refer to the revised Fig. 6 for the updated results.

      (9) Page 11, "This simulation demonstrates that, given a tentative connectivity matrix, an effective perturbation input (such as TMS or tDCS) can be designed by targeting the node with the highest weighted out-degree." "Demonstrates" seems strong. There is only one simulation, and that wasn’t totally convincing: a perturbation applied to node 7 did almost the same as a perturbation applied to nodes 1-6, even though it had a much higher outdegree than those nodes. It would be very helpful if you provided theoretical reasoning for why perturbing nodes with high-weight connections (Figure 6) minimises the prediction error. In particular, which properties of a hub node make it effective as a stimulation target? Why is it just its total output weights and not also its number of edges/centrality? How is the high output weight of a node related to its alignment with other eigenvectors, and is this always the case or just in this example?

      Demonstrates seems strong.

      Thank you for this important comment. We agree that "demonstrates" was too strong and have replaced it with "illustrates."

      It would be very helpful if you provided theoretical reasoning for why perturbing nodes with high-weight connections (Figure 6) minimises the prediction error.

      We agree with this comment. We have revised the simulation and accompanying text to connect the results directly to the main theory based on . As responded to Comments (7) and (8), we have removed the eigenvector-based analysis and instead plotted the reciprocals of the eigenvalues of Σ<sub>X</sub>, which are directly related to the estimation error via Equation (9). This change allows us to provide a clear theoretical explanation for why perturbing certain nodes minimizes the prediction error.

      Is this always the case or just in this example?

      We have also explicitly stated the limitations: whether a hub node or a specific subnetwork node is more effective depends on factors such as the outgoing edge weights from the hub and the individual damping rates of each mode. The revised text emphasizes that this simulation presents one example of perturbation location design, and the optimal strategy must be evaluated case by case using the theoretical framework of Equation (9).

      Results - Main Manuscript

      “As this simulation represents only one example of location design, its limitations and the corresponding countermeasures should be stated. In actual experiments, the most effective perturbation location depends on factors such as hub-node connectivity and modal damping rates, and B itself may not always be known a priori. In such cases, approaches such as iterative optimization of (described in a later section) and joint estimation of A and B (Supplementary Material B.3) provide systematic alternatives. Nevertheless, the results presented here provide an intuitive guideline: stimulation directed at hub nodes or at nodes driving heavily damped modes effectively excites the full set of dynamical modes and minimizes estimation error.”

      (10) What are the physical units for the time scales and stimulation amplitudes? For instance, on page 13, there is an argument that "In practical experiments, such long windows are unrealistic because neural states change rapidly over time." However, it is unclear whether T=100 a.u. or T=1000 a.u. etc. is realistic. By relating it to the eigenspectrum of A, which is supposed to be physiologically realistic, one can estimate the length of the stimulation window and support the above statement. Similarly, the impulse amplitude on page 13 is alpha=10<sup>20</sup> (a.u.). Also, Table 1 in the Supplementary has an extremely wide range of stimulation amplitudes. Is 10<sup>20</sup> a.u. a feasible amplitude in practice? And finally, please tell us which nodes the input was applied to.

      Thank you for your incisive comments. We have addressed each comment as follows.

      What are the physical units for the time scales and stimulation amplitudes?

      They don’t have physical units. This study is a theoretical investigation that focuses on the relative differences between passive observation and perturbation-based approaches, rather than providing precise predictions for specific experimental settings. The time scales and stimulation amplitudes are therefore expressed in arbitrary units.

      For instance, on page 13, there is an argument that "In practical experiments, such long windows are unrealistic because neural states change rapidly over time." However, it is unclear whether T=100 a.u. or T=1000 a.u. etc. is realistic.

      We agree that the original expression “unrealistic” was not appropriate given the arbitrary units. We have revised the text to clarify that the time scales are in arbitrary units and that the main point is about the relative difference in required data length between passive observation and perturbation-based approaches, rather than making an absolute claim about feasibility.

      Results - Main Manuscript

      “To obtain estimates under the passive condition that are comparable to those derived under perturbation, it is necessary to experimentally observe extensive time-series data. Figure 7f illustrates the LDA projection and classification accuracy for different time-series lengths. The leftmost LDA plot (T = 20) corresponds to the passive condition shown in Fig. 7c, indicating that the estimation performance in the passive condition becomes comparable to that in the perturbation condition only when the time window reaches approximately T = 200, a 10-fold increase compared to T = 20. While the absolute duration depends on the interpretation of the time unit, such time windows may not be prohibitive in some experimental settings. Nevertheless, our results consistently show that passive observation requires substantially longer recordings to achieve comparable performance, highlighting the efficiency of the perturbation-based approach when the available data length is limited.”

      Is 10<sup>20</sup> a.u. a feasible amplitude in practice?

      Although we have already stated that the stimulation amplitudes are in arbitrary units, we agree that the original value of 10<sup>20</sup> was excessively large and could be misleading. We have revised the impulse amplitude from 10<sup>20</sup> to 10<sup>2</sup>, as 10<sup>20</sup> is physically unrealistic—it would imply a stimulus intensity many orders of magnitude beyond any conceivable experimental setting. The revised value of 10<sup>2</sup> is more plausible: for reference, TMS stimulation voltages exceed typical EEG amplitudes by roughly 4–6 orders of magnitude. The classification accuracy decreased slightly; however, our main conclusion regarding the efficiency of the perturbation-based approach under limited data remains unchanged. See Table B.1.

      Finally, please tell us which nodes the input was applied to.

      The stimulus location was optimized to minimize the . We have clarified this procedure in the main text and added a visual indication of the selected stimulation site (red circles) in Fig.7.

      Results - Main Manuscript

      “The neural signals were simulated under five different task conditions and two stimulation conditions: passive observation and external perturbation. Perturbation was applied as impulse-type inputs, such as TMS. The stimulus location was determined for each task condition by applying an impulse to each node and selecting the one that minimized . The resulting time-series data are shown in Fig. 7b. Using this data, we estimated the underlying dynamical system via a control-based identification approach presented in Eq. 5, which corresponds to an estimation of functional connectivity.

      (11) Page 13: "The controlled transition test was run with T = 1." Previously, T referred to the number of time steps. Is that the case here? If so, that seems hard to justify. If not, please tell us what T is (and, ideally, use a different symbol).

      Is that the case here?

      No. In this context, T does not denote the number of time steps.

      If not, please tell us what T is (and, ideally, use a different symbol).

      Thank you for your helpful comment. We have standardized the notation throughout the manuscript. In this paper, T consistently denotes the number of time steps (i.e., data length). In the sections “Neural State Classification” and “Neural State Transitions,” we had mistakenly used T to refer to time length. To resolve this inconsistency, we have added a separate column labeled “Data Length (T)” to Table B.1 for clarification and replaced the previous usage of T with “data length” where appropriate.

      Results - Main Manuscript

      “To obtain estimates under the passive condition that are comparable to those derived under perturbation, it is necessary to experimentally observe extensive time-series data. Figure 7f illustrates the LDA projection and classification accuracy for different time-series lengths. The leftmost LDA plot (T = 20) corresponds to the passive condition shown in Fig. 7c, indicating that the estimation performance in the passive condition becomes comparable to that in the perturbation condition only when the time window reaches approximately T = 200, a 10-fold increase compared to T = 20. While the absolute duration depends on the interpretation of the time unit, such time windows may not be prohibitive in some experimental settings. Nevertheless, our results consistently show that passive observation requires substantially longer recordings to achieve comparable performance, highlighting the efficiency of the perturbation-based approach when the available data length is limited.”

      Results - Main Manuscript

      “The controlled transition test was run with T = 50. The control objective was to set nodes3 and 4 to 25 while keeping all other nodes at 0 without any movement. See Table B.1.”

      (12) Page 14: "where the matrix A (Fig. 10a) is designed to have 16 modes." What do you mean by has "16 nodes"?

      Thank you for pointing this out. By “16 modes,” we refer to 16 dynamical eigenmodes (i.e., 16 eigenvalue pairs). In the real-valued state-space representation used in the simulations, each complex conjugate pair corresponds to a 2-dimensional real block, resulting in a 32-dimensional system (32 nodes). We revised the wording to clearly distinguish between the number of dynamical modes and the dimensionality (number of nodes) of the state vector, to avoid confusion.

      Results - Main Manuscript

      “Iterative refinement of both the perturbation design and the estimation process progressively improves the accuracy of A. The time-series data is collected from 32 points, where the matrix A (Fig. 10a) is designed to have 16 oscillatory modes (i.e., 16 complex-conjugate eigenvalue pairs, yielding 32 eigenvalues in total).”

      (13) In Figure 10b, the y-axis should start at zero; otherwise, it’s a bit misleading how much the active perturbation helps. This will make it clear that the estimation error drops by about 33%. It would be worth commenting on whether this is typical; after all, potential users of this method would want to know how much improvement they’re likely to see.

      It would be worth commenting on whether this is typical; after all, potential users of this method would want to know how much improvement they’re likely to see.

      Thank you for this insightful comment. We agree with the reviewer that quantifying the expected improvement would be valuable for experimental practice. However, this simulation is a theoretical demonstration. Its primary purpose was to show that an iterative active perturbation approach can progressively converge to an optimal perturbation design even without prior knowledge of the true system, rather than to quantify a universally expected improvement rate.

      The y-axis should start at zero; otherwise, it’s a bit misleading how much the active perturbation helps.

      We believe that the y-axis should start at the accuracy with optimal perturbation because the primary purpose of this simulation was to demonstrate that an iterative active perturbation approach can progressively converge. If we had started the y-axis at zero, the message would be visually obscured.

      Potential users of this method would want to know how much improvement they’re likely to see.

      We acknowledge that this is one example of the application of our method, and the magnitude of improvement depends on various factors such as network structure, noise level, and stimulation design. We should not mislead the readers. We have therefore maintained the y-axis starting point and added a clarifying statement in the revised manuscript to indicate that the simulation is case-specific rather than universal.

      Results - Main Manuscript

      “These results should be interpreted as a case-specific illustration rather than a universal gain, as the magnitude of improvement depends on factors such as network structure, recording duration, noise level, and stimulation design. This simulation demonstrates that our perturbation design framework enables the step-by-step refinement of system identification even without prior knowledge of the system.”

      (14) Please provide a derivation for the factorised covariance in Equation S.2. This is the equation that underpins the main result of the paper - the error in dynamical system estimation. Currently, it appears to be taken from [Hamilton, J. D. Time Series Analysis], yet we were unable to find this result in the book. Most of the derivations in Chapters 8.1 and 8.2 assume diagonal noise covariance, and even isotropic noise (cov = sigma*2 * I), which simplifies the particular case of the derivations. Could the authors provide a reference or the derivation for the case with full covariance?

      As described in our response to the Weakness above, the relevant result is Proposition 11.1 of Hamilton (1994), not Chapters 8.1–8.2. Proposition 11.1 states the asymptotic distribution of the vectorised OLS estimator in a VAR model with i.i.d. innovations whose covariance matrix Ω is an arbitrary positive definite matrix. The factorisation in our Eq. (S.2) corresponds directly to the revised manuscript we explicitly cite “Hamilton (1994), Proposition 11.1” at Eq. (S.1) and in Hamilton’s Proposition 11.1, with and . In state the i.i.d. assumption on ξ(t) in the preceding paragraph, so that the connection to this result is unambiguous. See Theoretical Details Supplementary.

      (15) Strongly perturbing/Exciting fast-decaying modes to give them more ’runway’ and increase observed variability makes intuitive sense for a normal system. But what would happen in the case of non-normal dynamics, where stimulating one dimension only transiently amplifies it, but then excites other dimensions? Non-normality breaks the alignment between PC components and dynamic modes (see Kumar, Ankit, Loren M. Frank, and Kristofer E. Bouchard. "Identifying feedforward and feedback controllable subspaces of neural population dynamics." arXiv preprint arXiv:2408.05875 (2024)), so eigenvectors of A and Σ<sub>X</sub> won’t align for non-normal dynamics. This alignment appears to be a hidden assumption of this paper. Should the normal dynamics then be stated as an assumption/limitation of the framework? Is the example in Figure 6a highly non-normal? Does considering out-degree provide an empirical approach, an alternative to ’enlargement’, to dealing with non-normality?

      We thank the reviewer for this insightful comment regarding non-normal dynamics.

      Should normal dynamics be stated as an assumption/limitation of the framework?

      No. Our framework does not assume normal dynamics. For example, Figure 5 and the revised Figure 6 illustrate non-normal cases. In Figure 5, the row and column norms of A differ substantially (row norms ≈ [1.96, 4.05, 1.33, 4.70]; column norms ≈ [6.14, 1.94, 1.39, 0.87]). In Figure 6, Node 7 is a hub with only outgoing edges, so row 7 of A is zero while column 7 has nonzero entries. In addition, the Nodes 5–6 and Nodes 3–4 are strictly one-way, leaving the corresponding off-diagonal block upper-triangular. Both features break the symmetry required for normality.

      We additionally computed a commutator-based non-normality index which equals 0 for any normal matrix and approaches for the canonical maximally non-normal example (the 2×2 nilpotent Jordan block). The non-normality index is 1.34 for Figure 5 and 0.41 for Figure 6. These values confirm that both are clearly non-normal.

      Is the example in Figure 6a highly non-normal?

      Yes. As stated above, the network structure of Figure 6a clearly indicates non-normality.

      Does considering out-degree provide an empirical approach, an alternative to ‘enlargement’, to dealing with non-normality?

      No. In the previous manuscript, the results of out-degree and eigenvector structure were provided as supplementary intuition. However, the central contribution of our framework lies in maximizing the minimum eigenvalue µ of the observed state covariance (Equation 9), and for systems where non-normality is strong and subnetwork structure is less modular, the iterative optimization of provides a principled, assumption-free method. We clearly mentioned this point in the revised manuscript as follows. See Results in the main manuscript.

      (16) In the iterative experiment at the very end of the paper (Figure 10), what was the strategy for designing ‘u<sub>design</sub>´? How many stimulations were applied in an iteration? Are you stimulating along eigenvectors? Do you sample from components randomly, or perturb each of them individually, with the amplitude proportional to reciprocals?

      Thank you for this important question. We have clarified the design rule and stimulation protocol in the revised manuscript, and address each sub-question below.

      How many stimulations were applied in an iteration?

      One stimulation session was applied in an iteration.

      Are you stimulating along eigenvectors?

      The optimal stimulation is designed as the target node of the perturbation is determined through numerical optimization that minimizes

      Do you sample from components randomly, or perturb each of them individually, with the amplitude proportional to reciprocals?

      The optimal stimulation is designed as a composite-frequency sinusoidal input encompassing all modes of the estimated Â, and the target node of the perturbation is determined through numerical optimization that minimizes . See Results - Main Manuscript.

      (17) While it is clear that the proposed active method performs better than passive observation, some results lack a comparison with stimulating random directions/nodes with a comparable control energy (Figure 7e & Figure 10b).

      Thank you for your helpful comment. Although the optimally designed stimulation outperforms random stimulation, the previous stimulation settings were not configured to explicitly demonstrate this difference. Therefore, we modified the network structure, recording length, and stimulation intensity so that both the main messages and the difference from random stimulation can be shown simultaneously. Accordingly, we have added a comparison with random stimulation in both Figure 7e and Figure 10b.

      In Figure 7e, we included a random stimulation condition where the target node is selected randomly, and the results show that the optimized stimulation outperforms random stimulation.

      These additions strengthen the evidence for the effectiveness of our proposed method compared to non-optimized approaches.

      (18) Is it reasonable to assume full observability of the system? It would be interesting to consider biases arising from the partial observability of the system, in the spirit of Figure 9, which looked at partial controllability.

      We thank the reviewer for this suggestion. We agree that partial observability is an important consideration, however it is out of scope for the current work. Thus, we have added a future direction in the Discussion addressing partial observability. We note that Takens’ embedding theorem and Hankel DMD enable recovery of a system’s eigenvalues from partial observations, and since our framework relies on the eigenvalue structure of A, the proposed perturbation design remains applicable under partial observability.

      Discussion - Main Manuscript

      “Two directions warrant further investigation: extending the framework to partial observability, and validating it through stimulation experiments. In experimental neuroscience, recordings are often limited to a subset of neural populations, resulting in partial observability. A growing body of work has leveraged delay-embedding techniques, represented by Takens’ embedding theorem [25], to reconstruct hidden dynamics from partial observations [2, 3, 23, 11]. Applying such techniques enables the estimation of the full connectivity matrix, thereby extending our framework to settings with partial observability. The second direction concerns experimental validation. Validating a theoretical framework through experimental design is an essential in bridging the gap between theory and practice.”

      Recommendations for improving the writing and presentation.

      (1) The word ’state’ is overloaded in Figure 7. When talking about neural state classification, the ’state’ refers to a regime guided by a distinct dynamics A (should it be A<sub>i</sub>? Figure 7A bottom). However, each dynamical system also has a ’state’. A different word should be used in Figure 7A and the corresponding text.

      We agree that the terminology is potentially confusing. To avoid the confusion, we replaced the term “state” with “task condition” and “neural signal” in the main text, and revised Fig. 7.

      Results - Main Manuscript

      “We designed a neural network with clearly distinct task conditions and considered a simulation setting in which these conditions are classified using signals of a fixed duration. These distinct task conditions consist of five types, each defined by a unique linear dynamical system characterized by differing eigenvalue spectra and connectivity topologies of matrix A (Fig. 7a). These task conditions are intended to mimic different cognitive or behavioral contexts. For example, in a typical motor task experiment, such conditions could correspond to motor execution or imagery involving the left or right hand, or resting state [1, 24]. The neural signal was simulated under five different task conditions and two stimulation conditions: passive observation and external perturbation. Perturbation was applied as impulse-type inputs, such as TMS. The stimulus location was determined for each task condition by applying an impulse to each node and selecting the one that minimized . The resulting time-series data are shown in Fig. 7b. Using this data, we estimated the underlying dynamical system via a control-based identification approach presented in Eq. 5, which corresponds to an estimation of functional connectivity.”

      (2) It would be helpful to provide dimensions of matrices around Equation S.2, since vectorization makes dimensions hard to track.

      We thank the reviewer for this helpful suggestion. We agree that explicitly stating the matrix dimensions improves readability, particularly around the Kronecker product where vectorization can obscure the size of the resulting covariance matrix. We have revised the text as follows (the equation S.2 is 3 now). See Theoretical Details of the Supplementary Materials.

      (3) The background sections of the paper would benefit from referring to similar active perturbation methods: Wagenmaker, Andrew, et al. "Active learning of neural population dynamics using two-photon holographic optogenetics." Advances in Neural Information Processing Systems 37 (2024): 31659-31687. Minai, Yuki, et al. "MiSO: Optimizing brain stimulation to create neural activity states." Advances in Neural Information Processing Systems 37 (2024): 24126-24149.

      We thank the reviewer for these helpful suggestions. We have incorporated Wagenmaker et al. (2024) and Minai et al. (2024) into the Introduction. Their works focus on developing algorithmic approaches to active stimulation design for specific experimental platforms, while our work aims to establish a general theoretical framework for why and which perturbation inputs are effective has yet to be established. We cited these works and have clarified this distinction in the revised manuscript as follows.

      Introduction - Main Manuscript

      “In this paper, we propose a framework for designing the optimal perturbation input through control theory in neuroscience. We interpret neural dynamics as a control system [8, 6, 14, 12, 20, 24], and treat external perturbations as control inputs to design properties of neural stimulation (Fig. 1d). If the optimal perturbation input can be systematically designed, it becomes possible to steer the neural system toward states that are maximally informative (Fig. 1e), thereby enhancing the accuracy of the inferred connectivity (Fig. 1f). While recent studies have begun to develop algorithmic approaches to active stimulation design for specific experimental platforms [18, 28], a general theoretical framework for why and which perturbation inputs are effective has yet to be established. We first describe how to formulate neural dynamics as a control system and how to estimate the model parameters from observed data. Building upon this formulation, we derive a theoretical basis that enables us to design the optimal perturbation inputs for the neural system identification. We demonstrate the validity and utility of this theoretical basis by exploring its implications for optimizing parameters of common neurostimulation techniques and by applying it to practical examples, including neural state classification [4, 9, 1, 24] and control of neural states [8, 13, 14, 12]. In these demonstrations, we define concrete problems and apply the theory to validate its practical utility.”

      Minor corrections to the text and figures.

      (1) In Figure 3, it would be helpful to point out that the plots are in the subspace spanned by the first three principal components.

      Thank you for pointing out. We revised the caption of the Fig.3 as follows. See Results - Main Manuscript.

      (2) Figure 4 caption: "eigenvectors" –> "eigenvectors of Σ<sub>X</sub>", just to make it crystal clear (since A also has eigenvectors).

      Thank you for your suggestion. We have revised the caption of Figure 4. See Results - Main Manuscript.

      (3) Page 5, Equation (8): Capital xi should be introduced in the main text of the paper as the covariance matrix for the noise term xi. Now it can only be understood after reading the Supplementary. Or maybe call the covariance matrix Σ<sub>ξ</sub>rather than Σ<sub>ξ</sub>? That will probably make it clearer.

      Thank you for this suggestion. We have made both changes. First, we introduced the noise covariance matrix explicitly in the main text immediately after Eq. (1), defining. Second, we replaced the notation Σ<sub>ξ</sub> with Σ<sub>ξ</sub>(lowercase subscript matching the noise variable throughout the main text and supplementary.

      (4) Page 7, Equation (14): The description of the equation states "The covariance matrices of x(t) for impulse inputs can be written as:... ". We assume this is supposed to be the "state vector," not "covariance matrices".

      Thank you for catching this. We have corrected the wording to “state vector” instead of “covariance matrices". See Results - Main Manuscript

      (5) Page 10, after introducing Figures 6a-b, potentially a sentence is missing (’...’ in the first line of the last paragraph).

      Thank you for pointing this out. We have removed the placeholder along with updating the stimulation settings for Figure 6.

      (6) Page 10, Figure 6 (e): A clearer labelling would be helpful, e.g., a title for the legend (e.g. node index) and a more informative title (e.g. Eigenvector alignment with nodes), a caption (what is the take-home message?), and y-axis labels (what is the ’value’?).

      Thank you for this helpful suggestion. We have revised the Figure 6 taking your suggestion into account. The new figure includes a clearer title, axis labels, and an informative caption that highlights the key take-home message. See Results - Main Manuscript

      (7) Page 11, paragraph 1: The sentence "...(as discussed in Section)" is missing a section reference.

      Thank you for pointing this out. We have replaced the section reference placeholder with the equation reference to Equation (9), which is the relevant theoretical result.

      (8) Page 13, caption of Figure 8c: compputed -> computed.

      We have corrected this typo. We also reviewed the manuscript for any similar typographical errors and corrected them.

      (9) Page 14 Figure 9: The y-axis label for "Controlled State Process" plots is missing.

      Thank you for catching this. The y-axis was hidden by other elements in the figure. We have revised the figure layout to ensure that the y-axis label is visible.

      (10) Page 16: The sentence "As described in Section, a preliminary ... " is missing a section reference

      Thank you for pointing this out. We have corrected the missing reference. The sentence now reads “as shown in Fig. 10” rather than the incomplete “as described in Section.”

      (11) Page 16, Fig. 10c: It looks like the errors have inconsistent color ranges. A shared colorbar would help.

      Thank you for your suggestion. We have revised Figure 10c to use a shared colorbar across all subplots.

      (12) Page 17, Paragraph preceding eq. 16: "the eigenvectors of X" --> "the eigenvectors of Sigma_X".

      Thank you for catching this. We have revised the text. See Methods - Main Manuscript.

      (13) S.1: This is a GLS, not an OLS estimator, if this assumes full noise covariance.

      As stated in our response to the Weakness above, we assume the noise term ξ(t) is i.i.d. with covariance matrix Σ<sub>ξ</sub>, and the estimator we analyze is the OLS estimator. See Theoretical Details - Supplementary Material.

      (14) S.2: Unclear that⊗is the Kronecker product (not outer), as it is not defined.

      We thank the reviewer for pointing out this ambiguity. In the revised Supplementary Material, we have explicitly explained the Kronecker product with a reference to Hamilton [10, Appendix A.4, p. 732]. See Theoretical Details - Supplementary Material.

      (15) S.51: There is an accidental comma between alpha and A after "xdiff(t) ="

      Thank you for pointing this out. We have removed the accidental comma. See Theoretical Details - Supplementary Material.

      (16) Section B.2: It would be useful to have the simulation details for that section (as is given for the other section in B.1).

      We thank the reviewer for this helpful suggestion. We have added a dedicated “Simulation details” paragraph to Appendix B.5 (the section containing Fig. B.4) so that the setup is now described with the same level of specificity as the other appendix sections. See Experimental Conditions and Results - Supplementary Material.

      References

      (1) Irma N Angulo-Sherman, Marisol Rodríguez-Ugarte, Nadia Sciacca, Eduardo Iáñez, and José M Azorín. Effect of tDCS stimulation of motor cortex and cerebellum on EEG classification of motor imagery and sensorimotor band power. J. Neuroeng. Rehabil., 14(1):31, April 2017.

      (2) Hassan Arbabi and I Mezić. Computation of transient koopman spectrum using hankeldynamic mode decompoisition. APS, page G1.009, November 2017.

      (3) Steven L Brunton, Bingni W Brunton, Joshua L Proctor, Eurika Kaiser, and J Nathan Kutz. Chaos as an intermittently forced linear system. Nat. Commun., 8(1):19, May 2017.

      (4) Adenauer G Casali, Olivia Gosseries, Mario Rosanova, Mélanie Boly, Simone Sarasso, Karina R Casali, Silvia Casarotto, Marie-Aurélie Bruno, Steven Laureys, Giulio Tononi, and Marcello Massimini. A theoretically based index of consciousness independent of sensory processing and behavior. Sci. Transl. Med., 5(198), August 2013.

      (5) Karl Deisseroth. Optogenetics. Nat. Methods, 8(1):26–29, January 2011.

      (6) Shikuang Deng, Jingwei Li, B T Thomas Yeo, and Shi Gu. Control theory illustrates the energy efficiency in the dynamic reconfiguration of functional connectivity. Commun. Biol., 5(1):295, April 2022.

      (7) Shrey Grover, Renata Fayzullina, Breanna M Bullard, Victoria Levina, and Robert M G Reinhart. A meta-analysis suggests that tACS improves cognition in healthy, aging, and psychiatric populations. Sci. Transl. Med., 15(697):eabo2044, May 2023.

      (8) Shi Gu, Fabio Pasqualetti, Matthew Cieslak, Qawi K Telesford, Alfred B Yu, Ari E Kahn, John D Medaglia, Jean M Vettel, Michael B Miller, Scott T Grafton, and Danielle S Bassett. Controllability of structural brain networks. Nat. Commun., 6:8414, October 2015.

      (9) Mark Hallett, Riccardo Di Iorio, Paolo Maria Rossini, Jung E Park, Robert Chen, Pablo Celnik, Antonio P Strafella, Hideyuki Matsumoto, and Yoshikazu Ugawa. Contribution of transcranial magnetic stimulation to assessment of brain connectivity and networks. Clin. Neurophysiol., 128(11):2125–2139, November 2017.

      (10) James Douglas Hamilton. Time Series Analysis. Princeton University Press, Princeton, 1994.

      (11) Ann Huang, Mitchell Ostrow, Satpreet H Singh, Leo Kozachkov, Ila Fiete, and Kanaka Rajan. InputDSA: Demixing then comparing recurrent and externally driven dynamics. arXiv [q-bio.NC], November 2025.

      (12) Shunsuke Kamiya, Genji Kawakita, Shuntaro Sasai, Jun Kitazono, and Masafumi Oizumi. Optimal control costs of brain state transitions in linear stochastic systems. J. Neurosci., 43(2):270–281, January 2023.

      (13) Teresa M Karrer, Jason Z Kim, Jennifer Stiso, Ari E Kahn, Fabio Pasqualetti, Ute Habel, and Danielle S Bassett. A practical guide to methodological considerations in the controllability of structural brain networks. J. Neural Eng., 17(2):026031, April 2020.

      (14) Genji Kawakita, Shunsuke Kamiya, Shuntaro Sasai, Jun Kitazono, and Masafumi Oizumi. Quantifying brain state transition cost via schrödinger bridge. Netw. Neurosci., 6(1):118– 134, February 2022.

      (15) Hassan K Khalil. Nonlinear systems. Prentice-Hall, Upper Saddle River, NJ, 2002.

      (16) Paul K LaFosse, Zhishang Zhou, Jonathan F O’Rawe, Nina G Friedman, Victoria M Scott, Yanting Deng, and Mark H Histed. Single-cell optogenetics reveals attenuationby-suppression in visual cortical neurons. bioRxivorg, page 2023.09.13.557650, May 2024.

      (17) Andres M Lozano, Nir Lipsman, Hagai Bergman, Peter Brown, Stephan Chabardes, Jin Woo Chang, Keith Matthews, Cameron C McIntyre, Thomas E Schlaepfer, Michael Schulder, Yasin Temel, Jens Volkmann, and Joachim K Krauss. Deep brain stimulation: current challenges and future directions. Nat. Rev. Neurol., 15(3):148–160, March 2019.

      (18) Yuki Minai, Matthew Smith, Joana Soldado-Magraner, and Byron Yu. MiSO: Optimizing brain stimulation to create neural activity states. In A Globerson, L Mackey, D Belgrave, A Fan, U Paquet, J Tomczak, and C Zhang, editors, Advances in Neural Information Processing Systems 37, volume 37, pages 24126–24149, San Diego, California, USA, 2024. Neural Information Processing Systems Foundation, Inc. (NeurIPS).

      (19) Davide Momi, Zheng Wang, and John D Griffiths. TMS-evoked responses are driven by recurrent large-scale network dynamics. Elife, 12(e83232), April 2023.

      (20) Ali Moradi Amani, Amirhessam Tahmassebi, Andreas Stadlbauer, Uwe Meyer-Baese, Vincent Noblet, Frederic Blanc, Hagen Malberg, and Anke Meyer-Baese. Controllability of functional and structural brain networks. Complexity, 2024(1), January 2024.

      (21) Norman S Nise. Control Systems Engineering. John Wiley & Sons, 8 edition, 2020.

      (22) Katsuhiko Ogata. Modern Control Engineering. Prentice Hall, 2010.

      (23) Mitchell Ostrow, Adam Eisen, and Ila Fiete. Delay embedding theory of neural sequence models. arXiv [cs.LG], June 2024.

      (24) Yumi Shikauchi, Mitsuaki Takemi, Leo Tomasevic, Jun Kitazono, Hartwig R Siebner, and Masafumi Oizumi. Quantifying state-dependent control properties of brain dynamics from perturbation responses. J. Neurosci., page e0364252025, December 2025.

      (25) Floris Takens. Detecting strange attractors in turbulence. In David Rand and Lai-SangYoung, editors, Dynamical Systems and Turbulence, Warwick 1980, volume 898 of Lecture Notes in Mathematics, pages 366–381. Springer, Berlin, Heidelberg, 1981.

      (26) Liam C Tapsell, Matheus D Pinto, Ann-Maree Vallence, Casey Whife, Maria Luciana Perez Armendariz, Shaswat Senger, Jack Andringa-Bate, Dana Hince, and Myles C Murphy. What are the optimal transcranial direct current stimulation parameters and design elements to modulate corticospinal excitability? a systematic review and longitudinal meta-analysis. Neurol. Res. Pract., 7(1):86, November 2025.

      (27) Lei Tong, Shanshan Han, Yao Xue, Minggang Chen, Fuyi Chen, Wei Ke, Yousheng Shu, Ning Ding, Joerg Bewersdorf, Z Jimmy Zhou, Peng Yuan, and Jaime Grutzendler. Single cell in vivo optogenetic stimulation by two-photon excitation fluorescence transfer. iScience, 26(10):107857, October 2023.

      (28) Andrew Wagenmaker, Lu Mi, Marton Rozsa, Matthew S Bull, Karel Svoboda, Kayvon Daie, Matthew D Golub, and Kevin Jamieson. Active learning of neural population dynamics using two-photon holographic optogenetics. Adv. Neural Inf. Process. Syst., 37:31659–31687, 2024.

    1. eLife Assessment

      This work presents a software and hardware suite for targeted photostimulation that can be used in vivo. The package is a well-designed and documented hardware/software suite with a comprehensive build guide. This tool will likely promote important neuroscience advances through targeted real-time perturbation of the cerebral cortex. Overall, this manuscript makes a compelling case on how to design and make available power tools for the research community.

    2. Reviewer #1 (Public review):

      Lohse et al. describe an open-source system for laser scanning photostimulation (LSPS) in head-fixed animals. Although similar systems have been developed and used by different groups, Zapit provides an open-source solution requiring few custom parts and minimal coding. This tool can clearly facilitate and speed the adoption of LSPS, particularly for the increasingly used purpose of mapping the effects of focal cortical silencing during behavior. Other potential uses include mapping optogenetically evoked movements and selectively activating genetically labeled neuronal subtypes of interest in the cortex. The design is well thought through, and the presentation is mostly clear and well written.

      In general, the more modular such a system is, the better, in terms of compatibility with existing hardware and software that potential users may already have purchased - laser, galvo, and camera in particular. The system has struck a reasonable balance between allowing modularity and providing an integrated complete package, but even more flexibility would be welcome for potential users looking to cut costs, as would clearer presentation of such flexibility as already exists.

      Comments and suggestions are mostly minor, as follows.

      (1) Command signals:

      How is the relationship between analog voltage commands and laser power determined? Is this assumed (or required) to be linear (as Figure 7F implies)? Usability and modularity would be improved by an option to measure or provide a calibration curve for systems with a nonlinear mapping between command voltage and laser power.

      For the grid calibration step, how is the initial mapping from galvo voltage commands to image position determined? Presumably, some sort of initial guess or calculation based on the hardware specifications is needed for the grid calibration to be feasible. Also, how are the number of grid lines and the distance between them determined?

      Why is the mapping between analog outputs and hardware (galvos, laser, masking light) fixed? This would be trivial to make configurable and allow labs with existing setups to adopt Zapit without rewiring existing hardware.

      (2) Laser and optics:

      In Figure 1, the authors should consider explaining the scanning principle schematically, i.e., depicting how tilting of the scan mirrors translates via the scan lens into beam displacement in the specimen plane. Perhaps Zemax can be used for accurate rendering.

      Since the unexpanded beam greatly under-fills the back aperture of the lens, the z resolution is presumably terrible - which is good! That is, for the purposes of LSPS, this advantageously avoids focus-dependent effects, which might otherwise arise due to (e.g.) skull curvature. The authors should consider pointing this out, as well as providing an estimate of the z resolution.

      What is the working distance?

    3. Reviewer #2 (Public review):

      Summary:

      In this work, Lohse and colleagues develop a system for doing targeted photostimulation in mouse cortex. The system uses a camera image to target laser stimulation to stereotactically defined locations in mouse dorsal cortex.

      Strengths:

      The hardware is well designed, and the software is well documented and supported. The build guide and well-documented software package should allow for simple implementation of the technology. Without a doubt, this is a valuable community resource for the circuit neuroscience field.

      Weaknesses:

      No weaknesses were identified by this reviewer.

    4. Reviewer #3 (Public review):

      Zappit is an open-source implementation of arbitrary-access laser-scanning optogenetics for manipulation of neuronal activity in mice. As the method requires expertise ranging from optics, hardware control and programming, the authors make the point that this powerful strategy is underutilized in the field, and put forward a well-documented modular hardware and software platform aligned to the Allen Mouse Brain Atlas aimed at enabling the larger scientific community to use this approach (democratizing) for controlling cortical activity during behavior in mice.

      The authors favor a galvanometric approach to laser targeting. The system is inexpensive, easy to build, well-documented and user friendly (Matlab based GUI and GitHub repository). The photo-stimulation laser is directed into an X-Y galvo scanner targeted to the specimen using a dichroic mirror and focused on the sample using a Plössl lens as scan lens which is also used as an objective. The scan lens/objective images the specimen onto a camera via tube lens (also a Plössl lens) in a 0.5X magnification ensuring to fit the extent of the mouse brain onto the camera sensor (USB-3 Basler acA120-40um).

      The authors report short and reproducible onsite time (~ 0.5 ms) and block (mask) the stimulation source using the laser analog control (~0.5 ms). The system is reliable, aiming at up to 20 stimulation sites per sequence considered as quasi-simultaneous (10 ms). They minimize rebound by gentle ramping down of stimulation over 250 ms.

      The system is fast to calibrate by mapping scanner positions to pixel space in the camera space and mapping stereotaxic coordinate onto the image of the exposed skull. The theoretical x-y PSF is 70 µm (measured ~90µm) while the authors make the point that due to scattering the photo-stimulation spot size (lateral extent) is about 1 mm in diameter. This is what they also observe in electrophysiological recordings using silicon probes. The effective radius of inactivation depends on laser power, but was about 1 mm for laser powers (1-2-4 mW) on which the authors observed significant behavioral perturbations - in several tasks: 1) a delayed response somatosensory discrimination, 2) a visual detection task assessing changes in temporal frequency of a drifting visual stimulus; and 3) a visual discrimination (International Brain Laboratory task) in which mice were tasked to report the location of visual stimuli by turning a wheel. As proof of principle, the authors used a photo-stimulation set composed of 52 bilateral sites positioned at 0.5 mm interval covering a large network of frontal, motor and somatosensory cortical areas. Indeed, photo-inhibition of frontal motor cortex sites produced robust increases in reaction time. In contrast, stimulation at other motor and somatosensory sites produced modest, but significant decreases in reaction times.

      While the approach is not novel, it does serve the need of better disseminating this technique in the research community. Overall, the Zappit is well-documented and easy to build and use, and will have impact in increasing robust use of site directed photo-stimulation (exciting/inhibiting ensembles of neurons at particular ~1 mm size regions of interests across the dorsal surface of the brain). The authors also note that the axial resolution is ~1.5 mm.

      Concerns & comments:

      (1) While the authors argue that it offers the best utility to affordability trade-off - faster than motorized drivers and require much less power than DMDs (100X) and less expensive/easier to use compared to SLMs, in the current form, the manuscript does not clearly list the limitations of the approach. At such, in my opinion, the authors should include side by side comparisons (perhaps as a table). For example, clear statements should be included with respect to comparisons in lateral (x-y), axial (z) spatial resolution, as well as temporal sequential aspect of Zappit and other photo-stimulation techniques involving DMDs or SLMs.

      (2) Is power really a limitation in terms of the laser sources? Or is this a disadvantage mainly because using less power has beneficial effects on the tissue health? It may be useful to provide metrics of comparisons along these lines between Zappit and DMD-based approaches.

      (3) Arbitrary-scanning vs random scanning may be more appropriate to describe to strategy.

    1. eLife Assessment

      Mechanical transduction channels of sensory hair cells possess lipid scramblase activity. Membrane lipid disruption resulting from mechanical transduction is thought to be restored by flippase activities. This fundamental study provides compelling evidence that ATP8B1, a P4-ATP flippase and its subunit TMEM30B, are key in mediating this restorative function in outer hair cells of the mammalian cochlea.

      [Editors’ note, September 9, 2026: There is a potential confound with the findings shown in Figure 7 suggesting that specific targeting of ATP8B1 and TMEM30B subunits to the hair cell stereocilia is dependent on mechanoelectrical transduction channel activity. This does not affect the central findings or conclusions, but readers should be aware that a revised version of the paper is being prepared.]

    2. Reviewer #1 (Public review):

      Sensory hair cells of the inner ear convert mechanical sound vibrations into electrical signals through mechano-electrical transduction (MET). While the protein components of the MET machinery have been studied extensively, much less is known about how the surrounding membrane lipid environment contributes to hair cell function. The recent discovery that TMC1 and TMC2 also function as lipid scramblases has brought renewed attention to the importance of membrane lipid asymmetry and the mechanisms that maintain it in sensory hair cells.

      In this study, the authors identify the P4-ATPase ATP8B1 and its partner TMEM30B as key regulators of membrane lipid asymmetry in outer hair cells. Using complementary genetic models, HA-tagged knock-in mice, localization analyses, and functional experiments, they show that ATP8B1-TMEM30B is enriched in stereocilia and the apical membrane of outer hair cells and is required to maintain phosphatidylserine asymmetry, support hair cell survival, and preserve normal hearing. The parallels between the ATP8B1/TMEM30B loss-of-function phenotypes and TMC1 deafness-associated mutants with constitutive scrambling support a model in which ATP8B1-TMEM30B flippase activity maintains membrane lipid asymmetry and homeostasis, whereas constitutive TMC1-mediated phospholipid scrambling disrupts this balance and contributes to membrane instability.

      The authors have addressed the points raised during the initial review thoroughly. The revised manuscript includes clearer methodological details, additional physiological characterization, improved presentation and quantification of several datasets, and a more balanced interpretation of the localization and mechanistic findings. These changes improve both the clarity and rigor of the study while leaving its main conclusions unchanged.

      As with any study that opens a new area of investigation, important mechanistic questions remain. In particular, it will be interesting to determine how disruption of membrane lipid asymmetry ultimately impairs MET function and triggers hair cell degeneration, how flippase and scramblase activities are coordinated in vivo, and how these pathways are integrated with the broader molecular machinery underlying mechanotransduction. These questions highlight the exciting directions that this study opens for the field.

      Overall, this work provides evidence that ATP8B1-TMEM30B is a critical regulator of stereocilia membrane lipid asymmetry and represents an important contribution to our understanding of membrane homeostasis in auditory hair cells. I have no further major concerns and support publication.

    3. Reviewer #2 (Public review):

      Summary:

      Prior work identified TMEM30B (knockout mice) as well as ATP8B1 (human genetics and mouse model), ATP8A2 (knockout mice), and ATP811A (human genetics) as relevant for hearing. The authors also reasoned that given the recent discovery of TMC1 and TMC2's dual function as mechanotransduction channels of the inner ear and as lipid scramblases, a counterpart flippase should be in the sensory hair-cell stereocilia bundle where mechanotransduction happens. They use CRISPR/CAS to modify the endogenous mouse genes and add an HA tag at the N-terminus of the ATP8B1, ATP8A1, ATP8A2, and ATP11A proteins. Their experiments with these mice unambiguously localized ATP8B1 at the base of outer hair cell stereocilia bundles. Knockout of ATP8B1 results in loss of outer hair cells, deficient auditory function (ABR), and degeneration of outer hair cell stereocilia bundles. Similarly, hair cells from genetically modified mice with endogenous HA-tagged TMEM30B proteins show localization of this protein to outer hair cell stereocilia bundles. TMEM30B knock out mice phenocopy the ATP8B1 knock out model. Interestingly, the authors show that annexing V staining precedes hair cell loss in ATP8B1 and TMEM30B knockout mice and that proper localization of these proteins is lost in mice that lack CIB2, a protein essential for hair cell mechanotransduction.

      Strengths:

      (1) Use of knock-in HA-tagged proteins to unambiguously localize ATP8B1 and TMEM30B

      (2) Systematic characterization of auditory function (ABR), hair cell loss, and hair-cell stereocilia bundle morphology.

      (3) Advances our understanding of the role played by lipid homeostasis in auditory function.

      (4) Reports on mouse models that will be helpful to further understand the mechanistic role played by ATP8B1 and TMEM30B in normal hearing and hereditary deafness.

      Weaknesses:

      (1) Are the HA tags causing any functional issues? Function and localization of tagged proteins can sometimes be compromised. This is checked for TMEM30B and ATP8B1, but not for ATP8A1, ATP8A2, and ATP11A.

      (2) Following on the point above, is it possible that ATP8B1-HA is well localized, but localization for the other three flippases (ATP8A1-HA, ATP8A2-HA, and ATP11A-HA) is compromised by the tag? Is this potential miss-localization causing any functional phenotypes? I find surprising that there are flippases only in outer hair cells and only formed by ATP8B1. A possible explanation is that the tag is interfering with trafficking. If so, there should be a phenotype (ABRs), although this might be masked by redundancy among these flippases or caused by systemic issues (admittedly difficult to sort out).

    4. Author response:

      The following is the authors’ response to the original reviews.

      Summary of Revisions Performed:

      We have clarified the qPCR methodology in the methods section and stated the housekeeping gene GAPDH to address potential misunderstandings.

      We have assessed hearing in the generated HA-tagged mouse lines and included an adequately powered ABR measurements analysis in the revised manuscript as a supplemental figure.

      We have included powered DPOAE experiments in both ATP8B1 and TMEM30B KO mice to strengthen the findings of the ABRs.

      We have clarified the presentation of the z-stack in Figure 1F.

      We have elaborated on the analysis for Figure 7B to strengthen comprehension by readers.

      We have revised the statement to read: “No IHC stereocilia-enriched P4-ATPases were detected under the conditions examined.”

      While we appreciate the suggestion to examine TMEM30B localization on the ATP8B1 KO background, this is not feasible within a reasonable timeframe; we have clarified this limitation in the manuscript.

      We have incorporated relevant prior work (e.g., George and Ricci, 2026) demonstrating minimal Annexin V labeling prior to P6 and lack of PS externalization in TMC1/2 double knockout models.

      We have clarified that hearing thresholds for TMEM30B-HA and ATP8B1-HA lines were addressed in this study, while additional HA-tagged flippase lines (ATP8A1, ATP8A2, ATP11A) are part of ongoing work to be reported separately.

      We have softened statements regarding HA-tag insertion and clarified that, to our knowledge, localization and function are not disrupted, while acknowledging this as a potential limitation.

      We have revised the Methods section to clarify differences in fluorescence measurements across experiments.

      Public Reviews:

      Reviewer #1 (Public review):

      Figure1D.

      The authors should clarify how the qPCR data were normalized and specify the reference (housekeeping) genes used. This information is necessary to evaluate the robustness and comparability of the gene expression data.

      We thank the reviewer for this comment. qPCR data were normalized to GAPDH as the reference (housekeeping) gene. We have clarified this in the Methods section to ensure transparency and reproducibility.

      (2) Figure 1F.

      The lack of F-actin staining at the hair cell base raises the possibility that the permeabilization conditions may have limited antibody access to certain membrane regions. This is especially important given that the authors used a gentle permeabilization agent such as saponin to preserve membrane integrity. Because the authors conclude that ATP8B1 and TMEM30B are localized "almost exclusively to OHC bundles and the apical membrane, with minimal staining in the remaining plasma membrane," (line 128). Including co-labeling with a plasma membrane marker or more comprehensive F-actin visualization of lateral and basal regions would help ensure that the restricted localization is biological rather than technical. In the absence of such controls, the localization claim may be somewhat overstated and should be tempered accordingly.

      We thank the reviewer for this important point. The image shown represents a single z-slice from a larger stack, and the hair cell body lies outside the plane of this section. To clarify this, we revised the accompanying text.

      (3) Figure 7B.

      Although quantification of ATP8B1-HA intensity at the bundle appears similar between WT and Cib2 KO samples, the representative image suggests that some bundles lack detectable labeling. To better capture phenotype variability, it would be helpful to include an additional quantification showing the fraction or number of bundles with detectable ATP8B1-HA signal in Cib2 KO mice.

      We thank the reviewer for this suggestion. We have clarified the quantification of the fraction of hair cell bundles with detectable ATP8B1-HA and TMEM30B-HA signal per field of view. Although the representative images may give the impression that some hair bundles lack staining, this is due to changes in ATP8B1-HA and TMEM30B-HA distribution within the cell body. In all cases, detectable ATP8B1-HA and TMEM30B-HA signal remained present in the hair bundles.

      (4) Lines 346-349

      The manuscript suggests that IHCs lack stereocilia-enriched P4-ATPases. However, this conclusion is not directly supported by the presented data. The authors should either provide supporting localization or expression data for other P4-ATPases or soften the statement to indicate that no stereocilia-enriched P4-ATPases were detected under the conditions examined.

      We agree with the reviewer and have revised this statement to read: “No IHC stereocilia-enriched P4-ATPases were detected under the conditions examined.”

      Recommendations:

      (5) The authors convincingly demonstrate that TMEM30B loss results in ATP8B1 mislocalization. While not essential to the central conclusions, examining TMEM30B localization in ATP8B1 KO hair cells would clarify whether this interdependence is reciprocal, as described for other P4-ATPase-CDC50 complexes.

      While we agree that this experiment would provide valuable information, performing it would require generation of a compound mouse line carrying both the TMEM30B-HA allele and the ATP8B1 knockout allele. This work is beyond the scope of the current revision and cannot be completed within a reasonable timeframe.

      (6) Lines 359-374. The discussion of Annexin V labeling is careful and balanced. This paragraph would benefit from referencing other studies that showed minimal Annexin V labeling in healthy P6 organ of Corti, reinforcing that robust PS externalization in the present study is pathological rather than developmental.

      We thank the reviewer for this suggestion and have incorporated relevant prior work, including George and Ricci (2026), which demonstrates minimal Annexin V labeling prior to P6 and further supports our interpretation.

      (7) Lines 392-399.

      The proposed feedback model linking MET activity and ATP8B1-TMEM30B localization is compelling. The discussion could be strengthened by noting that in TMC1/2 double knockout hair cells, PS externalization is not observed, consistent with the idea that flippase activity becomes critical specifically when scrambling occurs. The mislocalization observed in Cib2 KO hair cells further supports the coupling between TMC-mediated scrambling and flippase-mediated membrane restoration.

      We agree and have revised the text to include that TMC1/2 double knockout hair cells do not exhibit phosphatidylserine externalization, supporting the idea that flippase activity becomes critical in the context of scrambling.

      Reviewer #2 (Public review):

      Weaknesses:

      (1) Are the HA tags causing any functional issues? Function and localization of tagged proteins can sometimes be compromised. It would be good to know, for each knock-in model (TMEM30B, ATP8B1, ATP8A1, ATP8A2, and ATP11A), whether the HA-tagged protein is causing any issues with the mice and particularly with hearing (ABRs). Are these mice normal? Can they hear? These data are missing.

      We thank the reviewer for raising this important point. In this study, we focus on TMEM30B-HA and ATP8B1-HA mouse lines, while additional HA-tagged flippase lines (ATP8A1, ATP8A2, ATP11A) are part of ongoing work to be reported separately.

      Both TMEM30B-HA and ATP8B1-HA mice are viable and exhibit normal breeding and ageing. We have included adequately powered ABR measurements of both TMEM30B-HA and ATP8B1-HA which indicate wild-type–like hearing thresholds.

      (2) Following on the point above, is it possible that ATP8B1-HA is well localized, but localization for the other three flippases (ATP8A1-HA, ATP8A2-HA, and ATP11A-HA) is compromised by the tag? Is this potential mislocalization causing any functional phenotypes? (ABRs of point 1). I find it surprising that there are flippases only in outer hair cells and only formed by ATP8B1. A possible explanation is that the tag is interfering with trafficking. If so, there should be a phenotype (ABRs), although this might be masked by redundancy among these flippases or caused by systemic issues (admittedly difficult to sort out). Given that this manuscript will likely become foundational, and that there is evidence that at least two of the other flippases are involved in hearing loss, it would be good to provide more information about the mice and HA-tagged proteins in the other knock-ins (ATP8A1-HA, ATP8A2-HA, and ATP11A-HA). Depending on the data available for the knock-ins, the authors may want to discuss these scenarios and soften the statement indicating that inner-hair cells may lack flippase activity altogether.

      We appreciate this concern. To our knowledge, the HA tag does not appear to disrupt localization or function of the tagged proteins. However, we agree that this cannot be fully excluded. We have therefore softened our conclusions about IHC flippases and clarified that additional flippases (ATP8A1, ATP8A2, ATP11A) are under investigation and will be described in a separate study.

      (3) Expression of ATP8B1 at P0 (Figure 1D), when there should not be protein in outer hair cells yet seems high. Does this mean that other cells in the cochlea also express ATP8B1? Is this a concern?

      We thank the reviewer for this observation. We interpret the elevated ATP8B1 transcript levels at P0 as reflecting transcription that precedes detectable protein accumulation in OHC stereocilia. While expression in other cochlear cell types cannot be excluded, we did not detect ATP8B1-HA immunolabeling outside hair cells in the knock-in model.

      (4) Fluorescence scales in Figure 6 B and D and Figure 7 B and D are very different. So are the values for WT. One would expect that the WT would be similar in all cases (at least within the same compartments), given that the methods section indicates that "All images were collected using identical acquisition parameters, including zoom and laser power, across genotypes". If WT shows such variability, how can we compare?

      We appreciate the need for clarification. Identical acquisition parameters were maintained within each experiment used for direct comparison (e.g., within a given panel). However, different panels (e.g., Figures 6B vs. 6D) were acquired on different days using different imaging settings.

      We have revised the Methods section to explicitly state this and clarify that comparisons are intended only within panels, not across experiments.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) Line 42: When discussing TMC similarity to TMEM16 scramblases, it may be helpful to mention that some TMEM16 family members (TMEM16A and B) function as ion channels, highlighting the dual ion/lipid functionality within the superfamily. The similarity to TMEM63/OSCA ion channels and lipid scramblases could also be noted. The fact that TMC, TMEM16, and TMEM63/OSCA belong to the same superfamily would provide a broader context. I also suggest referencing the work that initially suggested this relationship: (https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0192851)

      We have included this citation and expanded on this discussion in the revised version.

      (2) Line 45: Consider including the recent Cryo-EM structure of CeTMC2 (PNAS, 2023), which provides structural insight into TMC-lipid interactions.

      We have included this citation in the revised version.

      (3) Line 62: For precision, consider removing "calcium-activated," as caspase-activated scramblases also disrupt membrane asymmetry.

      For precision, we have removed “calcium-activated” in the revised version.

      (4) Line 71: Clarify whether this refers to "fusion of membranes" or "cell fusion."

      We have clarified this statement to mean cell-cell fusion.

      (5) Line 85: Consider citing studies showing constitutive PS externalization in TMC1 mutant mouse models linked to deafness.

      We have added a citation to show that constitutive PS externalization is linked to deafness (Ballesteros and Swartz, 2022, and Beurg et al. 2025).

      (6) Line 93: TMEM30C is not discussed. A brief comment on its expression or relevance in hair cells would provide completeness.

      We have added a brief statement regarding TMEM30C and cited prior work describing its expression pattern (Osada et al. 2007).

      (7) Figure 1A: Use distinct colors for the P4-ATPase and CDC50 subunit rather than a rainbow scheme to improve clarity.

      We have retained the original color scheme in this panel.

      (8) Figures 3C-D and 5C-D: Increase legend symbol size for clarity. Update Y-axis labels to "Number of OHCs/100 μm" and "Number of IHCs/100 μm." Correct "um" to "μm."

      We changed the legend to improve the presentation of these panels to be more legible and changed the measurement to μm.

      (9) Figures 3F, 5F, 5H: Add scale bars.

      We have added scale bars to these figures.

      (10) Figure 7: The confocal images (A, C) show the bundle on top and cell body below, but the quantification (B, D) is in the opposite order. Reorganizing the panels for consistent orientation would improve clarity.

      We have reorganized the panels to improve clarity.

      (11) ABR measurements: Please specify the sex of the mice tested or clarify whether both sexes were included.

      We have included both male and female mice in this study as there were no differences in hearing function. We have added this clarification to the methods section under hearing tests.

      Reviewer #2 (Recommendations for the authors):

      (1) In Figure 1A, the panels show CDC50. I would either change to TMEM30B or mention in the caption that TMEM30B is also known as CDC50 as labeled in the figure.

      We have changed CDC50 to TMEM30B.

      (2) Figures 1F and 1G are missing scale bars.

      We have added scale bars to these figures.

      (3) Figures 2 C, D, and 5 C, D - difficult to tell what's what in the legend. Perhaps make symbols larger in front of WT P17, KO P17, etc.?

      We changed the legend to improve the presentation of these panels.

      (4) Text under "TMEM30B is required for hearing and OHC maintenance". There is a difference in phenotype between the TMEM30B (Figure 5C) and ATP8B1 (Figure 3C) knockouts that is not discussed, as apical and middle cells seem to be okay. Should this be discussed?

      We appreciate this observation. We have elected not to expand the discussion of these regional differences because apical and middle hair cells also undergo degeneration at later ages (after P30), suggesting that the observed differences primarily reflect the timing of degeneration rather than distinct underlying mechanisms.

      (5) In the discussion text, under "Why do ATP8B1/TMEM30B-deficient OHCs die?", "Tmc1/2 or Cib2" should probably be "TMC1/2 or CIB2" or "Tmc1/2 or Cib2"

      We have changed this to read TMC1/2 or CIB2.

      (6) The methods section states "..., whereas non-significant comparisons are not shown." However, non-significant p values are shown in Figures 7B and D (bottom panels).

      We have removed the nonsignificant comparisons from Fig 7B and D to be consistent with the rest of the paper.

    1. eLife Assessment

      This study used several approaches (computational modeling, neuronal cultures, rodent epilepsy model, and human intracranial recordings) to address a significant question in epilepsy research, with additional relevance to EEG studies more broadly: Are high frequency oscillations (or "fast ripples", defined as >200 Hz by the authors) distinct from randomly occurring clustering of spikes? The results suggest fast ripples can occur by chance and how this may occur. The significance was considered important and the strength of evidence convincing, with minor limitations related to the need to address behavioral state, explaining the results in relation to epileptiform activity described by others, and discussing implications.

    2. Reviewer #1 (Public review):

      Summary:

      This is a study utilizing several types of analyses (computational modeling, neuronal cultures, rodent epilepsy model, and human intracranial multi-scale recordings) to address a highly relevant conceptual question: Are fast ripples (FRs) distinct pathological entities or largely emergent products of stochastic spike clustering? The results can potentially reshape current approaches to incorporating fast ripples into the epilepsy surgery evaluation.

      Strengths:

      The conceptualization of fast ripples as potentially arising by chance is highly novel and builds effectively on questions raised in prior studies that have never been satisfactorily resolved. Integration across biological scales and models provides a rigorous approach, now improved by addressing theoretical concerns regarding validity of the shuffling approach and state dependence. The discussion has been updated to provide a more nuanced interpretation of the study's findings.

      Weaknesses:

      The authors have satisfactorily and thoughtfully addressed the critiques provided in the first review. However, there remain two points that I would like authors to address:

      (1) Synchronized burst firing is a key feature of an epileptic site generating interictal discharges, and one that could generate either oscillatory or stochastic FRs as documented in multiple prior publications cited in the manuscript and/or in the prior review. Paroxysmal depolarization, for example, has been very well described, and consists of strong, disorganized burst firing (resulting in summated postsynaptic potentials strong enough to generate high gamma signal) in a neuronal population coinciding with a large low-frequency deflection. I would like to see the results described in this context, and to avoid blanket dismissal of stochastic FRs without a clear oscillatory component.

      (2) It would be highly useful to add a conclusion paragraph that spells out implications of the study for use of FRs as epileptic biomarkers in clinical invasive EEG recordings.

      Please address the above critiques in Discussion, or elsewhere as deemed necessary by the authors.

    3. Reviewer #2 (Public review):

      Summary:

      This paper asks an important question that has not been discussed much in the extensive literature on the High Frequency Oscillations (HFOs) that have been extensively studied in patients with epilepsy and experimental models of epilepsy. The question is whether the Fast Ripples (FRs), the HFOs in the 250-500 Hz frequency band, represent a pathological phenomenon or represent a physiological phenomenon that occurs in the healthy brain but happens to be more frequent in epileptic tissue. It is an important question that has not been systematically addressed until now. The authors conclude, from very extensive simulations, from extensive experimental animal studies (the systemic kianate model of epilepsy in rats), and from a modest amount of human data, that FRs occur in healthy brains as a result of the chance occurrence of bursts of action potentials, and that in epileptic tissue, their frequency of occurrence is approximately 30% higher than what is expected by chance. They conclude that FRs are not a separate phenomenon of epileptic tissue. This finding is reinforced by the recent findings of FRs in experimental models of Alzheimer's disease.

      Strengths:

      This is a valuable study because it asks an important and original question and because it evaluates it from several angles (simulation, tissue culture, experimental animals, and human patients). The simulations and the analyses of real data are performed very carefully and with original and solidly documented approaches, using extensive simulations and extensive data sets in the cultured cell data and in the in vivo experiments. The paper is clearly written and well-illustrated.

      Comments on revised version.

      The authors have appropriately addressed the questions I raised in the first review.

    4. Reviewer #3 (Public review):

      Summary:

      An outstanding question in the field of high frequency oscillations (HFOs) in the context of epilepsy is how these oscillations emerge, considering that they occur at such high frequencies i.e., 250Hz well above the firing ability of single neurons. One hypothesis that has been suggested in the past is that neurons that fire in an out of phase fashion or rather at random intervals may contribute to a spectrum of HFOs ranging from 250-500Hz that observed in epilepsy. However, how possible it is that random action potentials could aggregate to the extent that they could give rise to HFOs in the so-called fast ripple (FRs) frequency range (>200 according to the authors) remains unclear. To test this hypothesis, they used computational modeling to randomly insert action potentials in a signal, and they found that this approach is sufficient to generate FRs. Some of the predictors of whether FRs could occur were neuronal count, firing rate and synchronization. Besides computational modeling, they used different model systems to test whether that would be possible to be observed in neuronal cultures, in epileptic rats (intrahippocampal kainic acid model), and human data. Neuronal cultures treated with picrotoxin did not show evidence that FRs could be generated more than chance aggregation of action potentials. They then asked whether synchronization and firing rate could play a role in the emergence of FRs. They found that changes in neural firing and synchronization, such as those occurring during differences phase of the sleep-wake cycle could affect the number of FRs occurring by chance aggregation, with more FRs seen during periods of wakefulness, a result that they replicated in human data.

      The authors largely achieve their proposed aims of demonstrating that random neuronal firing can, in principle, generate FRs. Results from this study could influence current thinking around mechanisms generating FRs in epilepsy. The use of different computational approaches and model systems could offer new analytical methodologies for the study of FRs in the context of brain disease.

      Strengths:

      (1) The authors used a multi-level approach combining computational modeling with experimental datasets, including neuronal cultures, a rat model of temporal lobe epilepsy and human data.

      (2) Identification of key parameters such as neuronal count, firing rate, synchronization and brain state in observed incidence of FRs generated through random aggregation of neural firing.

      (3) Cross-species validation increases the likelihood of generalizability of the findings.

      Minor weakness:

      (1)The analyses conducted in human data lack direct comparison with sleep data due to no available data, but would encourage future investigations directly comparing HFOs during wakefulness and nocturnal sleep.

      Comments on revised version.

      The authors have addressed my comments and I have no further suggestions.

    5. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This is a study utilizing several types of analyses (computational modeling, neuronal cultures, rodent epilepsy model, and human intracranial multi-scale recordings) to address a highly relevant conceptual question: Are fast ripples (FRs) distinct pathological entities or largely emergent products of stochastic spike clustering? The results can potentially reshape current approaches to incorporating fast ripples into the epilepsy surgery evaluation.

      Strengths:

      The conceptualization of fast ripples as potentially arising by chance is highly novel and builds effectively on questions raised in prior studies that have never been satisfactorily resolved.

      The integration across biological scales and models is a major strength. The state dependency analysis provides additional, strong support. The methodology and statistical approaches used are thoughtfully presented and rigorously applied.

      In particular, this paper provides a strong response to the findings from Gliske et al, Nat Commun 2018. This study utilized long-term data analysis to uncover low rates of FRs detected from most recording sites, suggesting spurious detections, although FRs were concentrated within seizure onset areas.

      We fully agree with this comparison. Although we had already cited this paper, we now further emphasize this observation in the Discussion:

      “Furthermore, the variability of FRs across time (Gliske et al., 2018) indicates that longer nocturnal recordings in humans are necessary. It also suggests that changes in excitability across time could explain this change in FR incidence.”

      Weaknesses:

      The authors clearly aimed to use a statistical rather than a mechanism-based approach in this work. However, the paper's framing of true fast ripples as oscillatory events with stochastic fast ripples considered as confounders does not take prior investigations into biological mechanisms, particularly prior studies that point to an important role for stochastic fast ripples in some contexts. Incorporating recognition of these mechanisms would strengthen the manuscript and provide a more complete and nuanced characterization.

      Some examples from the literature:

      Eissa et al, eNeuro 2016, a paper that closely parallels this manuscript but took a mechanistic rather than statistical approach, showed that fast ripples can arise from population paroxysmal depolarizations - a key feature of epileptiform discharges - as temporally clustered, jittered population firing, with FRs appearing in LFP or EEG due to summated postsynaptic potentials (which are slower than action potentials and can generate signals in the high gamma range).

      Foffani et al., 2007, Neuron, and Ibarz et al., 2010, J Neurosci, argue that FRs are pseudo-oscillations created by jittered neuronal populations in the setting of altered spike timing.

      Smith et al., 2020, Sci Rep, contrasts FR characteristics in different regimes, i.e., intact inhibition early in a seizure vs. implied collapse of inhibition after recruitment. Schlingloff et al., 2025, J Neurosci, reported analogous findings in an animal model.

      We agree with the reviewer that even stochastic events may be of biological importance and an increase in stochastic events will occur when there is an increase in synchronisation and excitability, two properties of pathological cortex. We also don’t disagree that FRs can occur as distinct entities, although our work indicates that most are due to chance.

      To address this point, we have clarified our claims in the Abstract:

      “This work does not rule out FRs as potential indicators of epileptogenic tissue, but it does challenge prevailing assumptions about their generation and specificity. Their higher prevalence in epileptogenic tissue is likely primarily due to increased excitation and/or neural synchronization, rather than peculiar abnormalities in network behavior.”

      In addition, we expand on these points at various junctures in the Discussion. In particular, we reiterate our assertion that FRs may still be a useful biomarker, but that their interpretation should be moderated to reflect the fact that they often occur by chance:

      “Importantly, we do not question the potential of FRs to delineate the seizure-onset zone. Instead, our results suggest that the observed increase in FRs within the epileptogenic zone is an emergent phenomenon – arising due to changes in secondary network properties such as excitability and synchronization, not as a direct result of some pathology that is specific to epilepsy. In addition, we show that long-durations FRs are more likely to be distinct oscillations than stochastic events; and so FR duration is a key parameter that should be considered in future studies.”

      The computational model and subtraction approach provide a strong case for the random emergence of clustered activity in the high gamma band, given its assumptions. However, any such modeling effort needs to account for inhibitory activity, including impaired inhibitory function that is expected in epileptic brain regions, which has a strong modulating effect on excitatory firing and is thought to play a significant role in FR generation.

      We appreciate the reviewer’s concerns, but we believe that the impact of inhibitory interneuron activity on excitatory firing rates and synchronisation is incorporated indirectly into our simulations, while keeping our model as parsimonious as possible by not directly incorporating interneuron activity into our simulations. We have addressed this point in the Methods section:

      “Varying synchrony allowed us to test the impact, on the network, of inhibitory cells, which have been shown to favour synchrony (Bocchio et al., 2024; Cobb et al., 1995).”

      The shuffling procedure aims to preserve the power spectrum but randomizes high frequency phase (>200 Hz). However, this procedure removes biologically meaningful spike timing correlations, as well as structured cross-frequency coupling. The subtraction method thus likely underestimates the incidence of structured "distinct" FRs, while perhaps overestimating "chance" FRs due to biologically infeasible activity, making the statement that most FRs are due to chance correlation too strong.

      We appreciate this concern, which is especially important given that our results depend crucially on the validity of our shuffling procedure (as described in the Discussion). To address this issue, we have implemented an additional shuffling algorithm that preserves cross-frequency coupling (see last section of the Results, especially Supplementary Fig. 10f). This method showed no qualitative difference, compared with other alternative methods presented in Supplementary Fig. 10. These new results are described in the Methods section:

      “Last, we also implemented a method based on wavelet-IAAFT with preservation of cross-frequency coupling, since fast ripples are typically locked to low-frequency phase (Sheybani et al., 2019). The code detects the highest phase-amplitude coupling (PAC) in the original signal between [300-6000 Hz] for amplitude and several low-frequency bands ranging from 2-20 Hz, bandwidth of 3 Hz. PAC is computed using the modulation index (Tort et al., 2008). Then, in the shuffled signal under construction and during convergence testing of PSD (see above), the PAC between high-frequency part of the signal (300-6000 Hz) and the identified low frequency for phase is normalized to that of the highest PAC identified earlier.”

      The kainate findings underscore this point: the increase in the number of FR detections could be, as the authors state, an increase in chance clustering due to increased network excitability generally. However, the likelihood of a parallel increase in pathological FRs cannot be ruled out, given likely pro-epileptic alterations in spike timing and circuit function.

      We appreciate the reviewer’s point but wish to re-emphasise our interpretation of these findings – that the observed increase in the incidence of FRs occurs as a result of increased network excitability/synchrony, secondary to the pathological mechanisms of epilepsy. We have updated the Discussion accordingly:

      “Importantly, we do not question the potential of FRs to delineate the seizure-onset zone. Instead, our results suggest that the observed increase in FRs within the epileptogenic zone is an emergent phenomenon – arising due to changes in secondary network properties such as excitability and synchronization, not as a direct result of some pathology that is specific to epilepsy. In addition, we show that long-duration FRs are more likely to be distinct oscillations than stochastic events; and so FR duration is a key parameter that should be considered in future studies.”

      To further emphasise this important point, we have also updated the Abstract:

      “This work does not rule out FRs as potential indicators of epileptogenic tissue, but it does challenge prevailing assumptions about their generation and specificity. Their higher prevalence in epileptogenic tissue is likely primarily due to increased excitation and/or neural synchronization, rather than peculiar abnormalities in network behavior.”

      Reviewer #2 (Public review):

      Summary:

      This paper asks an important question that has not been discussed much in the extensive literature on the High Frequency Oscillations (HFOs) that have been extensively studied in patients with epilepsy and experimental models of epilepsy. The question is whether the Fast Ripples (FRs), the HFOs in the 250-500 Hz frequency band, represent a pathological phenomenon or represent a physiological phenomenon that occurs in the healthy brain but happens to be more frequent in epileptic tissue. It is an important question that has not been systematically addressed until now. The authors conclude, from very extensive simulations, from extensive experimental animal studies (the systemic kianate model of epilepsy in rats), and from a modest amount of human data, that FRs occur in healthy brains as a result of the chance occurrence of bursts of action potentials, and that in epileptic tissue, their frequency of occurrence is approximately 30% higher than what is expected by chance. They conclude that FRs are not a separate phenomenon of epileptic tissue. This finding is reinforced by the recent findings of FRs in experimental models of Alzheimer's disease.

      Strengths:

      This is a valuable study because it asks an important and original question and because it evaluates it from several angles (simulation, tissue culture, experimental animals, and human patients). The simulations and the analyses of real data are performed very carefully and with original and solidly documented approaches, using extensive simulations and extensive data sets in the cultured cell data and in the in vivo experiments. The paper is clearly written and well-illustrated.

      Weaknesses:

      I found only one serious weakness in this study, but it is one that is of importance. Although the original work on FRs was done in an experimental model of epilepsy, the field really became prominent when ripples and fast ripples were found first in microelectrode recordings of epileptic patients and then in the intracerebral EEG of such patients. Numerous studies have been performed since then, with a valuable meta-analysis including 700 patients (Wang Z, Guo J, van 't Klooster M, Hoogteijling S, Jacobs J, Zijlmans M. Prognostic Value of Complete Resection of the High-Frequency Oscillation Area in Intracranial EEG: A Systematic Review and Meta-Analysis. Neurology. 2024 May 14;102(9). Although the consensus at this point is that FRs are not the ideal and totally specific marker of epileptic tissue that many thought it could be, FRs are nevertheless much more frequent in epileptic tissue than in non-epileptic tissue and are a solid biomarker.

      We agree with the reviewer, and do not intend to challenge the role of FRs as a marker of the seizure-onset zone, and potentially the epileptogenic zone. Instead, the aim of this study was to address the question of whether FRs are generated by intrinsic pathological mechanisms, or whether they arise due to the chance co-occurrence of action potentials that follow different dynamics in epileptogenic parenchyma. We have updated the Discussion accordingly:

      “Importantly, we do not question the potential of FRs to delineate the seizure-onset zone. Instead, our results suggest that the observed increase in FRs within the epileptogenic zone is an emergent phenomenon – arising due to changes in secondary network properties such as excitability and synchronization, not as a direct result of some pathology that is specific to epilepsy. In addition, we show that long-durations FRs are more likely to be distinct oscillations than stochastic events; and so FR duration is a key parameter that should be considered in future studies.”

      To further emphasise this important point, we have also updated the Abstract:

      “This work does not rule out FRs as potential indicators of epileptogenic tissue, but it does challenge prevailing assumptions about their generation and specificity. Their higher prevalence in epileptogenic tissue is likely primarily due to increased excitation and/or neural synchronization, rather than peculiar abnormalities in network behavior.”

      It is also well established that they are much more frequent in NREM sleep than in wakefulness, as reported in the original paper of Staba et al (Staba RJ, Wilson CL, Bragin A, Jhung D, Fried I, Engel J Jr. High-frequency oscillations recorded in human medial temporal lobe during sleep. Ann Neurol. 2004 Jul;56(1):108-15., not mentioned in this paper) and in the study of Bagshaw et al (2009). In this last paper, using SEEG in various brain regions, the average rate of FRs in NREM sleep is about 6 times that in wakefulness. In the paper by Staba, with microelectrodes in mesial temporal structures, it is about twice. As a separate issue, the paper of Fraucher et al (Frauscher B, von Ellenrieder N, Zelmann R, Rogers C, Nguyen DK, Kahane P, Dubeau F, Gotman J. High-Frequency Oscillations in the Normal Human Brain. Ann Neurol. 2018 Sep;84(3):374-385), which is not quoted, found that, in an extensive sample, non-epileptic human tissue sampled with SEEG generated extremely rare FRs (an average rate of 0.04/min/channel, i.e. 1 every 25 min).

      The results above are mentioned because they do not fit with the data provided in the present study: FRs are much more frequent in NREM sleep than in wakefulness in human epileptic patients, and they are much more frequent (not 30% more, but many hundreds of percent more) in epileptic tissue than in non-epileptic human tissue. The fundamental phenomenon of interest is, I believe, the FRs in epileptic patients. The animal experiments, tissue studies, and simulations are models to study the human phenomenon. With respect to the modulation by sleep and the differentiation between epileptic and non-epileptic tissue, it seems that the systems studied in this paper are not good models of the human condition. The human results presented in the study only reflect wakefulness recordings, which is not the condition in which most HFO studies have been done and in which most HFOs occur. The authors refer to the study of long-term fluctuations in HFO rates by Gliske et al. (2018) to say that one has to be careful with the results regarding sleep, for example, Bagshaw et al (2009), but the clear predominance in of HFOs in NREM sleep has been observed by many studies. The cautions regarding fluctuations over extended periods also apply to the awake human data analyzed in this study. The study's conclusions regarding the generation of FRs are therefore questionably applicable to the human condition. I do not dispute their validity for the models and situations in which they were studied.

      We looked at this in more detail. Our simulations were intended to test how the incidence of FRs can vary with different parameters of network activity (neuronal count, firing rate, synchronization). Indeed, since their incidence is known to vary across regions and within regions and across states, we wanted to test how FRs are controlled by different factors. As such, we do not wish to draw firm conclusions about the observed sleep-wake changes in FR incidence in rodents, and how it relates to humans – evidence shows that pathological FRs in rodents do not display state-specific preferential occurrence (Ewell et al., 2019). We have added new text to the Abstract and Discussion to emphasize this.

      Abstract:

      “Our simulations showed that chance aggregation can generate fast-ripples and that their incidence changes depending on brain state, an observation that we confirmed in our rodent data.”

      We acknowledge that previous publications have reported higher rates during sleep, although with shorter recordings than in our rodent recordings (Staba, 2004: one night; Bagshaw, 2009: 10 min; Frauscher, 2018: 20 min – only sleep recordings). We have rewritten the part of the Discussion on the effect of the sleep-wake cycle on FRs incidence:

      “In our rodent data, we were initially surprised to find a higher rate of FRs during wakefulness, which contrasts with previous reports in humans (Bagshaw et al., 2009; Staba et al., 2004). However, previous studies only indicate that physiological vs pathological FRs are more easily distinguished during NREM sleep (von Ellenrieder et al., 2016) and that their incidence varies during sleep (Von Ellenrieder et al., 2017), but in hours-long recordings, no differences in incidence have been reported in the mesial temporal lobe (Dümpelmann et al., 2015). Furthermore, the variability of FRs across time (Gliske et al., 2018) indicates that longer nocturnal recordings in humans are necessary. It also suggests that changes in excitability across time could explain this change in FR incidence. Last, but not least, another report did not find a state-dependent expression of FRs in the kainate rat model of temporal lobe epilepsy (Ewell et al., 2019), thus indicating that the variability of FRs across sleep and wake is still an open question, at least in rodents. Hence, the main conclusion on the effect of sleep-wake transitions is that these transitions impact the likelihood of stochastic events, more than dictating the direction (increases vs decreases) of change. It also highlights that the specificity of FRs to epileptogenic parenchyma could vary across the sleep-wake cycle, which would be crucial in epileptology (Dimakopoulos et al., 2024; Roehri et al., 2018; Sheybani et al., 2019, 2018; Zijlmans et al., 2012, 2009). Hence, FRs reflect and are highly susceptible to changes in network excitability.”

      Reviewer #3 (Public review):

      Summary:

      An outstanding question in the field of high-frequency oscillations (HFOs) in the context of epilepsy is how these oscillations emerge, considering that they occur at such high frequencies, i.e., 250Hz, well above the firing ability of single neurons. One hypothesis that has been suggested in the past is that neurons that fire in an out-of-phase fashion, or rather at random intervals, may contribute to a spectrum of HFOs ranging from 250-500Hz that are observed in epilepsy. However, how possible it is that random action potentials could aggregate to the extent that they could give rise to HFOs in the so-called fast ripple (FRs) frequency range (>200 according to the authors) remains unclear. To test this hypothesis, they used computational modeling to randomly insert action potentials in a signal, and they found that this approach is sufficient to generate FRs. Some of the predictors of whether FRs could occur were neuronal count, firing rate, and synchronization. Besides computational modeling, they used different model systems to test whether that would be possible to be observed in neuronal cultures, in epileptic rats (intrahippocampal kainic acid model), and human data. Neuronal cultures treated with picrotoxin did not show evidence that FRs could be generated beyond chance aggregation of action potentials. They then asked whether synchronization and firing rate could play a role in the emergence of FRs. They found that changes in neural firing and synchronization, such as those occurring during differences phase of the sleep-wake cycle, could affect the number of FRs occurring by chance aggregation, with more FRs seen during periods of wakefulness, a result that they replicated in human data.

      The authors largely achieve their proposed aims of demonstrating that random neuronal firing can, in principle, generate FRs. Results from this study could influence current thinking around mechanisms generating FRs in epilepsy. The use of different computational approaches and model systems could offer new analytical methodologies for the study of FRs in the context of brain disease.

      Strengths:

      (1) The authors used a multi-level approach combining computational modeling with experimental datasets, including neuronal cultures, a rat model of temporal lobe epilepsy, and human data.

      (2) Identification of key parameters such as neuronal count, firing rate, synchronization, and brain state in observed incidence of FRs generated through random aggregation of neural firing.

      (3) Cross-species validation increases the likelihood of generalizability of the findings.

      Weaknesses:

      (1) Some of the simulated FRs appear short in duration and may not meet standard detection and definition criteria, potentially influencing validity.

      We thank the reviewer for raising this important concern. To address this issue, we quantified and compared the duration of FRs in original and shuffled rodent data. Consistent with the reviewer’s suspicions, we found that FRs in shuffled signals are shorter than FRs in original signals. This is important because it shows that: (i) a longer duration should be considered a core feature of genuine FRs; and (ii) depending on the basal duration of FRs, the shuffling procedure will lead to different ratios of genuine to stochastic FRs. We have updated the Results accordingly:

      “These findings demonstrate the challenge of identifying distinct FRs within a composite population of distinct and stochastic events. One parameter that could help disentangle these events is their duration. Indeed, one might expect stochastic events to be more likely to be short-lived, since the probability of consecutive APs continuing to co-occur across neurons decreases over time. Hence, we next compared the distribution of FR durations between original and shuffled rodent data and found that FRs in shuffled data are shorter than those in original data (Supplementary Fig. 9). This makes duration a key feature that could help identify distinctly generated FRs.”

      And Discussion accordingly:

      “In addition, we show that long durations FRs are more likely to be distinct oscillations than stochastic events; and so FR duration is a key parameter that should be considered in future studies.”

      (2) The neuronal culture approach does not directly test random insertion of action potentials, limiting interpretation.

      Neither the neuronal culture approach, the rat data or the human data directly test random insertion of action potentials. The insertion of random action potentials is only performed in the simulated data to test if FRs can arise from the chance insertion of action potentials. Once this was confirmed in the simulations, we then used the shuffling procedure in biological data to test if FRs are more frequent than expected by chance.

      (3) Sleep is treated as a homogeneous state in the rat dataset, without accounting for stage-specific differences in synchronization, which may affect the results and interpretation.

      We agree with the reviewer, but our primary aim was to answer the question of whether FRs can arise by chance. Although it was interesting to see that, in our longitudinal rodent data, the incidence of FRs varies across the sleep-wake cycle, any sleep-stage-specific changes are beyond the scope of this work.

      (4) The analyses conducted in human data lack direct comparison with sleep data.

      We agree that it would have been useful to investigate variations in the incidence of FRs across the sleep-wake cycle in human microelectrode recordings. Unfortunately, however, such sleep recordings were not available. Hence, while we cannot compare variations in FR incidence across brain states between humans and animal models, our conclusions that FRs arise mostly by the chance co-occurrence of action potentials still holds.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) Please indicate where corrections for multiple comparisons were used.

      P-values corrected for multiple comparisons are indicated by the accompanying phrase: “adjusted p-value”. We had previously omitted to mention this once in the Results, which we have now corrected.

      (2) Delta amplitude is likely sufficient for detecting sleep-wake transitions, but the beta/delta ratio is better supported in the literature. Do the results change if beta activity is incorporated?

      We have now computed the beta (15-40 Hz) to delta (0.5-4 Hz) ratio and find that this is closely correlated with delta across time. We have updated the Results accordingly:

      “Importantly, these findings were robust to the specific method used to detect FRs (Supplementary Fig. 4c) (Padmasola et al., 2024; Sheybani et al., 2019, 2018). Also our use of delta power to identify periods of presumed wakefulness and sleep was highly (negatively) correlated with an alternative method of using the beta-to-delta ratio across time (another marker of increased vigilance; (Fraigne et al., 2023), see Supplementary Fig. 4d).”

      Methods:

      “We further verified that delta power across time displayed similar fluctuations to beta (15-40 Hz) to delta power ratio, another marker of vigilance (Fraigne et al., 2023).”

      And we updated Supplementary Fig. 4d

      “(d) Beta to delta power ratio across time is superimposed over delta power across time. There is a strong (inverse) correlation between the two time-series (inset), which is confirmed by the correlation coefficient across animals (right).”

      (3) Figure 2's axis labeling with the 3D plots is hard to read.

      We have enlarged the font size.

      (4) The scaling of the histogram in Figure 3 is unclear.

      This was on omission. The scale has now been added to Figure 3.

      (5) There is a risk of overfitting in the regression model. Was cross-validation used?

      We have now repeated this analysis with cross-validation, without any qualitative impact on the results (e.g. the model still performs well above chance). We have updated the Methods:

      “To further confirm the performance of GBT, we used a cross-validation procedure where the GBT is trained on 80% of data and then tested on the 20% remaining. The procedure is repeated 1000 times and the r<sup>2</sup> is saved at each round. We repeated the analysis with randomization of the outputs across 1000 rounds and saved this null distribution r<sup>2</sup>. We then compared the performance against original data.”

      Legend of Fig. 3:

      “(e) Performance of the GBT classifier using cross-validation (training: 80% of data; test: 20% remaining) using original (orange) and shuffled (blue) data. The difference is significant (paired t-test, p<0.0001).”

      And Results:

      “Furthermore, using a cross-validation approach with 80% of the data as training set and the remaining 20% as the test set, we obtained a significantly higher explained variance than when outputs were shuffled across the 125,000 solution points (paired t-test, p<0.0001, Fig. 3e), […]”

      Reviewer #2 (Recommendations for the authors):

      Maybe I missed it, but I did not find the length of human data analyzed or how the sections were selected.

      Apologies for this omission. The methods have been updated accordingly:

      “Microwire signals were selected based on high signal-to-noise ratio, as reflected by the detection of ≥ 1 single unit. Duration of recordings was of (median, interquartile range) 10 min and 17 s [3-13 min] and number of electrodes per patient was 4.5 [2.75-8].”

      The authors use the term "virtual simulation", which I find odd. I think the simulation is very real in the sense that it simulates reality, and I do not understand how a simulation can be virtual.

      We have updated the manuscript accordingly.

      Reviewer #3 (Recommendations for the authors):

      Major Comments:

      (1) In Figure 1, the authors suggest that random insertion of action potentials in a signal is sufficient to yield FRs. However, the observed FRs shown in panel 1b (also in supplemental Figure 5) seem pretty short in duration and may not meet the mentioned criteria in methods that require at least 4 cycles and ".whose amplitude is 3 times that of the surrounding baseline..". Moreover, in panel 1b, it seems that the FR shows a candle-like appearance, which has often been associated with filtering of sharp transients. How did the authors validate that the detected FRs were "real" FRs?

      Given the very large amount of data, it was not possible to visually verify all FRs. However, FRs were detected with published methods (Roehri et al., 2016; Roehri et al., 2017; and Sheybani et al., 2018 for confirmation of 24-hour variability in rodents) that have subsequently been used in several publications.

      Regarding the candle-like appearance of the spectrogram, the Delphos algorithm precisely looks for isolated “islands” of increased power (see Roehri et al., 2018, Ann Neurol), thus excluding any candle-like appearance. Similarly, the detector in Sheybani et al. (2018) J Neurosci first detects candidate FRs but then excludes those that are associated with a peak in lower frequencies, thus also limiting the risk of detecting candle-like events.

      Regarding duration, we have compared the duration of FRs in original and shuffled rodent data and found that FRs in original signals are indeed longer. This makes duration a key feature to identify distinct FRs. We have updated the Results accordingly:

      “These findings demonstrate the challenge of identifying distinct FRs within a composite population of distinct and stochastic events. One parameter that could help disentangle these events is their duration. Indeed, one might expect stochastic events to be more likely to be short-lived, since the probability of consecutive APs continuing to co-occur across neurons decreases over time. Hence, we next compared the distribution of FR durations between original and shuffled rodent data and found that FRs in shuffled data are shorter than those in original data (Supplementary Fig. 9). This makes duration a key feature that could help identify distinctly generated FRs.”

      (2) In the context of neuronal cultures, it is unclear how it could be deducted that the result relates to chance incidence of action potentials considering that no random action potentials were inserted, but only random shuffling of the high frequency component of the signal was attempted "Hence, neural networks with limited complexity (Kim et al., 2020; Saglam-Metiner et al., 2024; Sanchez-Vives and McCormick, 2000; Timofeev and Chauvette) fail to generate FRs beyond that expected from the chance coincidence of APs, even after increasing network excitability."

      FRs arise from series of action potentials occurring at a delay corresponding to their oscillatory frequency (250-500 Hz). Simulations demonstrated that FRs can occur by chance. When the EEG is shuffled, the only FRs that remain are those occurring by chance, because those occurring as individual entities have been broken up. Hence, if the original EEG displays more FRs than the shuffled EEG, then it means that these additional FRs were generated as individual entities. We have improved the Results section to clarify this:

      “We hypothesized that if FRs arise purely from chance firing, then temporally shuffling these recordings while conserving their spectral properties (Supplementary Fig. 3) would disrupt any oscillatory structure, leaving only FRs that occur due to chance.] Any additional FRs in the original data, compared to the number of FRs in the shuffled EEG, should thus be assumed to be individual entities.”

      (3) In the rat dataset, sleep was treated rather homogenously, without accounting for the sleep stage that is characterized by different synchronization and firing. An analysis of different sleep stages would be valuable.

      Although we agree that it would be scientifically interesting, we believe that our claim – that the ratio of genuine to stochastic FRs changes across the sleep-wake cycle – would hold. Unfortunately, lack of EMG prevents us from performing reliable sleep scoring. However, we do now include an alternative method for differentiating sleep from wake using the beta-to-delta ratio, which was highly correlated with delta activity, supporting our previous approach. Please refer to Supplementary Fig. 4d for further information.

      (4) The authors found that chance aggregation was highest during periods of wakefulness. Analyses of human data also confirmed that FRs could occur by chance aggregation during wakefulness. However, a comparison with sleep data would further strengthen this finding.

      We fully agree, but unfortunately, we do not have sleep data using microwires. Although our central claim – that FRs can occur by chance clustering of action potentials – would hold, we agree that it would have been scientifically interesting to add sleep data.

      (5) The statistics section would benefit from addressing how normality was determined and power analysis, as well as the inclusion of the exact sample size for all experiments.

      With large sample sizes, ANOVA and linear mixed models are robust to non-normality. Given the large sample sizes of our data, we thus used ANOVA and linear mixed model. For tests with small sample sizes where normality was violated, we used non-parametric tests, indicated by their name, e.g., Wilcoxon test for Supplementary Fig. 3b.

      (6) Greater discussion on the implications of this study for proposed in-phase or out-of-phase FR generation mechanisms is suggested.

      We have added further discussion on this. In the aim to keep the Discussion short and impactful, we could not elaborate too much. We have synthetized other parts of the Discussion to keep it within the right length. Here is the additional part:

      “It has been argued that the very high frequency that can be obtained during FRs are due to out-of-phase firing of excitatory neurons (Foffani et al., 2007; Ibarz et al., 2010), which is also consistent with our concept of stochastic firing. The conceptual difference is the degree to which there is any underlying organization of this firing. We argue that in the majority of cases there is no organization, although a substantial minority cannot be explained on a stochastic basis.”

      (7) More explanation around why wakefulness may drive chance aggregation and the clinical relevance of it, as often presurgical epilepsy recordings are being evaluated during sleep.

      We have profoundly rewritten the Discussion regarding the effect of the sleep-wake cycle on FRs incidence:

      “In our rodent data, we were initially surprised to find a higher rate of FRs during wakefulness, which contrasts with previous reports in humans (Bagshaw et al., 2009; Staba et al., 2004). However, previous studies only indicate that physiological vs pathological FRs are more easily distinguished during NREM sleep (von Ellenrieder et al., 2016) and that their incidence varies during sleep (Von Ellenrieder et al., 2017), but in hours-long recordings, no differences in incidence have been reported in the mesial temporal lobe (Dümpelmann et al., 2015). Furthermore, the variability of FRs across time (Gliske et al., 2018) indicates that longer nocturnal recordings in humans are necessary. It also suggests that changes in excitability across time could explain this change in FR incidence. Last, but not least, another report did not find a state-dependent expression of FRs in the kainate rat model of temporal lobe epilepsy (Ewell et al., 2019), thus indicating that the variability of FRs across sleep and wake is still an open question, at least in rodents. Hence, the main conclusion on the effect of sleep-wake transitions is that these transitions impact the likelihood of stochastic events, more than dictating the direction (increases vs decreases) of change. It also highlights that the specificity of FRs to epileptogenic parenchyma could vary across the sleep-wake cycle, which would be crucial in epileptology (Dimakopoulos et al., 2024; Roehri et al., 2018; Sheybani et al., 2019, 2018; Zijlmans et al., 2012, 2009).”

      Minor Comments:

      (1) Abstract, please include the frequency range of fast ripples explored in this study.

      The abstract has been updated accordingly.

      (2) Abstract, consider including the exact epilepsy model system in rats instead of "a rodent model of hippocampal epilepsy".

      The abstract has been updated accordingly.

      (3) Line 87, while Ylinen uses the term "high frequency oscillations" to refer to ripples up to 200Hz, which are different from the ones discussed here, better to rephrase or use another reference.

      The reference has been changed for Bragin et al. (1999), Epilepsia

    1. eLife Assessment

      This valuable study demonstrates molecular changes associated with age related impairment in oligodendrocyte differentiation and ability to myelinate. The identification of particular genes that are associated with this decline will provide potential future targets for therapeutic interventions. The reviewers felt that the quality of the evidence was convincing while identifying some minor weaknesses that were largely addressed in the review process.

    2. Reviewer #1 (Public review):

      Summary:

      This manuscript by Ghosh and colleagues investigates the transcriptional changes within the oligodendrocyte lineage that contribute to age-related declines in oligodendrocyte differentiation and myelination. Combining bulk RNA-Seq on acutely purified oligodendrocyte lineage cells with bioinformatic approaches, the authors identify groups of genes that show different patterns of dynamic regulation during differentiation (which they term "switch" genes, or "switches"). A subset of these switch genes are differentially regulated with age. The authors identify two transcription factors, Bcl11a and Foxm1 that are downregulated during differentiation, have predicted binding site enrichment at other switch genes and are downregulated in aged OPCs. Functionally testing Bcl11a, the authors show that Bcl11a knockdown inhibits the differentiation of young OPCs in culture, whereas overexpression promotes differentiation of aged OPCs. Viral expression of Bcl11a in Sox10 expressing cells accelerates the formation of Plp1+ oligodendrocytes in aged rodents following lysolecithin induced demyelination.

      Strengths:

      The work is clearly presented and addresses an important biological problem. The bioinformatic approaches used in the manuscript are powerful, and the identification of Bcl11a as a modulator of oligodendrocyte differentiation is a novel finding. The combined in vitro and in vivo approaches to assess the function of Bcl11a in oligodendrocyte differentiation are a substantial strength of the work.

      Comment on revised version.

      In the revised version the authors now provide analysis of expression of stage-specific markers for OPCs, preOls and OLs in their isolated cells. It is slightly concerning that the OPC markers show higher expression in the isolated preOLs than in the isolated OPCs, but the authors do provide some discussion on this point in the supplementary text.

    3. Reviewer #2 (Public review):

      Ageing poses a significant challenge to the regenerative capacity of oligodendrocyte precursor cells (OPCs). Myelin abnormalities accumulate with age, while the ability of OPCs to differentiate into myelinating oligodendrocytes progressively declines. This likely contributes to inefficient replacement of damaged myelin and oligodendrocytes, impaired remyelination following injury, and reduced adaptive myelination. Identifying the molecular changes associated with this decline is therefore important for understanding and potentially treating age-related deterioration of CNS white matter.

      This study sought to identify transcriptional regulators involved in oligodendrocyte-lineage progression whose expression is altered in aged OPCs. The authors developed gSWITCH, a computational tool that identifies genes showing defined dynamic expression patterns across ordered biological states. By combining this analysis with comparisons of young and aged OPC transcriptomes and transcription-factor-binding-site enrichment, they identified Bcl11a as a candidate regulator. Bcl11a transcripts are abundant in young OPCs, decline during oligodendrocyte differentiation, and are markedly reduced in aged OPCs.

      A major strength of the study is its combination of computational candidate identification with functional experiments. Bcl11a knockdown substantially impaired the differentiation of young OPCs without measurably affecting their proliferation. Conversely, transient Bcl11a overexpression increased the differentiation of aged OPCs in vitro. Oligodendrocyte-lineage-specific expression of Bcl11a in aged mice also increased the generation of PLP1-positive oligodendrocytes following focal demyelinating injury. Together, these complementary loss- and gain-of-function experiments support the conclusion that Bcl11a expression is functionally important for OPC differentiation and that restoring its expression can improve the differentiation competence of aged OPCs.

      While the transcription-factor-binding-site enrichment analysis predicts a Bcl11a-regulated network, the current study does not establish direct binding or identify the downstream genes responsible for its effect on OPC differentiation. Similarly, the upstream mechanisms responsible for the age-associated reduction in Bcl11a expression were not investigated. Further work may help establish a more complete mechanistic framework explaining how restoration of Bcl11a expression improves OPC differentiation.

      Overall, this study offers valuable insights into the age-related loss of regenerative capacity in the central nervous system and introduces a computational framework that may be broadly useful for investigating dynamic gene regulation in other biological contexts.

      Comments on revised version.

      The authors have addressed my previous comments, and the revised manuscript has been substantially strengthened by the inclusion of additional supporting data and an expanded discussion.

    4. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This manuscript by Ghosh and colleagues investigates the transcriptional changes within the oligodendrocyte lineage that contribute to age-related declines in oligodendrocyte differentiation and myelination. Combining bulk RNA-Seq on acutely purified oligodendrocyte lineage cells with bioinformatic approaches, the authors identify groups of genes that show different patterns of dynamic regulation during differentiation (which they term "switch" genes, or "switches"). A subset of these switch genes is differentially regulated with age. The authors identify two transcription factors, Bcl11a and Foxm1, that are downregulated during differentiation, have predicted binding site enrichment at other switch genes, and are downregulated in aged OPCs. Functionally testing Bcl11a, the authors show that Bcl11a knockdown inhibits the differentiation of young OPCs in culture, whereas overexpression promotes the differentiation of aged OPCs. Viral expression of Bcl11a in Sox10-expressing cells accelerates the formation of Plp1+ oligodendrocytes in aged rodents following lysolecithin induced demyelination.

      Strengths:

      The work is clearly presented and addresses an important biological problem. The bioinformatic approaches used in the manuscript are powerful, and the identification of Bcl11a as a modulator of oligodendrocyte differentiation is a novel finding. The combined in vitro and in vivo approaches to assess the function of Bcl11a in oligodendrocyte differentiation are a substantial strength of the work.

      We sincerely thank the reviewer for their positive assessment and for recognising the significance of our study, as well as the bioinformatics approach and tool developed as part of this work.

      Weaknesses:

      Although the PCA plots show distinct and reproducible global gene expression differences between the different isolated cell populations, the authors do not present a figure showing expression levels of typical stage-specific markers (e.g., Pdgfra, Pcdh15, C1ql1 for OPCs, Bcas1, Enpp6, Gpr17 for preOLs, Mobp, Mog, etc. for OLs) or confirm the absence of markers of other lineages (astrocytes, neurons, microglia, etc.). This makes it difficult to evaluate the success of their cell isolation strategy at different ages without reanalyzing the raw data.

      Thank you for this suggestion. We have presented markers expression in a new figure (Supplementary Figure 1) and included a description in the new Supplementary text.

      We observed elevated expression of Hes1 in OPCs as compared to both PreOL and OL, consistent with its role as a Notch effector that maintains the OPC progenitor state and inhibits oligodendrocyte maturation (PMID: 19104146, PMID: 21167918).

      Compared with PreOLs, adult OPCs isolated from 2–3-month-old rats did not show higher RNA expression of canonical OPC markers: Pdgfra, Pcdh15, and C1ql1. However, as expected, OPCs expressed higher levels of these markers than mature OLs.

      One possible explanation is the intrinsic heterogeneity of adult OPC populations. Adult OPCs exist in multiple transcriptional states, including quiescent-like and differentiation-primed states. During early differentiation, OPC markers such as Pdgfra are not immediately extinguished, and PreOLs may transiently retain these transcripts. The PreOL population captured in our study represents intermediate states transitioning from OPC to OL, potentially still carrying residual OPC-associated RNAs from activated OPCs. Therefore, comparing PreOLs with the total heterogeneous OPC pool, which includes quiescent-like OPCs, may give the appearance of higher canonical OPC marker expression in PreOLs.

      Among the PreOL-specific markers, Gpr17 clearly distinguished the PreOL state in our data, showing higher expression compared with both OPCs and OLs. Bcas1 and Enpp6 showed higher expression in PreOLs compared with OPCs. However, when PreOLs were compared with OLs, Bcas1 appeared to be lower in PreOLs, whereas Enpp6 expression remained largely unchanged.

      The OL markers Mobp and Mog showed significantly higher expression in OLs compared with OPCs, whereas their expression was not altered between OPCs and PreOLs. However, the canonical OL maturity marker Mbp showed a progressive and significant increase during differentiation, with expression levels clearly following the expected pattern OL > PreOL > OPC.

      We did not find any difference of astrocytes marker Gfap in those cell types comparison, suggesting similar level of unavoidable contamination which will not affect determination of differential gene expression. Regarding this please also see reviewer #2 major point 1.

      We now included this in the supplementary text:

      “Please see Supplementary Figure 1. We observed elevated expression of Hes1 in OPCs compared with both PreOLs and OLs, consistent with its role as a Notch effector that maintains the OPC progenitor state and inhibits oligodendrocyte maturation (Brosnan et al, 2009; Ogata et al., 2011).

      Compared with PreOLs, adult OPCs isolated from 2–3-month-old rats did not show higher RNA expression of canonical OPC markers: Pdgfra, Pcdh15, and C1ql1. However, as expected, OPCs expressed higher levels of these markers than mature OLs. One possible explanation is the intrinsic heterogeneity of adult OPC populations. Adult OPCs exist in multiple transcriptional states, including quiescent-like and differentiation-primed states. During early differentiation, OPC markers such as Pdgfra may not be immediately extinguished, and PreOLs may transiently retain these transcripts. The PreOL population captured in our study represents intermediate states transitioning from OPCs to OLs, potentially still carrying residual OPC-associated RNAs from activated OPCs. Therefore, comparison of PreOLs with the total heterogeneous OPC pool, which includes quiescent-like OPCs, may give the appearance of higher canonical OPC marker expression in PreOLs.

      Among the PreOL-specific markers, Gpr17 clearly distinguished the PreOL state in our data, showing higher expression compared with both OPCs and OLs. Bcas1 and Enpp6 showed higher expression in PreOLs compared with OPCs. However, when PreOLs were compared with OLs, Bcas1 appeared lower in PreOLs, whereas Enpp6 expression remained largely unchanged.

      The OL markers Mobp and Mog showed significantly higher expression in OLs compared with OPCs, whereas their expression was not altered between OPCs and PreOLs. In contrast, the canonical OL maturity marker Mbp showed a progressive and significant increase during differentiation, with expression levels clearly following the expected pattern: OL > PreOL > OPC.

      We did not detect any difference in the astrocyte marker Gfap across these cell-type comparisons, suggesting a similar level of unavoidable astrocytic contamination across groups. Therefore, such contamination is unlikely to confound the interpretation of differential gene expression among OPCs, PreOLs and OLs.”

      In the main text we have added the following text:

      “The expression patterns of cell-type-specific markers were consistent with their being distinct OPC, Pre-OL, and OL populations (Supplementary Figure 1, see Supplementary text for detailed description).”

      Please note that a detailed discussion of marker expression in the main text will disrupt the flow of the manuscript in manner we feel would detract from its clarity. We have therefore provided this discussion in the Supplementary Text.

      In addition, other publicly available datasets (e.g., the Barres lab bulk RNA-Seq datasets from PMID 25186741 or the Castelo-Branco lab single cell datasets from PMID 27284195) do not show downregulation of Bcl11a during OL differentiation as is described here - this apparent discrepancy is not discussed.

      Thank you for raising this point. We have now included new data as a Supplementary Figure 4. We performed RT-qPCR (reverse transcription followed by qPCR) to quantify Bcl11a expression and found that it was significantly lower in OLs than in OPCs, and significantly lower in aged OPCs than in young OPCs. These data were presented together with stage-specific markers.

      Regarding the comparison with PMID: 25186741: we extracted Bcl11a FPKM values from their dataset (GSE52564) and plotted, as shown in Author response image 1. We found that Bcl11a expression is downregulated during differentiation. However, the dataset contains only two replicates, and the SEM between the two OL replicates is very high, which may have contributed to the apparent lack of clarity. With such high SEM and only two replicates, the statistical power is poor, making robust statistical inference difficult.

      Author response image 1.

      Plotting of FPKM values of Bcl11a (obtained from GSE52564). mean+SEM shown along with individual data points. OPC: Oligodendrocytes progenitor cells, NFO: Newly formed oligodendrocytes, MO: myelinating oligodendrocytes. Dotted red line: linear regression line.

      Regarding comparison with PMID 27284195: we contacted the Castelo-Branco laboratory, and they kindly provided us with the analysis shown below in Author response table 1. This analysis showed that Bcl11a expression is lower in myelinating oligodendrocytes (MOLs) compared with OPCs. The apparent discrepancy observed in the web interface is likely because MOLs are displayed separately by subtype in the online resource. In single-cell datasets, particularly earlier pre-10x datasets with relatively lower cell numbers and sparser transcript detection, visualisations such as violin plots or t-SNE plots can be difficult to interpret when expression is distributed across multiple subclusters. Therefore, directly examining the differential expression statistics, including fold-change and significance values, provides a clearer and more quantitative assessment of the expression change.

      Author response table 1.

      Bcl11a expression difference in MOLs vs OPCs (dataset: GSE75330)

      FC: fold change, p_val_adj: adjusted p-value.

      Therefore, our bulk RNA-seq and RT-qPCR analyses presented in this paper are consistent with the Barres laboratory bulk RNA-seq dataset (PMID: 25186741) and the Castelo-Branco laboratory scRNAseq dataset (PMID: 27284195).

      Reviewer #2 (Public review):

      Aging poses a significant challenge to the regenerative capacity of oligodendrocyte precursor cells (OPCs) to differentiate and myelinate neuronal axons. Myelin abnormalities accumulate with age, and it is likely that the ability of OPCs to differentiate into myelinating oligodendrocytes becomes progressively impaired during aging, leading to inefficient turnover of damaged myelin and oligodendrocytes, as well as reduced adaptive myelination. Understanding the molecular mechanisms underlying the compromised capacity of aged OPCs is therefore critical for addressing age-related white matter decline.

      This study aims to decipher the intrinsic molecular changes that occur in aged OPCs. By profiling differentially expressed transcription factors (TFs) between young and aged OPCs, and by employing a novel bioinformatic tool to identify key TFs that undergo dynamic changes across distinct stages of OPC differentiation, the authors identify Bcl11a as a potential regulator. Bcl11a is highly expressed in young OPCs but markedly reduced in aged cells. Functional experiments further demonstrate that while Bcl11a does not affect OPC proliferation, it significantly promotes the differentiation of aged OPCs. Importantly, this effect is also observed in vivo following demyelinating injury in aged mice.

      While the study provides compelling evidence that BCL11A represents a limiting factor for OPC differentiation during ageing, the downstream targets and molecular mechanisms through which BCL11A exerts its effects are not directly addressed. As such, the work should be interpreted primarily as identifying a key regulatory node rather than a fully defined molecular pathway.

      Overall, this study offers valuable insights into the age-related loss of regenerative capacity in the central nervous system and introduces a computational framework that may be broadly useful for investigating dynamic gene regulation in other biological contexts.

      We are grateful to the reviewer for their supportive comments and for highlighting the broader relevance of our computational framework beyond our specific subfield.

      Major Points:

      (1) MACS mouse anti-A2B5 microbeads are not OPC-specific and may also label astrocyte precursor cells or immature astrocytes. How do the authors justify this caveat? Could some of the claimed "OPCspecific" switch genes in fact be enriched in astrocyte lineage cells?

      We thank the reviewer for raising this important point. While anti-A2B5 is a well-established and widely used antibody for isolating OPCs, we nonetheless agree that no technique can isolate a specific cell type with 100% purity, and this also applies to OPC-specific isolation using a validated anti-A2B5 antibody.

      To check whether astrocyte contamination could be an issue in determining differential expression, and specifically whether the OPC population was affected by astrocyte contamination, we checked the relative expression and statistical significance of the astrocyte marker Gfap. We refer to our new Supplementary Figure 1 and Supplementary text. This suggests that no difference exists in Gfap levels when comparing OPC, PreOL and OL populations. Therefore, we contend that it is unlikely that the differential expression observed in any cell population is actually due to astrocyte contamination, or that the OPC population is selectively contaminated by astrocytes.

      We now included the following in the Supplementary text:

      “We did not detect any difference in the astrocyte marker Gfap across these cell-type comparisons, suggesting a similar level of unavoidable astrocytic contamination across groups. Therefore, such contamination is unlikely to confound the interpretation of differential gene expression among OPCs, PreOLs and OLs.”

      (2) Overall, Figures 1 and 2 are not very informative in terms of biological insight. The authors should provide more detail in the main figures regarding the enriched gene sets associated with each of the Type 1-4 switch categories. For example, summarizing the top Gene Ontology terms for each switch type would greatly enhance interpretability.

      We agree that GO analysis can add further interpretability. We have now prepared a new Supplementary Figure 3A to summarise the significant top GO-term enrichment for switch Types 1– 4, for which gSWITCH-identified patterns are presented in Figure 1C. We also prepared a Supplementary Figure 3B to summarise the top significant GO-term enrichment for the 135 Type 3 switch genes affected in ageing, presented in Figure 2C. Please note that only 8 Type 4 genes overlapped with differentially expressed genes in ageing. Due to this small number, we could not identify any significant GO-term enrichment, and therefore this was not plotted.

      (3) A similar issue applies to Figure 3. The authors should explicitly specify the transcription factors in the main figure, particularly the 27 TFs identified through theENCODE/ReMap2 analysis.

      Thank you for raising this point. We have now prepared a new Supplementary Table 3, where we list 27 TFs and highlight, with light grey shading, the 5 TFs that overlapped with Type 3 switches.

      (4) Have the authors validated Bcl11a expression across different CNS cell types and between young and aged conditions using independent methods such as qPCR, immunofluorescence, or western blotting?

      Thank you for this suggestion. We performed qPCR and presented this data in a new Supplementary Figure 4. We found that Bcl11a expression is lower in OLs than in OPCs (Supplementary Figure 4A). We also observed reduced Bcl11a expression in aged OPCs compared with young OPCs (Supplementary Figure 4B). (see also response to Reviewer 1’s recommendations).

      (5) Regarding OPC aging, an open question is whether the reduced differentiation capacity of aged OPCs is an intrinsic property of the cells themselves or whether it results from prolonged exposure to an aging environment that induces non-cell-autonomous epigenetic or genetic changes, thereby rendering OPCs less efficient at differentiating. It would be helpful if the authors could expand on this point in the Discussion, with reference to relevant previous studies and experimental evidence.

      We thank the reviewer for suggesting this important aspect be discussed. We have now included the following paragraph in the discussion section:

      “The extent to which the reduced differentiation capacity of aged OPCs is intrinsically encoded within the cells themselves or induced by prolonged exposure to an aged tissue environment is an interesting question. Based on our previous work, we favour the view that loss of OPC function is primarily determined extrinsically since various manipulations of the aged environment such as heterochronic parabiosis (Ruckh et al., 2012), fasting and calorie restriction mimetics (Neumann et al., 2019), and niche biomechanics (Segel et al., 2019) can all alter the cell-intrinsic state, reverting aged cells to a ‘youthful state’. Significantly, when aged OPCs are transplanted into the neonatal CNS they proliferate and differentiate as if they were neonatal OPCs (Segel et al., 2019). The reversion of aged OPCs to a functional state by changes in their external environment necessarily operates through changes in cell intrinsic function, suggesting that the same intrinsic mechanisms could be targeted directly to restore declining OPC function—for example through epigenetic regulation of differentiation inhibitors (Shen et al., 2008) or overexpression of transcriptional regulators such as c-Myc (Neumann et al., 2021, Dimas et al., 2025).”

      (6) Do the authors observe a change in the number or density of OPCs between young and aged mice?

      Thank you for asking this important question. In 2002 we reported that there was no difference in the OPCS density between young adult and old adult rats, at least in the deep cerebellar white matter (Sim et al. 2002 - PMID: 11923409). We also refer the reviewer to Figure S1 of another previous study, published in Cell Stem Cell in 2019 (PMID: 31585093). We did not find any difference in OPC number between young and aged brains. Quantification was performed using FACS, where freshly isolated cells were stained with A2B5 (OPC marker), CD11b (microglia marker), and MOG (oligodendrocyte marker). Thus, we do not find any evidence for an age-related decline in OPC densities.

      (7) The in vivo characterization of Bcl11a overexpression using the AAV-based approach appears incomplete. Do aged mice overexpressing Bcl11a in Sox10⁺ cells exhibit reduced age-related myelin degeneration under baseline conditions? In the LPC model, do the authors observe differences in lesion size and/or remyelination efficiency?

      Again, we thank the reviewer for raising these interesting points. To assess whether Bcl11a overexpression in Sox10+ myelinating oligodendrocytes exhibit less age-related myelin degeneration would, we suspect, require long-term experiments. For this to be the case would require a role for Bcl11a in myelin maintenance – and interesting question but one we feel (and hope the reviewer agrees) is beyond the scope of the current study. We do not see any difference in lesion size (and would not expect the expression of elevated levels of Bcl11a to protect against the membrane-solubilising effects of LPC) but do see changes in remyelination efficiency as shown in Figure 6.

      (8) Are the authors presenting gSWITCH for the first time in this manuscript? Given that the gSWITCH framework is novel and central to the study, its conceptual contribution could be emphasized more strongly. A brief comparison with existing trajectory- or pattern-based methods-ideally in the main text around Figure 1-would help readers better appreciate its novelty.

      We thank the reviewer for this important suggestion. Yes, gSWITCH is presented for the first time in this manuscript as a new computational framework and web application. We agree that its conceptual contribution should be made clearer in the main text itself, although we explained its concept in detail in ‘Materials and Methods’ and in the supplementary Figure 2 (which was Supplementary Figure 1 in first version of this manuscript).

      We now included the following paragraph in the manuscript:

      “Existing computational tools such as Monocle (Trapnell et al., 2014), tradeSeq (Van den Berge et al., 2020) and maSigPro (Nueda et al., 2014) are highly valuable for identifying genes with dynamic expression changes across pseudotime or time-course data. gSWITCH addresses a different question. It does not aim to infer trajectories. It works with a user-defined ordered series of biological states or time points and asks a more specific question — does this gene show a statistically supported "switchlike" change in expression as cells move through these states, and if so, what shape does that change take? It combines GLM-based statistical testing with criteria that capture where a gene reaches its highest or lowest expression and whether its expression changes steadily in one direction across the ordered series. To our knowledge, no existing tool combines significance testing with this type of explicit, shape-based classification into discrete, interpretable switch categories. gSWITCH sorts genes into four biologically meaningful patterns, rather than producing only a ranked list of significant genes based on pairwise comparisons between multiple conditions or states. gSWITCH also flags which of these switch genes are transcription factors, making it easier to prioritise candidates for follow-up experiments.

      This biologist-friendly tool is freely available as a web application requiring no programming, works with experimental designs containing three or more stages or time points with at least two replicates per stage (no upper limit on either), and can be applied to bulk RNA-seq or to single-cell RNA-seq data aggregated as pseudobulk.”

      (9) The evolutionary analysis also appears somewhat disconnected from the rest of the study. Could the authors leverage available public datasets to test whether a similar Bcl11a expression trajectory is observed in human oligodendrocyte lineage cells?

      We thank reviewer for mentioning this. We would like to clarify that the evolutionary analysis was included to examine whether Bcl11a sequences across vertebrates, including humans, show evidence of selective constraint, meaning that the sequence has been preserved during evolution because changes in it are likely to be disadvantageous. This analysis was therefore intended to provide broader evolutionary support for the functional importance of Bcl11a, rather than to stand as a separate or disconnected component of the study.

      For this analysis, we included Bcl11a DNA and protein sequences from 23 vertebrate species, including humans. We refer the reviewer to the Methods section of this paper, under “dN/dS analysis”, for further details. To provide further clarity regarding the different species used in this study, we have now prepared a new Supplementary Table 4, listing the 23 species together with their DNA and protein sequence accession numbers for Bcl11a.

      We also added this sentence in the main text:

      “We included twenty-three vertebrate species, including humans (Supplementary Table 4).”

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Given how central the isolated cells are to the subsequent analysis, the manuscript would be strengthened by a figure showing expression of stage and lineage-specific markers.

      Ideally, the authors would provide some sort of orthogonal experimental approach to confirm downregulation of Bcl11a during oligodendrocyte differentiation and loss with age (e.g., IF or RNAScope in conjunction with stage-specific markers in tissue, or western blot in culture).

      Thank you again. We have performed these. Please see the Reviewer #1 comment (above).

      Reviewer #2 (Recommendations for the authors):

      (1) Figure 1A: It should be 'anti-O4' instead of 'anti-04'.

      This is now corrected. Thank you.

      (2) Figure 1C: The authors should specify what the connecting lines indicate (e.g., gene sets or gene modules).

      Each coloured line represents one gene and connects its log2 fold-change values across the three oligodendrocyte lineage states: OPC, PreOL and OL. The connecting lines are used to visualize gene-wise patterns of expression change across these cell states. For example, in Type 1, each line shows a pattern in which gene expression increases progressively from OPC to PreOL to OL, with the highest expression change observed in OLs: OL > PreOL > OPC.

      We now included the following line in the figure legend:

      “Each coloured line represents one gene and connects its log<sub>2</sub> fold-change values across the three oligodendrocyte lineage states: OPC, PreOL and OL. The connecting lines are used to visualise gene-wise patterns of expression change across these cell states.”

      (3) Figure 2C: The authors should specify "DF genes" in the figure legend.

      Thank you for pointing this out. This was a typo: it was written as DF, but it should be DE (differentially expressed) genes. We have now corrected this in the figure and spelled out the abbreviation in the legend. Also, DE gene list is accessible through GEO accession: GSE303317. This also mentioned in the figure legend as:

      “DE: Differentially expressed. DE gene list is accessible through GEO accession: GSE303317.”

      (4) Figure 4C & Figure 5B: the title for the y-axis of the bar graph is confusing. The authors should specify what "#" indicates. Does it represent the counts? What are the thresholding criteria to judge whether an Olig2 cell is MBP-positive or not? It's unclear what the unit is here for the 0-100 scale.

      We apologise for the confusion. We used ‘#’, which is a common notation in mathematical and quantitative contexts, to denote counts, so you are correct. We now mentioned in the legend: “The symbol “#” indicates cell count.”

      We counted the number of MBP+OLIG2+ cells, divided this by the total number of OLIG2+ cells, and expressed the value as a percentage. For greater clarity, instead of writing #MBP+/#OLIG2+, we have now written #MBP+OLIG2+/#OLIG2+.

      Regarding the 0–100 scale, the unit of the Y-axis is percentage, as stated in both figure legends.

      The criterion for classifying an OLIG2+ cell as MBP+ was morphological: an OLIG2+ nucleus, shown in white, had to be surrounded by MBP+ staining, shown in red. Cells meeting this criterion were counted as MBP+OLIG2+ cells. Manual counting was performed blinded to sample identity.

      (5) Figure 6B: To discriminate from IF staining, the authors should use italic'Plp1' to indicate the RNA in situ results.

      Thank you for pointing this out; we have now corrected it.

    1. eLife Assessment

      This important work sets out to identify the neural substrates of associative fear responses in adult zebrafish. Through a compelling and innovative paradigm and analysis, the authors identify brain regions associated with individual differences in fear memory expression. While most findings are well supported, there is limited evidence that the four behavioral clusters represent discrete forms of associative fear-memory expression and the related interpretations would benefit from additional analysis or more cautious framing. Nonetheless, this study showcases the strength of zebrafish for systems-level neuroscience and will be of broad interest to the neuroscience community.

    2. Reviewer #1 (Public review):

      Summary:

      This work provides a comprehensive analysis of how adult zebrafish show fear responses to conspecific alarm substances (CAS) and retain their associative memory. It shows that freezing is a more reliable measure of fear response and memory compared to evasive swimming, and that the reactivity and the type of responses depend on the zebrafish strain. It further suggests neuronal substrates of different fear responses based on c-Fos mapping.

      Strengths:

      The behavioral part is the most comprehensive and detailed yet in the zebrafish field, providing strong support for the authors' claim. The flow from Figure 1 to Figure 4 is very smooth. They provide extremely detailed, yet complementary and necessary, analyses of how different categories of behavior emerge over time during the CAS exposure and memory retrieval. I'm convinced that neuro researchers who study fear/stress responses will always refer to this paper to plan and interpret their future experiments.

      Comments on revised version:

      The authors successfully addressed my comments, including the addition of Figure S6-2, which gives us some intuition into the relationships between c-Fos levels in individual areas and the behavioral outputs.

    3. Reviewer #2 (Public review):

      In this study, Fontana et al. develop a paradigm for associative conditioning by pairing exposure to alarm substance with a novel tank. Exposure to conspecific alarm substance (CAS) in the novel tank triggers freezing and what they characterize as evasive swimming behaviour, which are subsequently seen in a re-exposure to the novel tank without the CAS present. Importantly, these states are identified via automated processes including postural tracking and a random forest classification process, which could be very useful tools for subsequent studies.

      In their experiments they focus on the differences in behaviour among strains of zebrafish (both males and females), and among individual zebrafish. For males and females of different strains they find some differences, though the clearest message seems to be that the most robust measure of the behaviour in response to both the CAS and in the memory trials is the freezing behaviour, while evasive behaviour is more variable and not always seen. This may relate to their observation of significant "evasiveness" in vehicle control experiments (discussed further below).

      Moving on to individual variation from within this multi-strain male/female dataset, they first examine transition matrices between states, and find this is not dramatically altered by stimulus exposure. They then use clustering to identify 4 different "classes" of zebrafish that differ in their expression (or not) of two types of behaviour: freezing and/or evasive behaviour. They show that over the three exposure epochs of the experiment this classification is somewhat stable in an individual fish, though many fish change their behaviour -- e.g. evading + freezing -> only freezing.

      In the final set of experiments they move beyond behavioural analyses and perform whole-brain cFos mapping of these individual zebrafish, and perform analyses aimed at identifying correlations between individual behavioural expression and the number of cFos positive cells in different brain regions. Using partial least squares analysis they find areas associated with two types of behavioural contrasts, which differ in their weighting of different behavioural expression during the Memory trials. Covariation and network structure analysis within different classes of fish also find some differences in covariation among brain areas, providing hypotheses as to underlying network effects that may govern the expression of freezing and/or evasive behavior in the memory trial phases.

      Overall, I find this to be an interesting study that employs state of the art methods of behavioural analyses and whole-brain cFos analyses. The revision has clarified the take-home message considerably: the abstract is now more careful about which behavioural groups are memory-associated, and the causal language in the conclusions has been appropriately softened. Two of my three original main concerns have been addressed. The first is not and having looked at the data again I can now be more specific about what concerns me.

      Comments on revised version.

      (1) My first concern related to the claim that fear memory behaviour falls into four distinct groups, and specifically to the role of evasiveness in defining them. The authors give three reasons for retaining it, but I remain unconvinced.

      The first is that variable evasion in response to alarm substance is a long-standing observation (von Frisch; Suboski et al.), and that dissecting this individual variation is the purpose of the paper. I agree with the motivation, and it is a good reason to measure evasion. But it does not establish that evasion on memory day reflects fear memory, and memory day is the only day used for the clustering and neural activity mapping. The manuscript's own results point the other way: relative to pre-exposure, no strain or sex increased evasion on memory day, and relative to vehicle only female TUs did. The temporal profiles show evasion on memory day to be largely similar between vehicle and CAS-treated fish. Historical observations of variable evasion during CAS exposure do not carry over to the memory phase.

      The second is that the clustering itself reveals two kinds of freezing fish - one freezing between bouts of normal swimming, the other between bouts of evasion - demonstrating that a subset of fish increase evasion. In absolute terms, this does not match the data. In Figure 4B, evading freezers are below the population mean for absolute evasion, as are freezers. The text describes evading freezers as "high in freezing and evasive behaviors," and I do not think Figure 4B supports this.

      What actually separates the two freezing groups is the third measure, evasion as a percentage of active time. And this is where I have difficulty, because that measure is not an independent behavioural readout. The classifier assigns every window to normal, evasive or freezing, and active time is simply non-freezing time, so evasion-as-percent-of-active is fully determined once the other two are known.

      This matters for the clustering specifically. Distance-based methods weight each input dimension equally, so a variable that carries no information beyond the other two nonetheless contributes a full third of the distance between any two fish - and it contributes it in a way that counts freezing twice, once directly and once through the denominator of the derived measure. The space is nonetheless described as three-dimensional throughout, including in the Methods and the Figure 4 legend, when there are only two independent behaviours in it.

      The consequences fall hardest on exactly the animals at issue. Both freezing groups sit at 65-70% freezing, so there is very little active time to divide by, and small absolute differences in evasion - together with any noise in estimating them from a couple of minutes of non-frozen behaviour - are inflated into large differences on the rescaled measure. In terms of what the fish actually did, the two groups differ by a few percent of trial time. That is the boundary on which much of the rest of the paper rests.

      I recognise that evasion as a proportion of active time is in some respects the more biologically meaningful quantity, and the authors are right that a fish freezing 70% of the time has limited opportunity to do anything else. But that is an argument for reporting it as a descriptive measure, not for entering it into the clustering alongside the two variables from which it is computed.

      This impression is reinforced by Figure 4A itself. While the freezer group occupies a reasonably distinct region, the non-reactive, evader and evading freezer groups appear as a single continuous distribution with cluster boundaries drawn through it rather than around visible gaps. I appreciate that UMAP is a projection and that visual separation is not required for genuine structure, but this is the figure by which most readers will judge whether four discrete types exist, and it does not obviously support that reading - particularly given that the embedding is built from the same variables, including the rescaled measure, that most favour the separation.

      I would suggest that the authors re-run the clustering using only the two directly measured behaviours, percent freezing and percent evasion of total time, and report whether four groups still emerge and, in particular, whether the evading freezer / freezer split survives.

      The third is that the two groups have distinct functional networks despite equally high freezing, so the behavioural difference is real and is manifesting in the brain. This is the strongest of the three arguments, and I accept part of it: something about how a frozen fish spends its remaining active time does appear to be neurally meaningful, which is interesting in its own right. But it does not establish that these are two distinct types, nor that the difference has anything to do with the conditioning. Fish taken from either side of a cut through a continuous distribution will differ neurally if that continuum tracks brain state, so the network result is equally compatible with graded variation. More importantly, Figure 5A shows that a substantial proportion of fish are classified as evaders in the vehicle condition and at pre-exposure, before any CAS has been given. This suggests a pre-existing individual tendency toward evasive behaviour that is independent of the alarm substance, and one would expect such a tendency to persist into the memory trial. If so, the distinction the network analysis is drawing between freezers and evading freezers may simply reflect that baseline trait, and its neural correlates would be correlates of the trait rather than of fear memory. I am therefore not convinced that this distinction is related to CAS or to memory.

      (2) This concern is fully resolved. I had misread the CAS preparation: it was pooled from eight donors spanning all four strains and both sexes, so every fish received identical material and the strain and sex differences cannot be attributed to donor variability. The clarification now added to the Results will prevent other readers making the same error. The addition of FDR correction to the Figure 2 comparisons also addresses my related concern about multiple testing.

      (3) Somewhat resolved. The conclusion no longer states that behavioural variation is "driven by" activity in particular regions, and the added caveat that neural activity was not directly manipulated sets the right expectation for a mapping study. The scatterplots in Figure S6-2 are a useful addition and give a much better intuition for what the PLS contrasts represent. My remaining reservation is the one above: a great deal of the neural story rests on the evading freezer / freezer contrast, and I am not persuaded that this contrast marks a boundary relevant to fear memory.

    4. Reviewer #3 (Public review):

      This revised manuscript by Fontana et al. aims to study how animals respond to fearful stimuli, with a specific focus on brain regions involved in predicting animals that passively freeze or those that actively evade the threat. I continue to be enthusiastic about the study. The study addresses an important question regarding individual variation in fear-related behavior and links these behavioral phenotypes to whole-brain activity patterns in adult zebrafish. The combination of a contextual fear conditioning paradigm, strain/sex comparisons, behavioral clustering, and AZBA-based c-Fos mapping makes this a valuable contribution to the field, not just in answering the question posed by the authors, but also in formulating a framework for using adult zebrafish for whole brain analysis of complex behaviors. Overall, I find the authors have responded to my concerns:

      (1) I still think that separating memory acquisition and consolidation is an interesting question, and further use of the framework will need to eventually solve that; however, I also appreciate that this may be beyond the scope of the current study, and I appreciate the authors acknowledging this in the manuscript.

      (2) Regarding Figure 3, I also agree that this is difficult to present differently, and I appreciate the authors adding text to the body to clarify things. My one request is that the sentence (lines 214-215) that reads: "This increase in evasion in the vehicle group likely represents a response to the water disturbance that occurs when solution is added to the tank." Be changed to: "This increase in evasion in the vehicle group may represent a response to the water disturbance that occurs when solution is added to the tank." While it is entirely possible, there are no concrete data to support that this is "likely."

      (3) I appreciate the clarification regarding the PLS-derived contrasts in Figure 6A and in the body.

      Overall, this is a really interesting paper that will have a wide-ranging impact. All of my concerns have been addressed.

    5. Author response:

      The following is the authors’ response to the original reviews.

      We thank the reviewer’s for their thoughtful comments that have significantly strengthened the paper. Below, we have outlined our responses to both the public reviews and recommendations.

      In addition to the alterations to the manuscript based on the reviews, during our review of the data analysis we uncovered some small errors that we have now corrected. In looking back over the image registration, we identified three animals whose olfactory bulbs did not register properly and one with poor cell counting in the telencephalon. To account for these issues, we imputed the missing data using an iterative soft-threshold singular value decomposition (described on lines 779-783 of the updated manuscript). This update had little impact on the results. We also identified a small error in how we determined ‘unique’ and ‘overlapping’ edges in the network analysis (Figure 8). In the previous analysis we had incorrectly noted that all ‘unique’ edges did not have an overlapping confidence interval with the two other networks (i.e., the networks for evading freezers, freezers, and non-reactive). Instead, the ‘unique’ edges in the prior version of the manuscript did not have an overlap with at least one other network. We have now updated the analysis so the reader can distinguish between edges that are truly ‘unique’ versus those with ‘1 overlapping confidence interval’ or ‘2 overlapping confidence intervals’ with other networks. As before, this update and change to the analysis does not materially affect the results or conclusions.

      Public Reviews:

      Reviewer #1 (Public review):

      Weaknesses:

      The neural analysis part is very comprehensive. Figure 5 and Figure 6 are independent but complement each other very well. They together support that the cerebellar system is the key brain component for a freezing response. Their extreme focus on high-level analyses, however, came at the expense of biological intuitions. I suggest adding some figure panels and result/discussion paragraphs to help with that aspect.

      Thank you for the suggestion. We have made extensive edits to the manuscript to include additional discussion and biological intuition. Specifically:

      We added a supplemental figure (Figure S6-2) that has scatterplots showing how cfos levels vary with the different behavioral contrasts. Although the PLS analysis is multivariate, this univariate analysis should help give readers a better intuition of how the behavior relates to brain function.

      We have also rewritten the results sections for both the PLS analysis (lines 303-361) and network analysis (lines 396-437) to incorporate more of a discussion about the biological context of different regions identified. Thank you for this suggestion, we feel that this significantly strengthens the biological interpretation of the data for the reader.

      Reviewer #2 (Public review):

      (1) My first concern relates to the claim in the abstract that "We found that fear memory behavior fell into four distinct groups: non-reactive, evaders, evading freezers, and freezers".

      In my opinion, the "freezing" aspect is well supported as being both triggered by the CAS and for memory effect upon re-exposure to the tank, but I am less convinced about the "evasive" behaviour. In Figure 2, it appears that "evasiveness" is generally not increased in both the Exposure or Memory phases for many groups, and in Figure 5, it appears that "evasiveness" is expressed by nearly 50% of the fish in the pre-exposure condition before CAS addition and in all phases in the vehicle condition. Therefore, it appears that most of the expression of this behaviour is independent of any memorybased effect.

      We thank the reviewer for this suggestion and we agree that this line in the abstract was unintentionally misleading. We have now altered this line in the abstract (lines 34-36) to read:

      “We also found that that behavior fell into four distinct groups: non-reactive, evaders, evading freezers, and freezers with the evading freezer and freezer groups most clearly associated with memory formation.”

      On the larger point of the inclusion of evasion as part of the fear response, we believe this is warranted for the following reasons: (1) evasive behavior has long been acknowledged as a highly variable aspect of how fish respond to alarm substance where some fish exhibit evasion and others do not. This observation goes back to the original work from Karl von Frisch in minnows (von Frisch, 1938), and others in zebrafish (e.g., Suboski et al, 1990). One goal of our paper (and the work from the lab in general) is to try dissecting out this individual variation that can get lost when only considering population averages. (2) The unsupervised clustering also suggests that there are two distinct types of freezing clusters (Figure 4B) where some fish freeze intermittently with normal swimming and others freeze intermittently with evasive behavior. This suggests that evasion is increased in response to CAS, but only in a subset of fish. (3) The brain networks from the evading freezer and freezer groups are distinct (Figure 8A) despite having equally high levels of freezing behavior (Figures 4B and C). This means the difference we’re able to distinguish behaviorally is also manifesting in the brain, suggesting that it is not anomalous. Thus, while we agree that freezing is definitely the strongest and clearest behavioral response to CAS, we believe the analysis of this large dataset supports the interpretation that, in a subset of fish, increased evasive behavior in response to CAS is also a part of the response.

      (2) My second concern relates to the claim in the abstract that "background strain and sex influenced how fish respond to CAS, with males more likely to increase evasive behaviors than females and the TU strain more likely to be non-reactive."

      My understanding, based on the introduction and on the methods, is that it is likely important that the CAS be prepared from conspecifics of the same strain and sex, and for this reason, they prepared different CAS specific for each strain and each sex. Therefore, the "CAS" that is applied is necessarily different for each condition, and I am concerned about if the differences observed could relate more to variation in the quality, purity, concentration, etc. of the specific CAS samples for different groups, rather than their reactivity to the substance or their ability to form memories based on such experiences.

      The CAS was prepared by mixing extracts from all four strains and both sexes (so 8 fish per batch). Thus, all the fish were exposed to the same CAS mix derived from the same donors. This is described in the methods (lines 626-629). However, to ensure that this is clear to readers, we’ve now included a line indicating this in the results section (lines 123-124).

      (3) My third concern relates to the interpretation of the cFos data.

      As I mentioned above, I feel as though the behavioural analysis is perhaps more complex than is warranted via the inclusion of evasiveness, and I wonder if the conclusions from the experiments would be simpler if analyzed only from the perspective of freezing.

      We agree that the freezing response is driving the majority of the neural cfos response that we are seeing (e.g., Figure 6A-C). However, we feel that the network analysis (Figure 8) justifies the distinction between freezers and evading freezers. This is because the brain networks for these two groups (freezers and evading freezers) are quite distinct, even though these groups both have the same levels of freezing behavior (Figure 4). This stark difference in patterns of neural activity suggests the brain of a freezer and an evading freezer are engaging with the world in two distinct ways that is worth noting. We’ve updated the abstract to make this point clearer (abstract: lines 39-48) and discuss the biological interpretations of patterns of brain activity unique to evasion or evading freezers in more depth (lines 303-361; lines 396-437).

      Reviewer #3 (Public review):

      (1) The three-day contextual fear paradigm, as implemented - one CAS pairing on day 2 followed by a single recall test on day 3 - inevitably conflates acquisition and long-term memory, making it impossible to know whether strains like TU truly recall the association poorly or simply learn it more slowly. For example, given that TU fish extinguish fear faster than AB or TL strains in extended protocols, they may simply require additional or repeated CAS pairings to achieve the same asymptotic performance. To disentangle learning kinetics from recall strength, the assay could be revised to include multiple acquisition trials (e.g., conditioning on two or more consecutive days) with an immediate post-conditioning probe to assess acquisition independent of consolidation, and continuous measurement of freezing and evasive behaviors across each trial to fit learning curves for each strain. Such refinements - even if on a subset of the strains - would reveal whether "non-reactive" phenotypes reflect genuine recall deficits or merely delayed acquisition.

      We thank the reviewer for this thoughtful comment. We agree that it is difficult to disentangle acquisition from consolidation. Indeed, the TU fish do appear to have lower levels of freezing in response to the CAS (Figure 2A), supporting the idea that reduced performance at memory day could be due to some sort of deficit at acquisition. However, pursuing a detailed examination of strain dependent differences in fear memory acquisition versus consolidation is beyond the scope of the current paper where we primarily focus on individual differences in behavior. Nonetheless, we have included this important point in the discussion (lines 470-471).

      (2) My second major question is with respect to Figure 3 panel B. This is a complex figure, and I can understand the gist of what the authors are attempting to show, but it is difficult to understand as it is. Can this be represented in a way that is clearer and explained a bit more easily?

      We agree that this figure is one of the more complex in the paper. However, we’ve struggled to come up with a better way to present it. We have improved the presentation based on other reviewer comments by making the vehicle and CAS groups more easily distinguishable by using open versus closed circles. We’ve also included additional interpretations of the data in the results, which we hope will help guide readers through this figure better (lines 208-223).

      (3) The brain mapping is by far one of the most interesting aspects of this study, and the methods that the group used are interesting. The brain mapping, however, relies on generating "contrasting" groups (Figure 6A), and I was not clear as to how these two groups were formed. Could the authors elaborate a bit?

      These contrasting groups (contrast 1, contrast 2) arise analytically from the partial least squares (PLS) analysis; they are not defined by the experimenter. In brief, PLS is a multivariate technique that identifies latent variables that capture axes of maximal covariation between two datasets: behavior and brain activity. As an analogy to a more widely known technique, principal components analysis (PCA) uncovers axes of maximal variance within a single dataset. PLS, in contrast, simultaneously analyzes the covariation in two datasets. The contrast groups in Figure 6A represent the behavioral weights of the latent variables that capture the most covariance, which illustrates how the four behaviors load onto these top two contrasts.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Major points:

      (1) The c-Fos analysis in Figure 5 is very comprehensive and convincing, but lacks intuitive presentations. In my understanding, the increase in c-Fos expression in red areas means increased freezing behavior for Contrast 1 for the PLS analysis? Do you have representative c-Fos expression images between different groups of fish?

      We decided not to include a representative cfos image because the data is derived from a large number of fish (N=87) and thus it can easily be cherry-picked to choose images that match the narrative. Instead, to more accurately capture the breadth of the data while providing a more intuitive presentation, we have included an additional supplemental figure that includes scatterplots of scaled cfos data against behavioral scores for each of the two contrasts (S6-2). We believe this more fully and accurately captures the relationship between behavior and brain activity. We included six different example brain regions and scatterplots for cfos activity against behavioral scores for contrasts 1 and 2, demonstrating a range of relationships. However, we should note that PLS is a multivariate technique, and so this univariate analysis does not fully capture the subtleties of the PLS analysis. Nonetheless, we think this will help give a more intuitive interpretation of the data to readers. We have also referenced this additional data in the manuscript (lines 307-309). We thank the reviewer for this excellent suggestion that improves the ability of readers to understand the paper.

      (2) Also related to Figure 5, the result section only describes the PLS statistics and does not try to describe the biological interpretation. Do the authors think the c-Fos expression directly represents lowlevel behavior, such as swimming, or a high-level behavioral state or learning? Maybe different areas mediate different aspects?

      For example, the medullary locomotor areas, which are usually highly correlated with swimming in terms of neural activity, seem to have higher c-Fos expression in freezing fish. I'm not saying this shouldn't be the case. c-Fos expression in this area was not elevated in larval fish during OMR in Shainer et al., 2023, indicating that it doesn't linearly reflect neural activity. But discussing a bit of intuition on the connection between c-Fos expression and biological process, rather than just saying "the cerebellum could regulate emotional states", would help us guide through this highly complex analysis.

      We have now added more interpretation of the data in both the PLS and network analysis sections (lines 303-361 and lines 396-437). Again, thank you for this excellent suggestion. This helps make the biological interpretation of the data clearer.

      Minor points:

      (1) Figure 2B titles: please write "memory" on the right side.

      We considered writing ‘memory’ on the right-hand side, but we thought this may add confusion because it would not apply to both graphs in the row. The left-hand graphs are the responses during ‘exposure’ and the right-hand graphs are the responses during the ‘memory’ phase. This is indicated by the titles above the left and right-hand sets of graphs.

      (2) Figure 2C: needs legend lines.

      We have now moved the legend lines from the top of the graphs to below the graph to make them more visible to readers.

      (3) Line 187: "aggregated" data.

      This has now been changed to ‘aggregated’ (now line 194).

      (4) Line 371: I'm not sure what "Beyond" means.

      We have now significantly changed this part of the paper and we no longer use the word ‘beyond’ here.

      Reviewer #2 (Recommendations for the authors):

      (1) Regarding point (1) in the Public Review:

      I would encourage the authors to consider whether this study might be better focused exclusively on the freezing behaviour, which does appear to be reliably expressed during CAS exposure and in the memory phases, and would significantly simplify the subsequent analyses of neural activity, and perhaps may lead to a more coherent conclusion.

      As noted in our response to the public review, we appreciate this suggestion, but we have decided to keep the inclusion of the evasive behavior. This is because (1) evasive behavior has long been acknowledged as a highly variable aspect of how fish respond to alarm substance where some fish exhibit evasion and others do not. This observation goes back to the original work from Karl von Frisch in minnows (von Frisch, 1938), and others in zebrafish (e.g., Suboski et al, 1990). One goal of our paper (and the work from the lab in general) is to try dissecting out this individual variation that can get lost when only considering population averages. (2) The unsupervised clustering also suggests that there are two distinct types of freezing clusters (Figure 4B) where some fish freeze intermittently with normal swimming and others freeze intermittently with evasive behavior. This suggests that evasion is increased in response to CAS, but only in a subset of fish. (3) The brain networks from the evading freezer and freezer groups are very distinct (Figure 8A) despite having equally high levels of freezing behavior (Figures 4B and C). This means the difference we’re able to distinguish behaviorally is also manifesting in the brain, suggesting that it is not anomalous. Thus, while we agree that freezing is definitely the strongest and clearest behavioral response to CAS, we believe the analysis of this large dataset supports the interpretation that, in a subset of fish, increased evasive behavior in response to CAS is also a part of the response.

      A more minor concern related to the analyses in Figure 2: in the figure legend, it is stated that "*-P < 0.05 compared to vehicle treated fish via t-tests". How are the authors dealing with the multiple comparisons problem? Would something like an ANOVA not be more appropriate?

      Thank you for bringing this point up. We did not initially correct for multiple comparisons because we considered each of these experiments across sex and strain separate since we did not compare across strains. However, the way we’ve grouped the data together in figure 2 makes it appear as if they are one large experiment. To alleviate any concern about multiple testing, we have now corrected for multiple comparisons using the false discover rate (FDR) correction. The statistics in the figure and captions have now been updated.

      (2) Regarding point (2) in the Public Review:

      If the authors agree with my concern regarding potential variability in the CAS samples, I would suggest either testing for differences among strains using the same batch of CAS, or including and explaining this caveat in the text.

      As noted in our response to the public review, the CAS was the same for all the fish. Each batch was derived from 8 donor fish, one fish from each strain and sex (described in lines 123-124 of the results and lines 626-629 of the methods).

      (3) Regarding point (3) in the Public Review:

      I feel like the standard in the field for such conclusions would be after

      (a) Direct analyses of the activity states in these areas. I was surprised not to see a direct analysis of the cFos stainings in the cerebellum relative to freezing behaviour, for example, ideally in a different animal cohort.

      The PLS analysis does relate activity in the cerebellum (and other brain regions) to specific behaviors via the the behavioral contrasts (Figure 6A). We believe this approach (instead of dividing fish into ‘high and low freezers’) is a more powerful way to leverage the data from all the animals tested (87 fish). However, we appreciate that the interpretation of the PLS analysis is not as intuitive as seeing scatterplots or bar charts comparing neural activity. For this reason (and in response to a comment from reviewer 1), we have included as a supplementary figure (Figure S6-2) scatterplots showing how standardized c-fos activity varies with the behavioral scores from the contrasts identified from the PLS analysis. Given that contrast 1 weights heavily in the positive direction on freezing, these figures can essentially be read as looking at cfos activity as a function of freezing levels. What can clearly be seen is that for regions of the cerebelleum (E.g., the LCa and CC) there is a clear positive relationship between cfos activity and the behavior scores for contrast 1.

      (b) Some kind of manipulation of the brain area resulting in the relevant behavioural modification.

      We completely agree with the reviewer. However, at the moment, we do not have the tools to do this in adult zebrafish. It is something we’re actively working on.

      Of course, I appreciate that such experiments might not be possible or feasible, and in which case I would suggest adjusting the claims accordingly and highlighting the caveats to their interpretations.

      We have incorporated the caveat that we have not directly altered neural activity into the discussion (lines 542-543) and adjusted how we discuss our findings in the abstract (lines 39-41) to more accurately represent the type of evidence we provide. Hopefully we’ll be able to do so in the near future!

      MINOR CONCERNS:

      (1) In Figure 3, how is the end of a behavioural epoch defined? I am surprised to see that you consider transitions between the same behavioural state. How does erratic swimming -> erratic swimming differ from a longer single epoch of erratic swimming? In general, I find this analysis confusing, and I am not sure if it adds significantly to the message of the paper.

      Thank you for this question as it prompted us to realize we were missing this in our methods section. We have now updated the methods to include how we calculated the behavioral transitions (lines 644-650). In short, we used a 750 ms behavioral epoch time that corresponds to the size of the sliding window we used for the random forest model.

      We have also updated the description of this analysis in the results to indicate the main finding from it (lines 207-223). In brief, the main finding is that exposure to CAS results in longer bouts of evasive behavior without increasing its frequency. Whereas CAS induced freezing arises from both longer bouts and likelihood of occuring. While we agree that this is a relatively minor finding in the paper, one of our goals is to provide as comprehensive analysis of fear behavior as possible to help guide future researchers interested in using fish for understanding different aspects of fear-related behaviors.

      (2) In the PLS analyses, two measures of evasion are used: evasion time, and evasion as a percent of active behavior. I don't understand the justification for both of these being used rather than one. Again, my overall recommendation is to reduce the focus on the analysis of evasion behaviour, but if you do not choose to do this, I think the rationale of how both measures are used and why needs explanation.

      We chose to incorporate two different measures of evasion throughout the study because the high levels of freezing in some animals results in little opportunity to express other behaviors (like evasion). Thus, to better capture what fish may be doing in the absence of freezing (i.e., when they are active) we also calculate the amount of active time spent performing evasive behaviors (instead of normal swimming). We have now included an explanation for this earlier in the results section when we first use this metric (lines 149-152).

      (3) In the methods, I don't understand this: "Animals that were assigned the wrong sex were removed from data analysis, as well as its paired fish (< 2%)".

      We determine the sex of fish when we set them up for dual housing. However, we occasionally make errors in sex determination. To ensure we properly sexed the fish, at the end of experiments, we euthanize the fish and check for the presence of eggs. If we incorrectly assigned the sex to a fish, they are removed from the experiment alongside the other fish they were dual housed with. This is because we want to ensure all fish are housed in the same way (i.e., a male fish with a female fish).

      (4) How was this determined differently from the first time, resulting in exclusion?

      After experiments, fish were euthanized and we checked for the presence of eggs (line 599-601). We’ve now added a line in this other part of the methods referring back to where we describe this (lines 676678).

      Reviewer #3 (Recommendations for the authors):

      Here are some minor concerns and errors found in the manuscript:

      (1) For Figure 2B and Figure 3B, can the group make the lines solid and dotted? The circle or triangle designation is difficult to see, and since the crux of the figure depends on comparing Veh and CAS, it would be easier to see if the lines were altered.

      Thank you for this suggestion. Instead of making the lines solid and dotted, we decided to make both the CAS and vehicle group circles and then have open and closed circles. We believe this solves the issue of being able to distinguish these groups and makes the data more readable.

      (2) Figure 2C: It appears that the line colors in the legend are missing.

      We have moved the line colors below the graphs to make them more obvious.

      (3) Figure 8A: Same thing here - could the text be enlarged? It's really difficult to make out each node, and when I zoom the text becomes pixelated. This is an important figure and one that will likely be referenced, and making it clear would be helpful.

      This one is difficult. We have made the network images as large as would fit on a page. We have now uploaded vectorized versions of the images so that they do not become pixelated when zooming in. As part of our supplemental materials we also include a cystoscope file that can be explored in greater depth as well.

      (4) The paper is really well written: I found a few typos, though:

      (a) Line 529: "Institutional Cara and Use Committee" should be "Institutional Animal Care and Use Committee" (Change cara to care and add animal).

      (b) Line 274: "hybdridization" should read hybridization.

      Thank you for catching these typos. They have now been fixed.

    1. eLife Assessment

      This study provides useful information for the Drosophila ageing community by characterising the auxin-based gene expression system (AGES) and identifying caveats associated with its use in adult flies. The authors provide solid evidence for sex-, age-, tissue- and dose-dependent variability in transgene induction, as well as effects of auxin feeding and AGES activation on stress resistance, metabolism, and lifespan. While the study would benefit from a broader characterisation of some of these limitations, the findings provide a helpful benchmark for researchers using AGES in ageing studies.

    2. Reviewer #1 (Public review):

      Summary:

      The authors set out to evaluate whether AGES, a recently developed auxin/TIR1-based conditional GAL4 expression system, is a suitable tool for Drosophila ageing research. They characterise induction efficiency across sex, transgene insertion site, auxin dose and age, then test whether AGES can replicate a well-established pro-longevity manipulation (dominant-negative insulin receptor expression).

      Strengths:

      The study is thorough and methodical. The authors use appropriate genetic controls throughout, which is required to properly interpret AGES-based experiments. They identify an important issue, in that activation of the AGES machinery itself (independent of any UAS-transgene) shortens lifespan and alters protein levels, while high-dose auxin independently affects body mass, and even a moderate dose (5 mM) impairs stress resistance across all genotypes. These findings are important for researchers when interpreting their experiments. The tissue and age mapping of induction efficiency (brain, fat body, gut) is also useful, and the inclusion of driver-only positive controls at each age (Figure 2) establishes that da-GAL4 activity itself is stable across the ages tested, ruling out declining driver activity as an explanation for the reduced induction seen in older flies (though, as noted below, reduced auxin ingestion with age remains a very plausible contributing factor alongside declining AGES efficacy).

      Weaknesses:

      Longevity and stress assays were conducted only in females, which, combined with the finding that males show weaker and less consistent induction, means the study cannot speak to whether the metabolic and survival costs of auxin/AGES activation observed here also apply to, or differ in, males. The KCl vehicle control matches the potassium cation (K⁺) content of K-NAA across conditions; therefore, chloride (Cl⁻) concentration differs between control and auxin-fed media (both minor weaknesses).

      Achievement of aims and impact:

      The authors achieve their stated aim. Rather than validating AGES as unambiguously suitable for longevity work, they set out to characterise its behaviour and limitations in this context, which they do convincingly. The data support their overall conclusion that AGES can be used to conditionally induce transgene expression at advanced ages, but that its use in longevity/healthspan studies requires caution and rigorous control genotypes. This is a useful contribution with direct practical value: it will help other researchers make informed decisions about whether and how to deploy AGES in ageing-related work, and the cautionary findings regarding auxin/AGES toxicity are likely to be of broad relevance to the growing community of AGES users beyond the ageing field specifically.

    3. Reviewer #2 (Public review):

      McGilvary et al. evaluate the recently developed auxin-based gene expression system (AGES) for use in aging studies of Drosophila melanogaster. This system is based on the widely used Gal4/UAS system that enables cell-specific expression of UAS-transgenes under Gal4 activator control. AGES uses an auxin-inducible degron-tagged Gal80 repressor that should prevent Gal4-dependent activation unless flies are fed auxin, providing a useful approach for temporal control of transgene induction - something that would be highly useful for aging studies. The authors perform a comprehensive analysis of AGES-dependent transgene induction in male and female flies at different ages with multiple controls, demonstrating some moderate induction in female flies only - albeit with some substantial background induction even in the absence of auxin.

      Overall, transgene induction appears to be both much lower with the AGES system compared to Gal4 driver controls and very leaky, with some tissue-specific differences in induction observed as well. Combined with their observations that auxin feeding has impacts on body mass, triacylglycerol and protein levels, and lifespan, these data raise some concerns regarding the interpretation of data obtained using the AGES system for aging or longevity studies in flies. This study provides well-needed validation for the recently developed AGES system and highlights critical caveats that will support future studies.

      Most conclusions of the paper are well supported by data, but additional controls and textual edits would strengthen and clarify the findings. In addition, the abstract and conclusions of this study should more accurately reflect the limitations of transgene induction using this AGES system in adult flies.

    4. Reviewer #3 (Public review):

      Summary:

      In this useful work, the authors characterize the auxin-based gene expression system (AGES) as a tool for studying ageing. They found that this system can be applied to ageing studies. In addition, they identified important drawbacks of the methods, including effects of insertion sites and sex on induction of the system, and that some auxin doses may have inadvertent effects on body mass and physiology. Overall, the study extends the AGES system for use in fly ageing studies and highlights some caveats. While the findings are solid pointers, more extensive characterization is needed to benchmark the extent of the caveats identified.

      Strengths:

      The study provides the first longitudinal evaluation of the AGES system's induction efficiency across the entire Drosophila lifespan. The authors also highlighted a number of caveats of the AGES system in ageing animals. These are all important points to be considered when using this system, and findings should be interpreted keeping these caveats in mind.

      Weaknesses:

      There were inconsistencies with auxin dosages between the figures.

      (1) The authors used a higher dose of auxin (20mM) compared to the original AGES paper (McClure, 2022) in Fig 3. The auxin dose-dependent effects are not linear for TAG and protein levels, highlighting that the genotype-dependent effects may be highly variable and may yield quite different results in other studies.

      (2) The results in Figure 4 showing the lack of induction in the brain are quite interesting; however, only 5mM auxin is tested. Characterizing dose-dependence for the variability of the induction in tissue types would be useful. At the very least, recapitulating prior results at 10mM should be done.

      (3) Food intake and hydration status were not measured alongside body mass, TAG, and protein endpoints. The changes seen could be an effect of decreased feeding or fluid balance rather than metabolic reprogramming.

    5. Author response:

      We thank the editorial team and all three reviewers for their time and attention to detail in reviewing our manuscript. We are particularly grateful for comments recognising the “direct practical value” and “comprehensive analysis” in our work (Reviewer #2) as well as its “thorough and methodical” approaches (Reviewer #1).

      We also appreciate the reviewers’ constructive comments to improve our manuscript, which we plan to address in a revised version. Specifically, we plan to more explicitly acknowledge some of the limitations of our study, including clearer highlighting of the experiments performed only in females (Reviewer #1); the inability of our experimental design to control for chloride concentration (Reviewers #1 and #2); the limitations of transgene induction using the AGES system in adult flies (Reviewer #2); and the rationale for using different concentrations of auxin across different experiments (Reviewer #3). We also note that Reviewers #1 and #2 have provided additional recommendations beyond the public reviews (largely relating to helpful ways to clarify our text and more explicitly acknowledge limitations), which we also plan to address.

      In addition, we plan to perform additional experiments to address specific points raised by reviewers in both their public reviews and additional recommendations. In the first instance, we plan to follow Reviewer #3’s suggestion to characterise induction in the brain using 10mM auxin, as well as Reviewer #2 and #3’s suggestions to explore feeding behaviour of flies fed auxin within our experimental setup.

      We look forward to submitting an improved manuscript guided by the reviews, with the aim of strengthening this “well-needed validation” study that also “highlights critical caveats that will support future studies” (Reviewer #2).

    1. eLife Assessment

      This important study integrates mouse genetics with sequencing, electrophysiological and behavioral tools to uncover the behavioral role and molecular profile of a developmentally defined subpopulation in the lateral septum. The data collected and analyzed are convincing, highlighting the role of this subpopulation in threat avoidance and stress, providing a framework by which developmental origin is linked to mature neuronal function. The work will be of interest to biologists and neuroscientists working on development and behavior.

    2. Reviewer #1 (Public review):

      This study investigates the role of a specific neuronal population in the lateral septum (LS) in balancing exploratory and defensive behaviors. The authors created a mouse model (cKO) lacking Nkx2.1-lineage neurons in the LS by deleting the Prdm16 gene. They discovered that this ablation specifically eliminated Crhr2-expressing neurons, which are normally targeted by urocortin-3 (UCN-3) inputs. Behaviorally, cKO mice did not show general changes in anxiety but displayed a significantly increased exploratory drive. In a predator odor test (using TMT), cKO mice spent more time investigating the aversive stimulus compared to controls, suggesting these LS neurons normally suppress exploration during threat. Furthermore, the study found that Nkx2.1-lineage neurons in the LS are specifically activated by acute stress (body restraint), as shown by an increased number of c-Fos-positive neurons. While the loss of these neurons caused some connectivity and electrophysiological changes, the remaining Nkx2.1-lineage neurons were more excitable. Therefore, the authors demonstrate that LS Nkx2.1-lineage/Crhr2+ neurons are a distinct population crucial for calibrating behavioral responses to stress, acting to inhibit exploration in favor of defensive strategies.

      This work provides new insights into the neural circuitry underlying anxiety and threat avoidance. However, some of the methods and data analyses require revision for greater clarity, and additional experiments and analyses are needed to further substantiate the conclusions.

      Some of my specific questions and concerns are as follows:

      (1) The authors showed a reduction in the size of LS and a specific decrease in Crhr2+ neurons in cKO mice. I would suggest examining whether the density of other types of neurons (e.g., Crhr1+ cells or other known cell types in LS) was altered in the cKO mice.

      (2) For the single-cell sequencing experiment (Figure 2), it is unclear whether tissues from the 3 male and 3 female mice within each genotype were pooled together or processed individually (i.e., as 6 separate samples). This information is not clearly stated in the manuscript. Given that male and female mice exhibited behavioral differences, it would be valuable to examine sex-dependent effects in the analysis shown in Figure 2.

      (3) Previous studies have shown that LS neurons exhibit distinct firing patterns, including regular spiking, bursting, complex-bursting, and phasic spiking. Since the authors recorded from both tdTomato-positive and -negative LS cells, it would be interesting to determine whether the positive cells display a unique firing pattern, thereby representing a distinct electrophysiological cell type within the LS.

      (4) More detailed descriptions of the electrophysiological data analysis should be provided in the Methods section. Some LS neurons display spontaneous firing without current injection; therefore, it should be clarified how the resting membrane potential was measured in these cells. The amplitude and onset latency of the first spike are presented in the figures; however, it is unclear how the first spike was selected-whether from spiking responses to rheobase current or to a specific current pulse. I would suggest defining the first spike based on responses at a certain firing frequency. The method used to determine the spike voltage threshold should also be specified.

      (5) Could the authors analyze the single-cell sequencing data to examine whether changes in ion channel expression might explain the observed alterations in spike waveforms?

    3. Reviewer #2 (Public review):

      Summary:

      The manuscript "Selective loss of Nkx2.1-lineage neurons in the lateral septum alters the balance between novelty seeking and threat avoidance" is an interesting study by Miguel Turrero García and colleagues. Here, the authors report a novel mouse model allowing complete ablation of neurons pertaining to the Nkx2.1 lineage by conditionally ablating the transcriptional regulator Prdm16 from the Nkx2.1 lineage. The authors combined single-nucleus RNA sequencing, histological and electrophysiological approaches, as well as behavioral analyses to demonstrate that a large portion of LS neurons are profoundly altered by Prdm16 deletion from the Nkx2.1 lineage. This manipulation preferentially impacts Crhr2-expressing neurons, leading to electrophysiological defects. At the behavioral level, this cell population is preferentially recruited in a stressful situation, and ablation of Prdm16 from this lineage leads to enhanced exploratory behavior even in the presence of a perceived threat.

      Strengths:

      The strengths of this manuscript are (i) leveraging a transcriptional regulator within a specific cell lineage and restricted to an early stage of ontogeny (ii) disrupting the developmental trajectory of a discrete neuronal population identified with an elegant snRNAseq approach and (iii) without obvious compensation (iv) and its impact on behavior in adult mice, with a special emphasis on exploratory drive in the presence of an acute stressor. The manuscript is well written; the experiments are well conducted, organized, and presented in a logical framework. Statistical analyses are well described. Each experimental group includes a sufficient number of subjects, allowing robust statistical comparisons.

      Overall, I very much enjoyed this manuscript and the elegant mouse model bridging developmental biology with systems neuroscience. Insights generated from this line of work could illuminate how discrete perturbations in gene expression programs at early stages of ontogeny could have a profound impact on the development, organization, and function of select neural circuits and how they may impinge on behavior at later stages of ontogeny.

      Weaknesses:

      Some comments and suggestions:

      (1) General-

      Photoinhibition of LS Crhr2-expressing neurons has no effect on anxiety-like behaviors in the absence of a stressor (Anthony et al., Cell, 2014). It would thus be interesting to reappraise the behavioral experiments performed with cKO mice in response to an acute stressor. The authors duly acknowledge this important point in the discussion section.

      (2) Specific-

      (a) Figure 2J: Was the increase in Crhr2 expression observed in tdTom-cells from cKO mice in the snRNAseq as well? If so, was Crhr2 expression enhanced in a specific cluster that did not belong to the Nkx2.1 lineage, or was it randomly enhanced across distributed clusters?

      (b) Figure 3 and S3: Does the lack of UCN3+/ENK+ terminals reflect a downregulation of UCN3 and ENK, or does it reflect the absence of innervation? Restricting a retrograde viral vector in iLS that expresses a fluorophore to illuminate the UCN3+ cell bodies (and lack thereof) in PefAH of cKO mice could address this question.

      (c) Figure 3: Does immunostaining for UCN3 in the PefAH area reveal cell bodies in cKO mice? In other words, is the loss of UCN3 terminal-specific or does it reflect a general downregulation of UCN3 in the PefAH?

      (d) Figure 4: The remaining tdTomato+ neurons are more excitable in cKO mice. To what extent can alterations in the electrophysiological properties of tdTomato+ neurons lacking Prdm16 be related to their survival? Is it a general response to Prdm16 deletion that is unrelated to survival? Is it a compensation mechanism in surviving cells? Or, alternatively, is it a unique property of these specific cells that favored their survival despite Prdm16 deletion?

      (e) Figure 5 and S5: Really nice figures. Great use of MoSeq with the predator odor test.

      (f) Figure 6: Interesting that the decrease in cFos induction in NeuN+ cells of cKO mice is more prominently observed in LSd when tdTomato+ cells are prominently found in the LSi/LSv (Figure S2B). Could this be related to intra-septal connectivity?

      (g) Figure 6: If tdTomato+ cells consist of 10-30% of neurons, and these tdTomato+ cells are preferentially found in the LSi/LSv, shouldn't we expect a decrease in c-Fos+NeuN+ density in the LSi/LSv (since there are generally fewer neurons in cKO mice)? If I am not mistaken, this could suggest that another unrelated LS population that is tdTom- displays an increase in cFos expression in cKO mice compared to WT mice. Could be interesting to see if the Crhr2+ neurons that are tdTom- are preferentially recruited in cKO mice as a compensation mechanism in LSi/LSv.

    4. Reviewer #3 (Public review):

      In the current work, Turrero Garcia et al. investigate the molecular, electrophysiological, and behavioral outcomes of conditionally knocking out cells from a unique developmentally-defined subpopulation in the lateral septum (LS). The authors focused on targeting cells from the Nkx2.1 developmental lineage that also expressed the transcriptional regulator Prdm16, which is uniquely upregulated in LS postmitotic neurons. They observed that this mutant line (Nkx2.1Cre;Prdm16fl/fl;Ai14; cKO) resulted in complete ablation of neurons from the Nkx2.1-lineage exclusively in the LS and not in the medial septum, making it an ideal model to study a lineage-defined subpopulation in the LS. They performed single-nucleus RNA-sequencing and uncovered four neuronal subtypes missing in the cKO. Furthermore, the authors validated this expression loss in the subtype expressing Crhr2 and observed a reduction of UNC-3 inputs (a neuropeptide with high affinity for Crhr2), highlighting additional disruptions in circuit connectivity. Loss of Prdm16 in Nkx2.1-lineage cells only resulted in mild electrophysiological changes in the LS. Finally, the authors performed a battery of behavioral assays to study anxiety-like behaviors and threat avoidance and observed increases in exploratory behaviors in some but not all assays. It should be noted that the cKO line also results in 30% loss of cortical interneurons, which could be contributing to the behavioral phenotype and not be exclusively due to LS loss. Overall, this manuscript takes a novel perspective by providing unique insights into how embryonic origin gives rise to mature molecular identity and distinct behaviors in adults. Therefore, it elegantly links developmental origin to mature molecular identity and function in the LS, an important question understudied in the field.

      The conclusions of the work are overall supported by the data and limitations discussed, but some of the findings, in particular the histological and behavioral results, need to be extended.

      Strengths:

      (1) Utilizing developmental origin as a marker for mature neuronal identity and function is a valuable approach which remains under-utilized in the field and serves to provide a deeper understanding of how circuits are shaped to allow for appropriate behavioral responses.

      (2) The authors perform a comprehensive analysis of the cKO mutant to determine the role of the developmentally defined Prdm16 in Nkx2.1-lineage cells at the molecular, cellular, electrophysiological, and behavioral level.

      (3) The sn-RNAseq dataset in the LS of WT and cKO mice will be valuable to the neuroscience community.

      (4) The authors perform an extensive array of behavioral paradigms investigating the balance between threat avoidance and exploratory behavior, performing all experiments in both male and female mice to determine whether the same developmental origin can lead to sex-specific differences.

      Weaknesses:

      (1) It remains unknown whether the reduction of UCN3 inputs to the LS is due to loss of the Nkx2.1-lineage in the LS itself or due to reductions in the number of UCN3 cells that provide innervation to the LS in the cKO (Figure 3). The authors speculate and include anecdotal observations that the perifornical region of the hypothalamus (PeFAH) provides inputs to the LS and could be the region driving the differences in UCN3 inputs in cKO. The authors should expand on this histological data and directly test whether the UCN3 inputs are indeed originating from PeFAH and whether the loss of Prdm16 in the Nkx2.1 lineage leads to a reduction in cell numbers in these inputs. These would disentangle the authors' claims on whether it is due to loss of Prdm16 in the Nkx2.1 lineage cells in the LS or whether it is due to loss of Nkx2.1 lineage neurons in upstream regions.

      (2) The authors claim that there is an increase in exploratory drive in cKO mice, even though the dark-light test showed increases in time spent in the dark side for cKO mice in comparison to controls (Figure 5C). They discuss that this could be due to the mice being placed first in the light compartment of the chamber during the light hours, so they spend more time exploring the dark compartment of the chamber instead, which would be considered 'novel'. If this were the case, to make the results more solid, authors should place cKO mice in the dark compartment of the chamber during the dark hours and then record time spent in both chambers. If there was indeed an increase in exploratory drive in new environments, the authors should see increases in time spent in the light compartment.

      (3) The result that cKO mice spend more time than controls exploring the inlet with TMT is of interest (Figure 5E, F). It needs to be highlighted that this is primarily the case in male mice and there is a trend in females. To confirm that the increases in exploration time of the inlet are not due to overall increases in general arousal, locomotion (e.g., velocity; pixels/frame) should be assessed in cKO vs controls.

      (4) Statistical analysis correcting for repeated testing should be performed when running multiple t-tests in the same dataset, such as when analyzing histological results in Figure 1 and Figure 6 to increase confidence in the presented results.

      (5) An important consideration that the authors address in the discussion is that the loss of Prdm16 in Nkx2.1-lineage cells is, for the most part, restricted to the LS, but other regions such as the cortex also show decreases in this population. Therefore, to strengthen the authors' conclusions that Nkx2.1-lineage neurons in the LS are indeed directly responsible for balancing threat avoidance and exploratory drive, targeted manipulation experiments, or at least additional c-Fos experiments assessing activity of Nkx2.1-derived cells in the LS will need to be performed in the future.

    1. eLife Assessment

      This study examines how prefrontal population dynamics tune social behaviors, providing important findings that oxytocin receptor-expressing interneurons support sociosexual choice in female mice. The multidisciplinary evidence is convincing and could be strengthened by more comprehensive reporting of the methodology and a clearer distinction between inferential and directly tested insights from modelling. This work will be of interest to those interested in prefrontal cortex, oxytocin, and/or social behavior.

    2. Reviewer #1 (Public review):

      Summary:

      Amadei et al investigate how excitation/inhibition balance in the prefrontal cortex plays a role in social behavior. To address this question, they developed a behavioral task where adult female mice can choose between a social reward (e.g., an adult male for sociosexual choice, or an adolescent female mouse) and a non-social reward (e.g., milk). They found that optogenetic inhibition of inhibitory neurons expressing oxytocin receptors (OXTR neurons) in the prefrontal cortex (PFC) reduces choice for sociosexual interaction compared to non-social reward and to a greater extent in sexually receptive females. They also found that this manipulation increases pyramidal neuron activity. Specifically, the authors identified a neuronal ensemble which represent the male option. Inhibition of OXTR disrupts the ability of the neuronal ensemble to represent the male option during decision-making in the behavioral task. Thus, using computational modeling, the authors proposed that OXTR neurons promote male choice by letting a male-representing pyramidal ensemble outcompete other pyramidal populations in the mPFC.

      Strengths:

      The study addresses an important topic in social behaviour and reward neuroscience with a focused hypothesis. The combination of behavioral testing and circuit manipulation combined with calcium imaging is a clear strength, and the work has the potential to make a solid contribution.

      Weaknesses:

      The main weaknesses are limited methodological clarity and details.

    3. Reviewer #2 (Public review):

      The authors aim to understand how inhibitory circuitry within the medial prefrontal cortex regulates the selection of sociosexual behaviour. Rather than studying social interaction in isolation, they develop an elegant behavioural paradigm in which female mice repeatedly choose between interacting with a male and obtaining an appetitive non-social reward. This task allows the authors to examine behavioural choice under conditions that more closely resemble natural decision-making. They combine optogenetic inhibition of oxytocin receptor-expressing interneurons, large-scale calcium imaging of pyramidal neurons, slice electrophysiology, and computational modelling to investigate how inhibition shapes cortical representations that ultimately bias behavioural choice.

      The study has several notable strengths. The behavioural paradigm is novel and well-designed, allowing repeated choice measurements while controlling for general social motivation by including both male and juvenile female stimuli. The integration of multiple experimental approaches is particularly impressive. The behavioural effects of optogenetic inhibition are complemented by population imaging demonstrating elevated pyramidal activity, electrophysiological recordings confirming monosynaptic regulation of pyramidal neurons, and a computational model that provides a mechanistic interpretation of the observed circuit dynamics. The work therefore spans multiple levels of analysis, from synaptic interactions to behaviour, and the individual datasets are generally of high technical quality.

      The imaging analyses identifying a putative "MALE" ensemble are particularly interesting. The observation that a relatively small subset of pyramidal neurons preferentially represents the male option before behavioural commitment provides an attractive framework for understanding how inhibition can stabilise specific behavioural representations. The temporal analysis suggesting that disruption of this representation precedes impaired behavioural choice is especially compelling, as it moves beyond simple correlations between neural activity and behaviour.

      Several aspects of the mechanistic interpretation remain somewhat speculative. The central conclusion relies heavily on the computational competition model, which assumes an asymmetric competition between a relatively small male-selective ensemble and a much larger default pyramidal population. While the model successfully reproduces several experimental observations, many of its architectural assumptions are inferred rather than experimentally demonstrated. In particular, the designation of the remaining pyramidal neurons as a functional "OTHER" population representing the non-social alternative is not directly established experimentally. Alternative circuit architectures may be capable of producing similar behavioural and population-level effects, and the current data do not fully distinguish among these possibilities.

      Similarly, although the identification of MALE cells is thoughtfully performed, the classification depends on an operational threshold derived from ROC analysis and correlated activity. It remains uncertain whether these neurons constitute a stable functional ensemble across sessions or merely reflect one end of a continuous representational spectrum. Longitudinal analyses examining the stability of these ensembles across days or across changes in behavioural state would strengthen the claim that they represent a dedicated neuronal population.

      An additional limitation concerns the specificity of the behavioural interpretation. The reduction in male choice is interpreted primarily as impaired sociosexual decision-making. While the inclusion of juvenile female stimuli substantially improves the experimental design, it remains difficult to completely separate altered sociosexual motivation from broader changes in motivational salience, valuation, or action selection. The observed changes could reflect alterations in multiple components of the decision-making process, and this distinction deserves a somewhat more balanced discussion.

      The interaction with the oestrous state is a very interesting aspect of the work and is consistent with previous studies of oxytocin-dependent sociosexual behaviour. However, this analysis is based on relatively modest numbers of animals and sessions, making it difficult to judge the robustness of these effects. The conclusions regarding hormonal modulation would therefore benefit from a more cautious interpretation.

      Overall, the authors achieve their primary objective of demonstrating that oxytocin receptor-expressing interneuron-mediated inhibition contributes to the selection of sociosexual behaviour while regulating pyramidal population dynamics in the medial prefrontal cortex. The behavioural, imaging, and electrophysiological datasets provide convincing evidence that inhibition shapes cortical activity during decision-making. The computational model offers a plausible mechanistic framework linking these observations, although some aspects of this framework remain hypothetical and await further experimental testing.

      The work is likely to have a significant impact on the fields of cortical circuit function, social neuroscience, and decision-making. Beyond its specific findings, the study introduces a behavioural paradigm that should prove broadly useful for investigating how competing behavioural options are represented within prefrontal circuits. The combination of behavioural neuroscience, population imaging, and computational modelling represents a valuable resource for the community and provides an important foundation for future studies examining how excitation-inhibition balance shapes flexible social behaviour.

    4. Reviewer #3 (Public review):

      Summary:

      Using a combination of Miniscope imaging and optogenetic manipulation, Amadei et al. reveal how oxytocin receptor neurons in the prefrontal cortex of mice control pyramidal subpopulations and socio-sexual behavior. This work was planned and executed carefully and provides a novel and important angle to study the oxytocin system in the cortex. According to their results, oxytocin receptor neurons help discriminate between sexual and non-sexual stimuli, most likely by controlling different pyramidal subpopulations that are either most active during trials that include a sexual stimulus or that include non-sexual stimuli. I highly appreciate this article; however, I have one major concern related to the modeling part.

      Strengths:

      (1) Well-designed experiments.

      (2) Rigorous analysis.

      (3) Generates a new avenue to study socio-sexual decision making and creates a hypothesis about the connectivity of oxytocin-sensitive circuits.

      Weaknesses:

      (1) Major

      In their last figure (Figure 4), the authors generated a computational model that, according to the authors, reveals a potential network mechanism in which oxytocin receptor (OXTR) neurons are connected to both pyramidal populations with certain connectivity rules. Although the model seems to reproduce the experimental results, some assumptions of the model seem to be poorly supported. If I understood correctly, the authors simply assumed that the strength of the connections between OXTR neurons and MALE neurons is the same as the strength of the connections between OXTR neurons and OTHER neurons. The authors neither discuss literature supporting such connectivity nor provide experimental evidence for this. I also could not find information about the magnitude of the synaptic weights to each of these populations. I guess these parameters are critical for the outcome of the simulation, and it may be worth exploring the outcome of simulating the different combinations of connectivity and synaptic weights between OXTR neurons and pyramids, as well as the degree of recurrent connectivity within the pyramidal subpopulations. Further, the authors should at least discuss in depth how inhibitory OXTR neuronal subtypes (they have different properties that could potentially be implemented in the modeling) may match their computational model best. If the current model remains the most promising, the authors should clearly discuss which experimental trajectory should be taken next to actually provide proof for its correctness (e.g. whether and how it would be possible to determine the predicted connectivity experimentally).

      (2) Minor

      The authors state regarding counterbalancing in Figure 1 and Figure S4C: "The social presentation order (male or female first), as well as the locations of the social and milk options (left or right relative to start arm), were fixed over sessions within a given subject, but varied over subjects (Figure S4C)". In Figure S4C, it looks as if there are fewer animals in which the male was always presented first than animals in which the female was always presented first (~ 16 vs. 20). While the difference is not very big, it may be influential. To fully exclude a sequence effect, I would suggest adding male-first animals until both groups are the same size.

    1. eLife Assessment

      This important paper provides evidence that visual representations, as observed through drawing, preferentially preserve topological features over metric features. Across three experiments with different age groups and creative converging methods, they demonstrate that holes and T-junctions are relatively preserved as compared to topologically irrelevant L-junctions, and Euclidian features like length or precise angle. Together, these data provide convincing evidence for topology as an organizing principle for spatial representation, but the work would be strengthened by considering a broader range of topological features, or by ruling out alternative accounts, like preserved memory for complex features, motor production limitations, or noisy compression, that may produce a similar pattern without a topological prior.

    2. Reviewer #1 (Public review):

      The paper presents novel evidence that spatial representations prioritize coarse topological features (T‑junctions, holes, crosses) over precise Euclidean metrics like angle and length, using drawing-based memory tasks with adults and children. The study is interesting and well‑motivated, and the importance of topological relations is clear, but stronger and more nuanced evidence is needed before concluding that topological relations are more important than metric details, as task difficulty and the potentially distinct roles of metric and topological information in spatial representation have not yet been fully disentangled.

      Introduction<br /> (1) P.5: Please explain in more detail what you mean by "What is relevant is the relative prioritization of each of these features."

      Results<br /> (2) P.8: Please clarify how the "proportion of drawings with angles biased towards 90{degree sign}" was computed. Specify the criterion for counting a drawing as biased (e.g., a certain absolute deviation toward 90{degree sign} from the original angle), and explicitly state in the Results that absolute degrees of deviation were used, as described in Methods.

      (3) It would help to spell out whether the findings imply that obtuse angles are typically drawn smaller (closer to 90{degree sign}) and acute angles larger (closer to 90{degree sign}). Also, would angles be more biased toward 90{degree sign} or 180{degree sign} (or 0{degree sign}) depending on the angle? (e.g., 175{degree sign} is seen more as 180{degree sign} while 95 is seen more as 90{degree sign})

      (4) Figure 4B: The statement that "positive values indicate bias in the direction of 90 degrees" needs a more precise explanation. Please explain exactly how the bias metric is computed (e.g., signed difference between drawn and original angle, with the sign indicating movement toward or away from 90{degree sign}) and what the y-axis values represent. Given that the Methods refer to absolute deviations, it would be useful to reconcile where the positive/negative signs come from in this plot.

      (5) Figure 4C: The description in the Results seems to use a different metric than what is plotted. Please ensure that the measure in the text matches the measure shown in the figure, and adjust labels or wording so they align clearly.

      (6) P.11: Consider briefly justifying why the authors predicted that participants would also add L‑junctions, rather than only remove them.

      (7) P.12: The last sentence: Weren't the overall rates of feature preservation 'higher' in the adult sample?

      Methods<br /> (8) Experiment 1: Please clarify whether the angles associated with T‑ and L‑junctions were equated or differed systematically. A short description of stimulus generation (e.g., angle ranges, line lengths, junction configurations) would be helpful.

      (9) It would also be helpful to specify the statistical tests used (e.g., t‑tests, ANOVAs, mixed‑effects models), including the main factors and any random effects, so readers can clearly follow your analysis pipeline.

      Discussion<br /> (10) It may be important to note that task difficulty likely differs across feature types: junctions involve presence/absence or counting, whereas angle and length reproduction require finer metric precision. The authors' claim of "prioritization" and possible difficulty effects should be disentangled.

      (11) Furthermore, would it be possible that people retain relative order/comparison of different angles/lengths rather than computing precise values?

      (12) I agree that topological relations are extremely important. However, for above reasons, it seems like stronger/stricter evidence is needed to claim that topological relations are 'more' important than metric details. They also might serve different roles in spatial representations

      (13) The Discussion would benefit from a short paragraph on where different junction types (T, L, crosses) typically appear in everyday scenes and objects (e.g., as cues to occlusion, surface intersections, 3D structure) and what functions they serve. This would help connect your experimental findings to the ecological importance of these features for natural vision and spatial cognition.

    3. Reviewer #2 (Public review):

      Summary:

      This is an interesting study that uses drawings to evaluate the extent to which visual representations of letter- and graph-like figures (preferentially) include topological features, like junctions and holes.

      The main claim is based on the observation that when participants are asked to draw presented figures from memory, they tend to (1) regularise angles towards 90deg and lengths towards the average length of the lines in the figure, while (2) preserving topological features like T-junctions more assiduously than non-topological features like L-junctions. A third experiment with 'serial reproductions' in which participants copy drawings made by other participants (like a visual version of the 'broken telephone' game) reproduce these patterns in exaggerated form. These findings were also reproduced in children (Experiment 4).

      These findings are consistent with the idea that memory representations are low-bandwidth or noisy approximations to the original figure. I would suggest that when participants are asked to reproduce the figure, it is if they combine the noisy stored representation, with generic priors about angles and the average line length. The preferential preservation of T- over L-junctions indicates that they are somehow more salient or memorable. This is not inconsistent with the authors' preferred interpretation of an explicit representation of topological structure. However, it is also not inconsistent with the idea that in order to compress the visual signals for storage, high-information (complex) components of the source are given preferential treatment. This would be compatible with optimal use of limited resources when compressing the information. Additional comparisons and control conditions would help tease these alternatives apart.

      Strengths:

      + Innovative use of drawing methods to probe internal visual representations<br /> + Experiments spanning both adults and children

      Weaknesses:

      - Failure to consider alternative hypotheses that are consistent with the findings

    4. Reviewer #3 (Public review):

      Kittur et al. ask whether human spatial memory is organized around topological relations (meaning coarse structural properties such as T-junctions, crosses, and holes) rather than around Euclidean properties such as angle and length. Across four experiments, adults and children studied letter-like figures and reproduced them from memory by drawing. The authors report two complementary patterns: metric features are systematically distorted, with angles pulled toward 90 degrees and line-length ratios compressed toward an average, while topologically critical features are comparatively well preserved. The central test contrasts T-junctions with L-junctions, which are visually similar but topologically distinct, since an L-junction reduces to a straight line, whereas a T-junction does not. A serial reproduction experiment amplifies both patterns across chains of participants, and a fourth experiment extends the findings to children aged five to eight.

      Strengths:

      The question is a good one and sits at a productive intersection of topics. It bears on debates about the representational format of cognitive maps, on proposals about the primitives of visual perception, and on a classic developmental claim from Piaget and Inhelder that has rarely been tested directly.

      The drawing paradigm is well chosen and offers something that the group's earlier forced-choice work could not. Because participants produce an open-ended response, distortion of metric detail and preservation of structure can be observed within a single response, and the relationship between them can be examined directly. The serial reproduction experiment is a particularly effective use of this affordance. The choice of the T-junction versus L-junction contrast as the primary test is well-motivated, since it holds the number of junctions constant and varies only topological relevance. It is also worth noting for readers that the central claim of a representational privilege for topologically distinct features was previously established by this group using forced-choice paradigms in both adults and children. That a similar conclusion emerges from free generation is a genuine strength, since the two methods have very different sources of error.

      The work is carefully executed. Sample sizes, dependent variables, and analyses were preregistered; stimuli were purpose-built for each question, including the deliberate exclusion of 90-degree angles so that no reference angle was available; drawings were double-coded; and the full set of raw drawings is being released publicly.

      Weaknesses:

      Drawing is treated as a transparent window onto representation, and motor limitations are not considered. Drawing is a motor act, drawing skill varies widely across individuals, and the manuscript does not discuss motor limitations at any point. As the study is designed, representational imprecision cannot be separated from difficulty of precise reproduction. The clearest way to resolve it might be asking adults to copy the figures exactly while the stimulus remains visible. If the biases persist under direct copying, then some portion of the effect is production rather than memory. Because the topological findings have already been demonstrated in keypress-only paradigms, this concern affects the metric distortion results most heavily, which are the novel contribution of the present paper.

      Three distinct claims are treated as one, and the data speak mainly to the weakest of them. The paper moves between a claim about mnemonic robustness (topological features survive degradation better than metric features), one about representational architecture (topology is a base layer with metric detail superimposed on top), and a claim about priority (topology is encoded prior to metric detail). The experiments show evidence for robustness, which is a claim about what is lost first. Robustness does not entail architecture: an encoder with a single layer, whose loss happens to spare structure, produces the same pattern with no layered format and no claim about encoding order. Earlier work does address format, because false "same" judgments to topologically matched but metrically different items show that topology plays a role in what the system treats as equivalent. Preservation counting measures robustness, rather than equivalence.

      Some alternative hypotheses to consider/address: (a) A capacity-limited memory that reconstructs from a prior produces the metric distortions with no commitment to topology. A literal absence of angle encoding, which the authors invoke, predicts noisy and unconstrained recall rather than recall pulled toward a particular value. The observed pattern reflects a structured prior. (b) The result that does discriminate might be confounded with local salience. A memory that adds uniform noise to all parts of a figure does not predict that T-junctions are preserved better than L-junctions; however, a three-way branch point is plausibly more locally distinctive than a corner, so a salience-weighted-but-topology-free account predicts the same ordering. (c) Motor simplification also predicts the same ordering, since omitting an L-junction converts a bend into a straight line, which is easier to draw, whereas omitting a T-junction requires dropping a stroke.

      The better a feature works as a topological marker, the less variance it produces and the harder it is to test, so the method is best powered where the theoretical signal is weakest. Holes are the textbook case of a topological invariant and are reported to disappear from drawings less than one percent of the time, but they are excluded from formal analysis because they are at ceiling and have no matched comparison feature. It would be useful to see bidirectional rates for holes (both how often a hole disappears and how often participants spuriously close an open figure into a loop).

      In Experiment 4, the conclusion of developmental stability rests on a nonsignificant effect of age, which is failure to detect a change rather than evidence of stability. An equivalence test or an estimate of the precision of the null is better support for developmental stability. Motor skill is confounded with age throughout. So this is an experiment that shows that the effect generalizes to childhood, but cannot adjudicate a developmental question.

      In the serial reproduction experiment, chains were intermixed so that each participant contributed one drawing to each of ten chains. The final drawings are therefore linked through shared intermediate participants, and an individual with an idiosyncratic drawing style influences ten chains at once, so the reported degrees of freedom are somewhat generous. The analysis also focuses on the final drawings and sets aside the nine hundred intermediate ones, which are the data that would show where in a chain metric detail collapses and whether structure ever breaks.

      To formalize the topology is to strengthen the argument: each figure is a one-dimensional complex (its underlying graph), treated intrinsically and up to homeomorphism. The homeomorphism type is what remains after suppressing all degree-2 vertices. Under this definition, every feature in this paper's taxonomy becomes one kind of object, namely a homeomorphism invariant of the graph: number of components, first Betti number, and the degree sequence of three or greater with its adjacency structure. Relatedly, the term "metric" needs to be unpacked. The 90-degree bias concerns angle, whereas the 4:2:1 result concerns length ratios, which are affine rather than strictly metric. The stronger statement available is that distortion appears at every level above topology in the transformation hierarchy while preservation occurs at the topological level, which connects directly to Chen's (2005) invariance hierarchy that is already cited.

      Appraisal and impact:

      The authors aimed to show that topological structure is preferentially retained in memory while metric detail is lost, and in the sense of relative preservation they succeed. The dissociation is real, replicates across two stimulus sets, amplifies under serial reproduction, and appears in young children. What the data do not establish is the stronger architectural claim that topology is a base representational layer, nor that the metric distortions specifically implicate topology rather than general properties of reconstructive memory. Separating these claims would communicate the well-supported result better. Conceptually, the work strengthens a growing case that coarse relational structure deserves a place alongside Euclidean properties in accounts of spatial representation. Practically, the public release of the full set of adult and child drawings, including excluded ones, is a resource that will support analyses well beyond those reported in this paper, and the serial reproduction design is a method that others will want to borrow.

    1. eLife Assessment

      This important study introduces a Bayesian method to determine bacterial counts that accounts for the experimental noise inherent to dilution and plating methods and distinguishes it from biological uncertainty. The evidence supporting the conclusions is compelling, combining simulated data and experimental data. The method will be of interest to microbial ecologists, and potentially to the broader community interested in inference from biological data, even more so if the domain of application and the limitations are further clarified.

    2. Reviewer #1 (Public review):

      Summary:

      The authors developed a novel theoretical/computational procedure to count bacterial populations without introducing artificial randomness effects due to dilution. Surprisingly, this very important aspect of studies of bacterial systems has been overlooked. The proposed method provides a simple and transparent approach to eliminate the randomness of bacterial accounting procedures, allowing now to fully concentrate on the intrinsic effects of the studied systems.

      Strengths:

      A very simple and clear procedure is introduced and explained in full detail. This elegant approach finds an excellent compromise between mathematical rigor and computational efficiency, which is important for practical applications. The provided examples are convincing beyond a doubt, clearly indicating the potential strong impact of the proposed framework. Various complications and possible issues are also discussed and analyzed. This seems to be a very powerful novel method that should significantly advance the analysis of complex biological systems.

      Weaknesses:

      The only minor weakness that I found is the assumption of independence of bacterial species, which is expressed as the well-stirred approximation. One could imagine that bacterial species might cooperate, leading to non-uniform distributions that are real. How to distinguish such situations?

      I believe that this method can be extended to determine if this is the case or not before the application. For example, if the bacteria species are independent of each other and one can use the binomial distributions - then the Fano factor would be proportional to the overall relative fraction of bacterial species. Maybe a simple test can be added to test it before the application of REPOP. However, I believe that this is a minor issue.

      Comments on revised version.

      I am satisfied with the correction proposed by the authors. The method is already quite impressive, and there is no need to complicate it at this stage.

    3. Reviewer #2 (Public review):

      I appreciate the thorough responses from the authors, which address my concerns. The expansion of Appendix B as well as the addition of text discussing how the REPOP method interacts with data collection efforts are very useful. These new sections show that relative error decreases with increasing samples, as expected, yet these error metrics, including KL divergence, describing the fit of the full distributions not just the modes, drop off fairly quickly with increasing number of samples showing that REPOP likely minimizes discrepancies between estimated and true distributions even at lower sampling efforts.

      Additionally, the extension of the REPOP method to the quantification of multiple bacterial species or phenotypes shows the potential utility of the method in contexts beyond basic plate counts. Between this example and the additional information on how to implement REPOP, I believe this workflow will be attractive and accessible to the audience.

    4. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      (R1C1) The only minor weakness that I found is the assumption of independence of bacterial species, which is expressed as the well-stirred approximation. One could imagine that bacterial species might cooperate, leading to non-uniform distributions that are real. How to distinguish such situations?

      I believe that this method can be extended to determine if this is the case or not before the application. For example, if the bacteria species are independent of each other and one can use the binomial distributions, then the Fano factor would be proportional to the overall relative fraction of bacterial species. Maybe a simple test can be added to test it before the application of REPOP. However, I believe that this is a minor issue.

      This is an interesting point raised by the reviewer.

      First, we need to clarify an important point: we do not make a well-stirred assumption. Samples can be drawn and plated from any region of space however small and that region’s population can be quantified using our method. The stirring only occurs after we collect a sample in order to dilute the contents and pour the solution homogeneously over the plate.

      As such, learning multiple independent species is possible and not impacted by the dilution (“well-stirred” assumption). In the new first paragraph of the methods section, we made it clear that this assumption concerns the dilution process. REPOP is designed to recover the true underlying heterogeneity in species abundance (even from limited data) by leveraging a Bayesian framework that remains valid regardless of whether species are independent or correlated.

      If the method is applied to multiple species as currently implemented, REPOP can recover the marginal distribution of each species, provided that the species are either selectively cultured or produce sufficiently distinguishable colonies on the same plate. To demonstrate this, we have added a new Results subsection with a synthetic two-species example in which the species abundances are correlated across samples.

      However, in order to learn the joint distribution and capture correlations between species within samples, the method would need to be extended. At present, in Eq. 5 we sum the likelihood over all values of n, using a data-driven cutoff (twice the largest naïvely estimated count times the dilution factor). Extending this to multiple species adding up to (n<sub>1</sub>,n<sub>2</sub>), while retain the generality of the method, would require quadratically scaling memory with this cutoff in the population number. For this reason while we comment on this in the new paragraph in the conclusion, it is not implemented as part of REPOP.

      Reviewer #2 (Public review):

      (R2C1) A more thorough discussion of when and by how much estimated microbial population abundance distributions differ from the ground truth would be helpful in determining the best practices for applying this method. Not only would this allow researchers to understand the sampling effort necessary to achieve the results presented here, but it would also contextualize the experimental results presented in the paper. Particularly, there is a disconnect between the discussion of the large sample sizes necessary to achieve accurate multimodal distribution estimates and the small sample sizes used in both experiments.

      That is a great suggestion from the reviewer. To address it, we expanded Appendix B. We know report (1) the relative error in the estimated means (as already done for Fig. 4 formally 3), and (2) the Kullback-Leibler (KL) divergence between the reconstructed and ground-truth distributions. These metrics will are show as a function of the size of the dataset, for the examples in Fig 3. enabling a direct assessment of how the sampling effort affects the precision of the inference.

      That said, we now highlight in the Conclusion that, by explicitly modeling the dilution process within a Bayesian framework, REPOP extracts the maximum information available from each individual sample at a given sample size. This strategy therefore enables more accurate inference with fewer measurements, which is particularly important in applications such as plate counting, where data acquisition is labour-intensive.

      Reviewer #3 (Public review):

      (R3C1) While the study is promising, there are a few areas where the paper could be strengthened to increase its impact and usability. First, the extent to which dilution and plating introduce noise is not fully explored. Could this noise significantly affect experimental conclusions? And under what conditions does it matter most? Does it depend on experimental design or specific parameter values? Clarifying this would help readers appreciate when and why REPOP should be used.

      We agree with the reviewer that this is an important point, and we expanded Appendix B to include a quantitative analysis using simulated data (Fig. 3, formely 2), reporting both relative error and KL divergence as a function of dataset size. This complements our response to R2C1 clarifying when REPOP offers the greatest benefit.

      In addition, we will expand the discussion on how modeling dilution noise becomes essential when learning population dynamics. In particular, we emphasize? the role of Model 3, especially relevant when working with multiple plates and approaching the asymptotic regime; an aspect that was alluded to in Fig. 3 but not fully explored.

      (R3C2) Second, more practical details about the tool itself would be very helpful. Simply stating that it is available on GitHub may not be enough. Readers will want to know what programming language it uses, what the input data should look like, and ideally, see a step-by-step diagram of the workflow. Packaging the tool as an easy-to-use resource, perhaps even submitting it to CRAN or including example scripts, would go a long way, especially since microbiologists tend to favor user-friendly, recipe-like solutions.

      In the new paragraphs of the introduction, we made clear that REPOP is written in Python (PyTorch), installable via pip, and designed for ease of use. We are also expanding the tutorials to include clearer guidance on data formatting and common workflows. The new workflow figure (Fig 2) better illustrates the full process.

      (R3C3) Third, it would be great to see the method tested on existing datasets, such as those from Nic Vega and Jeff Gore (2017), which explore how colonization frequency impacts abundance fluctuation distributions. Even if the general conclusions remain unchanged, showing that REPOP can better match observed patterns would strengthen the paper’s real-world relevance.

      We thank the reviewer for this interesting suggestion. We agree that applying REPOP to additional existing datasets would make REPOP’s relevance clearer. However, the Vega and Gore datasets lack the information required. REPOP requires the plate count measurement process to be specified, including the dilution factors used for each measurement. Furthermore, we can leverage on additional information about the experimental procedure when the colony cutoffs and dilution schedules used are reported. Without the dilution factors, the likelihood connecting the observed colony counts to the underlying population size is not possible. We hope this clarification will help make future datasets made available publicly more useful for purposes of uncertainty propagation.

      (R3C4) Lastly, it would be helpful for the authors to briefly discuss the limitations of their method, as no approach is without its constraints. Acknowledging these would provide a more balanced and transparent perspective.

      We agree with the reviewer. We have added two new paragraphs to the conclusion highlighting important current constraints and future development directions of the framework. In particular, we now discuss that, in its present implementation, REPOP focuses on the population distribution that maximizes the posterior, rather than returning posterior uncertainty over the reconstructed distributions themselves. We also note the computational demands of the method, making GPU acceleration highly beneficial and more complex multi-population inference computationally challenging. This discussion synthesizes points raised throughout our response to R1C1 and the reviewers and provides a more balanced perspective on the current scope of the method.

    1. eLife Assessment

      This valuable study examines how the prelimbic cortex represents learned and generalized threat over time and identifies potentially distinct stable and dynamic subnetworks that may support these functions. The work is conceptually interesting and is strengthened by the longitudinal calcium imaging approach and the inclusion of key control groups. However, the evidence supporting the claims is incomplete, particularly because the interpretations regarding inference, time-dependent representational change, and the dissociation of neural activity from freezing behavior extend beyond what is currently established by the data.

    2. Reviewer #1 (Public review):

      Summary:

      The authors combine discriminative auditory fear conditioning with longitudinal in vivo calcium imaging to ask how prelimbic (PL) representations of learned and generalized threat evolve across recent and remote memory time points. Using two different CS+ frequencies and a no-shock control group, they report that PL population activity tracks graded behavioral generalization, that population similarity is highest for tones eliciting strong threat responding, and that distinct subnetworks can be identified that appear to encode tone-specific sensory features versus learned threat-related response structure.

      To my knowledge, this may be the first study to comprehensively examine neural encoding of fear generalization in prelimbic cortex (PL). The manuscript is ambitious and technically interesting, and several aspects are potentially important. In particular, the suggestion that neurons showing graded, learning-related response patterns become selectively stabilized over time is intriguing. The inclusion of two CS+ training conditions and a no-shock control also strengthens the case that at least some of the reported effects are related to associative learning rather than simple sensory differences. However, in its current form, the manuscript does not yet fully support the strength of the conceptual claims. Several issues limit confidence in the interpretation, including the possibility that repeated testing itself contributes to changes across days, uncertainty about the relationship between neural activity and freezing behavior, limited quantitative documentation of longitudinal cell registration, and a number of problems in figure clarity and statistical framing. Overall, the study contains promising observations, but the claims should be narrowed, and several analyses or controls would be needed to fully support the proposed framework.

      Detailed Comments

      (1) A general concern is that the repeated test procedure itself may contribute to extinction. Because the animals are exposed to multiple CS frequencies across multiple test days, and each tone is presented three times per session, some of the reported changes in behavior and neural activity across days could reflect extinction or repeated nonreinforced retrieval rather than the passage of time per se. This is especially relevant given that the manuscript makes claims about recent versus remote representations and representational drift over 30 days. At a minimum, the authors should discuss this limitation explicitly and temper claims about time-dependent changes. Ideally, they would include a control group in which animals are tested only once or twice (e.g., at an early and later time point with fewer CS frequencies), or a reduced-frequency testing design that minimizes extinction while still allowing evaluation of recent versus remote memory.

      (2) More generally, some of the reported learning-related neural differences may be driven by behavioral differences, particularly freezing, rather than by learning or generalization per se. For example, animals that freeze more to certain frequencies may show corresponding neural response differences simply because freezing alters PL activity. The authors should examine this possibility more directly. Analyses testing whether recorded cells encode freezing behavior, or whether tone frequency-related neural differences remain robust when comparing high- and low-freezing epochs, would help determine whether the reported effects reflect learned stimulus value rather than behavioral state differences.

      (3) A central feature of the manuscript is the analysis of neural response properties over an extended period of time, up to 30 days after learning. However, aside from a brief mention in the Methods that spatial registration was used, the manuscript provides very little quantitative information about this critical aspect of the study. The paper would be strengthened by including explicit metrics describing longitudinal cell tracking, such as the number and proportion of ROIs retained across all sessions, distributions of spatial-footprint correlations or centroid distances across days, and representative examples of matched imaging fields over time. Without this information, it is difficult to assess how strongly the longitudinal claims are supported.

      (4) The text states that "Figs. 1c and 1d show GCaMP6f expression in PL, representative calcium footprints, and activity traces". However, the figure as presented does not clearly show all of these elements, at least not in a way that matches the description in the Results. The correspondence between text and figure should be corrected.

      (5) The labeling of Figure 2a is insufficient for interpretation. The legend states that the panel shows raster plots of sound responsiveness, but the axes and scaling are not clearly defined. It is not clear from the figure what the x-axis represents, whether the y-axis corresponds to individual neurons, where the CS period occurs, or what the activity scale at the right denotes. Also, the term 'rasters' implies that spikes were analyzed. It seems that the spike inference approach (CASCADE) was only used for later analyses. Perhaps 'heat-plot' would be more accurate here? Generally, this figure should be annotated more clearly so that the reader can understand it without referring back to the Methods.

      (6) In relation to Figure 3, the analysis of population-averaged responses across tone frequencies is useful, but the manuscript would be stronger with additional statistical analyses across time and across groups. For example, if the authors want to argue that learning induces graded changes in neural responses and that these evolve across time, they should directly compare within-group responses across days and also compare matched frequencies between the conditioned groups and the no-shock controls. These analyses would help establish whether the observed differences are genuinely learning dependent and whether they change significantly over time.

      (7) The inclusion of two different CS+ frequencies and a no-shock control is a strength of the study and substantially improves the interpretation that graded neural responses are related to learning and generalization rather than to simple sensory processing or passage of time. That said, I am not entirely comfortable with the use of the term "inference" throughout the manuscript. What is being measured here appears closer to sensory generalization than inference in a stronger cognitive sense. The current task does not clearly require that animals infer hidden structure or stimulus value through abstract reasoning; rather, the generalized stimulus may simply be treated as similar to the conditioned cue. The terminology should therefore be reconsidered or softened.

      (8) I also found the use of the term "valence" somewhat problematic. The manuscript appears to use valence to refer to graded responding across tones with different aversive significance, but valence typically refers more broadly to distinctions between appetitive and aversive value. Here, terms such as "threat value," "aversive value," may be more precise. The authors should consider revising this language throughout.

    3. Reviewer #2 (Public review):

      Summary:

      The following points are those that occurred to me across readings of the paper. They are listed in what I take to be the order of their significance. Many of the points relate to the loose use of language and invocation of concepts that are not warranted, given the study design and results obtained.

      Major Comments:

      (1) The concept of ensemble turnover is interesting - the way it is introduced and discussed implies some type of spontaneous change in the neural underpinnings of fear discrimination and generalization in the PL. But, of course, every trial involves an opportunity to learn about the threat CS or the generalization test stimuli, and I am troubled by the thought that stability in the neural underpinnings of fear discrimination and generalization will actually reflect the level of defensive behaviours evoked on different trial types and/or the discrepancy between those behaviours and the outcome of a given trial in the generalization test. That is, stability in the neural underpinnings may be related to an animal's certainty or uncertainty in the contingency between a stimulus and danger; or, put another way, an animal's confidence that danger will or won't occur given the presence of some stimulus. This is not uninteresting. It is, however, not considered anywhere in the paper, which is overloaded with references to inferred threat values and integration of information across different types of stimuli. The protocol is not one that requires inference about anything or integration across anything.

      (2) I appreciate the link to Gu and Johansen in paragraph 3 of the Introduction, but the type of generalization under investigation here is not the same as the type of 'generalization' studied by Gu and Johansen [who used a sensory preconditioning protocol]. Nonetheless, the authors have forced the language used by Gu and Johansen into their paper, and this has created tension [at least for this reader] as the concepts introduced by Gu and Johansen [inference, integration] are simply not relevant given the generalization protocol used here. Here are a few examples of points where the tension might interfere with a reader's understanding:

      a. 'We hypothesized that generalization to novel stimuli depends on stable subnetwork organization that enables comparisons between learned and inferred valence, as well as population-level features that reduce variability across related representations.'

      I understand the words in the hypothesis, but can't form a representation of what is being said because of the reference to terms that stand in need of clarification [inferred valence, variability across related representations], but, ultimately, won't be clarified. This needs to be re-expressed so that the reader can appreciate what is being said.

      b. 'Our results show that stable cortical subnetworks integrate the emotional "gist" of memory and inferred valence for novel cues over time, despite ongoing ensemble reorganization, and that population-level firing rate similarity across stimulus presentations determines threat generalization.'

      Again, what does this mean? How is the gist of a memory integrated with inferred valence for novel cues over time? The statement simply doesn't make sense. This needs to be rewritten for clarity.

      c. 'In CS⁺15 mice, positively modulated sound-responsive neurons exhibited graded tone activity reflecting the contingency learned valence as well as the inferred valence of novel tones across testing days...'.

      Can this be rewritten as 'In CS⁺15 mice, positively modulated sound-responsive neurons exhibited graded activity to the tone CS and its variants that were used to assess generalization.'? The overloading of the text with references to 'contingency learned valence' and 'inferred valence' is unnecessary and makes it much harder to understand what has been shown in the results.

      (3) Re the same passage of text as in 2c:

      Is it the case that these neurons are simply tracking the expression of freezing to the various tones? The same question applies to the results obtained for the CS+3 mice. If this is the case, then why should the results be taken to support the banner statement that 'Sound-modulated PL population responses encode learned and inferred valence' - these analyses do not support that statement. And, as indicated, I don't believe that the language of learned and inferred valence is appropriate to such statements, given the nature of the protocol used and results obtained. It is a study looking at how populations of neurons in the PL respond during presentations of auditory stimuli that were subject to discriminative conditioning, and during tests of generalized freezing to other [intermediate] auditory stimuli.

      (4) It is stated that:

      'In no-shock controls, although both positive and negative responses were present, population activity was not modulated by tone frequency or valence'.

      What does this mean? I can understand that population activity was not modulated by tone frequency. But what does it mean to say that it was not modulated by valence? Why should it have been when none of the tones were conditioned in this group and, hence, mice were responding to all the tones equally? And given that this is true, I don't understand the use of 'valence' here, or the subsequent statements in this paragraph that 'graded responses require associative learning' and that 'PL population responses encode graded sound-valence associations that reflect both learning and inference, closely matching behavioral generalization.' The latter statement is particularly unwarranted and, again, highlights a major issue with the paper. It could and should be rewritten as 'PL population responses reflect behavioral generalization.' There is nothing in the additional language that adds to the reader's understanding of what has been shown. The reference to 'graded sound-valence associations that reflect both learning and inference' is completely unwarranted, given the nature of this study. It is anathema to the vast literature on stimulus generalization. If the authors wished to make statements of this sort, they should have taken a different approach, perhaps using protocols like those featured in Gu and Johansen.

      (5) The section titled, 'Consistently active neurons preserve valence representations as newly recruited neurons sharpen remote memory traces' ends with the following summary:

      'Together, these results indicate that consistently active neurons maintain stable representations of learned and inferred sound associations across time, whereas neurons recruited after conditioning progressively acquire graded tuning at later retrieval stages. This dynamic refinement suggests that cortical memory representations become increasingly selective during systems consolidation, while a stable neuronal subpopulation preserves the core emotional content of the memory.'

      Once again, the summary is not in keeping with the results obtained. The 'dynamic refinement' of representations is far more likely to reflect the repeated testing across days 1, 15, and 30 rather than anything to do with systems consolidation - at the very least, it is the simplest interpretation of the results. The impact of repeated testing is evident in the sharpening of generalization gradients over time, which is contrary to what is otherwise observed in the literature - the incredibly well -documented broadening of generalization gradients with time. Given this impact of repeated testing, surely the changes in the neuronal population that underlie performance are more likely to reflect the learning that occurs on days 1, 15, and 30, which is reflected in reduced freezing to the non-conditioned tones. If this is a reasonable take on the results, then I don't see the basis for invoking systems consolidation at all, and I don't see the basis for inferring a stable neuronal subpopulation that preserves the emotional content of the memory. Rather, non-reinforced presentations of 'never-reinforced' tones result in recruitment of additional neurons that result in suppression of freezing responses to those stimuli.

      (6) In the section titled, 'Population vector similarity at stimulus onset determines degree of generalization', it is stated that:

      'Because population similarity peaked shortly after stimulus onset, we quantified similarity during the first 5 s after tone onset relative to the CS⁺. In CS⁺15 mice, population similarity was highest for 15/15 and 15/11 tone pairs with no differences between them.'

      Isn't this consistent with the view that the population response in the PL simply reflects the level of freezing? Freezing to the 15-15 and 15-11 tones is most likely to be similar on their first presentation prior to the effects of extinction on the 11 Hz tone; hence the results obtained. That is, these results appear to clearly indicate that neuronal responses in the PL reflect the degree of stimulus generalization, as evidenced in freezing behavior. Given all that we know about the involvement of the PL in expressing fear responses, it is not appropriate to claim that 'population vector similarity at stimulus onset *determines* the degree of generalization. The PL responses simply reflect the varying levels of performance displayed to the different types of tones. What have I missed that could be taken to support additional statements?

      Later in the same section, it is stated that 'population-level similarity at stimulus onset scales with behavioral threat generalization and is maximal for tones associated with robust threat responses.' For simplicity and, therefore, clarity, this should be rewritten as 'population-level similarity at stimulus onset reflects behavioral threat generalization.'

      (7) In the section titled, 'Different subnetworks encode acoustic versus learned properties of sound association', it is stated that:

      'Our previous analyses show that learned and inferred associations are represented at the population level. However, these results do not resolve whether graded responses arise from pooled activity of frequency-selective neurons or from subnetworks encoding integrated learned valence across tones.'

      What does it mean to say 'integrated learned valence across tones'? As it presently stands, the meaning of the phrase is unclear. It only makes sense if one supposes that generalized freezing responses to the 11 and 7 kHZ tones reflect separate associations between those tones and the aversive foot shock US. This supposition is inconsistent with the rich literature on generalization of Pavlovian conditioned fear responses. Specifically, it is inconsistent with the many theories of fear generalization, which attribute the reduction in fear as one moves away from the specific conditioned stimulus to a decrement in the ability of the test stimulus to activate the trained CS-US association. My strong impression is that the authors would do well to ground their findings in theories of stimulus/fear generalization, of which there are many. This would better serve the results obtained [and the reader's appreciation of them] - at present, the unnecessary invocation of concepts does very little to enhance the reader's appreciation or understanding of what has been found in the study.

      (8) Another example of what has been a common theme in this review :

      '...we hypothesized that the PL active ensemble segregates into functionally distinct subnetworks: one encoding tone-specific sensory features with dynamic characteristics, and another responding to all frequencies encoding stable core memory content and inferred emotional valence.'

      What does it mean to say 'all frequencies encoding stable core memory content and inferred emotional valence'? Do the authors mean to say '...and another that tracks freezing/defensive responses regardless of whether they were elicited by the trained CS or one of the generalization test stimuli'?

      (9) It is stated that - 'Graded clusters encode emotional valence but constitute only a fraction of the active population; yet valence coding at the population level remains accurate and precise. This indicates that neurons newly recruited into the population-likely frequency-selective and organized within learning-independent clusters-can be shaped by associative processes through modulation of firing activity.'

      What does this mean? Are the authors trying to say that - 'Some clusters of PL neurons track freezing responses. In spite of the fact that these are only a fraction of the total active neuronal population, the population-level response of PL neurons also tracks the levels of fear to the trained tone and its variants used in the test for generalization.' If this is what one wants to say, then the final statement in the reproduced section does not follow. That is, there is no indication that 'neurons newly recruited into the population-likely frequency-selective and organized within learning-independent clusters-can be shaped by associative processes through modulation of firing activity.' As noted, the characteristics of other ensembles that become active across the repeated tests on days 1, 15, and 30 are more likely to reflect learning from non-reinforcement that occurs within and across those sessions. Perhaps this is what is meant by the phrase, 'shaped by associative processes'? If so, it should be stated explicitly instead of left to the reader to work out.

      (10) The following points all relate to the Discussion and reiterate many of the points above.

      a. 'A subset of neurons remains consistently active across sessions, preserving core components of the memory trace and supporting inference of emotional valence for novel sounds, while neurons recruited after conditioning progressively acquire valence selectivity at remote time points.'

      'Inference of emotional valence' is unclear and unwarranted for all of the reasons provided above regarding the use of language.

      b. '...Our data reconcile these views by demonstrating that cortical representations of emotional valence emerge rapidly after learning and persist within stable subnetworks, even as the broader population undergoes substantial turnover. This architecture preserves core mnemonic content while allowing flexibility in the surrounding ensemble.'

      These statements assume that the PL neuronal responses reflect something more than the levels of freezing behavior to the different stimuli; what are the grounds for this assumption?

      c. 'Importantly, these subnetworks encode both learned contingencies and the inferred valence of novel stimuli along a graded representational axis, suggesting that strong recurrent connectivity provides a stable scaffold for emotional memory representations.'

      What is a graded representational axis, and what part of the first statement suggests that 'strong recurrent connectivity provides a stable scaffold for emotional memory representations'? If the authors' goal was to make statements about emotional memory representations vis-à-vis emotional memory content, they should have used protocols that allowed them to probe such content. The auditory fear conditioning protocol used here [followed by tests for generalization to other auditory stimuli that differ in frequency from the conditioned tone] is not one that lends itself to analysis of emotional memory representations or content.

      d. 'Dynamic tone-selective responsive neurons emerge independently of learning, as they are present in both control and experimental mice, reflecting pre-existing PL sensory-driven properties (Hockley & Malmierca, 2024; Zikopoulos & Barbas, 2006).'

      Maybe. They are also likely to have developed as a consequence of the repeated testing on days 1, 15, and 30, which involved intermixed exposures to the tones of different frequencies. That is, rather than 'pre-existing PL sensory-driven properties', the responses of these neurons might reflect the emergence of discrimination between the various tones across testing, and greater suppression of freezing to the non-trained tones compared to the trained tone across the various test intervals.

    4. Reviewer #3 (Public review):

      Summary:

      Normandin et al. explore the coding of stimuli predicting an aversive event in the prelimbic cortex. Stimuli could either be explicitly paired, explicitly unpaired, or novel but with an inferred association with the aversive event (generalization). Long-term tracking of GCaMP-positive neurons allowed them to examine how coding evolves out to a month following training. In general, they found two types of ensemble codes. One was ensembles coding for each stimulus independently, but with enhanced responding to the one eliciting a freezing response. The other was ensembles that responded to all stimuli in proportion to their similarity to the stimulus paired with the aversive event, either increasing or decreasing their activation with the degree of freezing elicited by a stimulus. Importantly, this second set of ensembles was more stable across days, potentially providing a memory trace.

      Strengths:

      (1) The authors track ensembles in prelimbic cortex over long time scales, providing valuable information on the consolidation of neural codes.

      (2) Neural coding of generalization is examined, which is under-examined in the field.

      Weaknesses:

      (1) Difficult to determine if responses treated as encoding stimulus valence are driven instead by the behavior that the stimulus elicits, freezing.

      (2) The study implies that the identified ensembles are causally related to valence memory, but no experimental interventions are performed to justify this.

    5. Author response:

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The authors combine discriminative auditory fear conditioning with longitudinal in vivo calcium imaging to ask how prelimbic (PL) representations of learned and generalized threat evolve across recent and remote memory time points. Using two different CS+ frequencies and a no-shock control group, they report that PL population activity tracks graded behavioral generalization, that population similarity is highest for tones eliciting strong threat responding, and that distinct subnetworks can be identified that appear to encode tone-specific sensory features versus learned threat-related response structure.

      To my knowledge, this may be the first study to comprehensively examine neural encoding of fear generalization in prelimbic cortex (PL). The manuscript is ambitious and technically interesting, and several aspects are potentially important. In particular, the suggestion that neurons showing graded, learning-related response patterns become selectively stabilized over time is intriguing. The inclusion of two CS+ training conditions and a no-shock control also strengthens the case that at least some of the reported effects are related to associative learning rather than simple sensory differences. However, in its current form, the manuscript does not yet fully support the strength of the conceptual claims. Several issues limit confidence in the interpretation, including the possibility that repeated testing itself contributes to changes across days, uncertainty about the relationship between neural activity and freezing behavior, limited quantitative documentation of longitudinal cell registration, and a number of problems in figure clarity and statistical framing. Overall, the study contains promising observations, but the claims should be narrowed, and several analyses or controls would be needed to fully support the proposed framework.

      Detailed Comments

      (1) A general concern is that the repeated test procedure itself may contribute to extinction. Because the animals are exposed to multiple CS frequencies across multiple test days, and each tone is presented three times per session, some of the reported changes in behavior and neural activity across days could reflect extinction or repeated nonreinforced retrieval rather than the passage of time per se. This is especially relevant given that the manuscript makes claims about recent versus remote representations and representational drift over 30 days. At a minimum, the authors should discuss this limitation explicitly and temper claims about time-dependent changes. Ideally, they would include a control group in which animals are tested only once or twice (e.g., at an early and later time point with fewer CS frequencies), or a reduced-frequency testing design that minimizes extinction while still allowing evaluation of recent versus remote memory.

      We agree with the reviewer that repeated testing is an inherent limitation of longitudinal memory studies and may itself contribute to some neural changes across sessions. However, several aspects of our behavioral design and results argue against extinction or repeated nonreinforced retrieval as the primary drivers of the observed effects. Importantly, discrimination ratios remained stable or increased across time rather than progressively diminishing as would be expected under extinction (this new analysis will be added to the resubmission). Nevertheless, we will address this important point in the Discussion and explicitly acknowledge that repeated retrieval may contribute to some component of the observed representational changes.

      (2) More generally, some of the reported learning-related neural differences may be driven by behavioral differences, particularly freezing, rather than by learning or generalization per se. For example, animals that freeze more to certain frequencies may show corresponding neural response differences simply because freezing alters PL activity. The authors should examine this possibility more directly. Analyses testing whether recorded cells encode freezing behavior, or whether tone frequency-related neural differences remain robust when comparing high- and low-freezing epochs, would help determine whether the reported effects reflect learned stimulus value rather than behavioral state differences.

      We thank the reviewer for raising this important point, which was also noted by the other reviewers. To address this issue, we will implement Reviewer 3’s suggested Generalized Linear Model (GLM) analysis using inferred spiking activity derived from the Ca2+ signals, with both tone identity and freezing behavior included as predictors. Because freezing behavior varies across trials whereas stimulus identity is fixed, this approach will allow us to dissociate their respective contributions to neuronal activity. If, after accounting for freezing behavior, responsive neurons continue to exhibit graded coding consistent with inferred threat value, this would strengthen the interpretation that the identified ensembles reflect generalization gradients related to aversive value rather than freezing behavior alone. Otherwise, we will adjust the conclusions according to the interpretation that freezing itself drives the generalization gradients.

      (3) A central feature of the manuscript is the analysis of neural response properties over an extended period of time, up to 30 days after learning. However, aside from a brief mention in the Methods that spatial registration was used, the manuscript provides very little quantitative information about this critical aspect of the study. The paper would be strengthened by including explicit metrics describing longitudinal cell tracking, such as the number and proportion of ROIs retained across all sessions, distributions of spatial-footprint correlations or centroid distances across days, and representative examples of matched imaging fields over time. Without this information, it is difficult to assess how strongly the longitudinal claims are supported.

      We thank the reviewer for this suggestion. We will include measures of registration quality in the resubmission.

      (4) The text states that "Figs. 1c and 1d show GCaMP6f expression in PL, representative calcium footprints, and activity traces". However, the figure as presented does not clearly show all of these elements, at least not in a way that matches the description in the Results. The correspondence between text and figure should be corrected.

      We will correct correspondence between text and Figure.

      (5) The labeling of Figure 2a is insufficient for interpretation. The legend states that the panel shows raster plots of sound responsiveness, but the axes and scaling are not clearly defined. It is not clear from the figure what the x-axis represents, whether the y-axis corresponds to individual neurons, where the CS period occurs, or what the activity scale at the right denotes. Also, the term 'rasters' implies that spikes were analyzed. It seems that the spike inference approach (CASCADE) was only used for later analyses. Perhaps 'heat-plot' would be more accurate here? Generally, this figure should be annotated more clearly so that the reader can understand it without referring back to the Methods.

      Thank you for this suggestion. We will clarify the labelling of the Figure 2a and call the graphs “activity-plots”.

      (6) In relation to Figure 3, the analysis of population-averaged responses across tone frequencies is useful, but the manuscript would be stronger with additional statistical analyses across time and across groups. For example, if the authors want to argue that learning induces graded changes in neural responses and that these evolve across time, they should directly compare within-group responses across days and also compare matched frequencies between the conditioned groups and the no-shock controls. These analyses would help establish whether the observed differences are genuinely learning dependent and whether they change significantly over time.

      We will redo the Statistics of Figure 3 to take into account the following variables: group (CS15, CS3, no shocks), frequency (3, 7, 11, 15), and day of testing (2, 15, 30).

      (7) The inclusion of two different CS+ frequencies and a no-shock control is a strength of the study and substantially improves the interpretation that graded neural responses are related to learning and generalization rather than to simple sensory processing or passage of time. That said, I am not entirely comfortable with the use of the term "inference" throughout the manuscript. What is being measured here appears closer to sensory generalization than inference in a stronger cognitive sense. The current task does not clearly require that animals infer hidden structure or stimulus value through abstract reasoning; rather, the generalized stimulus may simply be treated as similar to the conditioned cue. The terminology should therefore be reconsidered or softened.

      We thank the reviewer for appreciating the strengths of the experimental design and for this thoughtful suggestion regarding terminology. We agree that the term “inference” may overstate the cognitive processes engaged by the current task. Accordingly, we will revise the terminology throughout the manuscript to describe these effects as graded generalization of threat value across stimuli.

      (8) I also found the use of the term "valence" somewhat problematic. The manuscript appears to use valence to refer to graded responding across tones with different aversive significance, but valence typically refers more broadly to distinctions between appetitive and aversive value. Here, terms such as "threat value," "aversive value," may be more precise. The authors should consider revising this language throughout.

      We will correct the language and use “threat value”.

      Reviewer #2 (Public review):

      Summary:

      The following points are those that occurred to me across readings of the paper. They are listed in what I take to be the order of their significance. Many of the points relate to the loose use of language and invocation of concepts that are not warranted, given the study design and results obtained.

      Major Comments:

      (1) The concept of ensemble turnover is interesting - the way it is introduced and discussed implies some type of spontaneous change in the neural underpinnings of fear discrimination and generalization in the PL. But, of course, every trial involves an opportunity to learn about the threat CS or the generalization test stimuli, and I am troubled by the thought that stability in the neural underpinnings of fear discrimination and generalization will actually reflect the level of defensive behaviours evoked on different trial types and/or the discrepancy between those behaviours and the outcome of a given trial in the generalization test. That is, stability in the neural underpinnings may be related to an animal's certainty or uncertainty in the contingency between a stimulus and danger; or, put another way, an animal's confidence that danger will or won't occur given the presence of some stimulus. This is not uninteresting. It is, however, not considered anywhere in the paper, which is overloaded with references to inferred threat values and integration of information across different types of stimuli. The protocol is not one that requires inference about anything or integration across anything.

      We thank the reviewer for these important points, which we address in further detail below.

      Ongoing learning during test sessions: The reviewer correctly notes that unreinforced test presentations may constitute extinction-learning trials and that some neural changes across days could therefore reflect ongoing learning rather than spontaneous ensemble reorganization. However, new analyses indicate that extinction is unlikely to be the primary driver of our findings. Discrimination ratios do not decay over time; instead, they either sharpen or remain stable across sessions (new analyses to be included in the resubmission). These results argue against robust extinction as the primary source of the neural changes observed across sessions. This interpretation is also consistent with the strength of our conditioning protocol, which used 10 CS+ shock pairings and 10 CS− no-shock pairings specifically to minimize extinction across repeated testing sessions. Nevertheless, we acknowledge that the current design cannot fully dissociate time-dependent consolidation from retrieval-induced plasticity, and we will explicitly discuss this limitation in the revised Discussion.

      Stability reflecting behavioral consistency: We agree this alternative cannot be fully excluded. However, the cluster stability analyses assess identity at the level of response profile across all four frequencies, not response magnitude alone. Tone-selective clusters, which also show consistent behavioral correlates (firing rate correlates with threat-value, Fig. S8), do not show equivalent profile stability, suggesting that the stability of graded clusters is not simply a consequence of behavioral consistency. This point will be added to the Discussion in the resubmission.

      Language of "inference" and "integration": The reviewer is correct that responses to novel tones are consistent with graded stimulus generalization. We will substantially revise the manuscript to replace "inference" and "integration" with more precise language describing graded frequency generalization gradients.

      (2) I appreciate the link to Gu and Johansen in paragraph 3 of the Introduction, but the type of generalization under investigation here is not the same as the type of 'generalization' studied by Gu and Johansen [who used a sensory preconditioning protocol]. Nonetheless, the authors have forced the language used by Gu and Johansen into their paper, and this has created tension [at least for this reader] as the concepts introduced by Gu and Johansen [inference, integration] are simply not relevant given the generalization protocol used here. Here are a few examples of points where the tension might interfere with a reader's understanding:

      We thank the reviewer for these specific and constructive criticisms. We will revise the manuscript throughout to remove or redefine terms like "inferred valence" and "integration," replacing them with clearer, more accurate descriptions of gradient generalization of threat value. Below we address each point raised by the reviewer regarding terminology clarifications.

      (a) 'We hypothesized that generalization to novel stimuli depends on stable subnetwork organization that enables comparisons between learned and inferred valence, as well as population-level features that reduce variability across related representations.'

      I understand the words in the hypothesis, but can't form a representation of what is being said because of the reference to terms that stand in need of clarification [inferred valence, variability across related representations], but, ultimately, won't be clarified. This needs to be re-expressed so that the reader can appreciate what is being said.

      The hypothesis will be rewritten as: "We hypothesized that generalization to tones acoustically similar to the CS+ and CS− depends on the emergence of stable ensembles encoding threat value, and that population-level response similarity across stimuli would correlate with the degree of behavioral fear generalization, consistent with prior work in auditory cortex [1]."

      (b) 'Our results show that stable cortical subnetworks integrate the emotional "gist" of memory and inferred valence for novel cues over time, despite ongoing ensemble reorganization, and that population-level firing rate similarity across stimulus presentations determines threat generalization.'

      Again, what does this mean? How is the gist of a memory integrated with inferred valence for novel cues over time? The statement simply doesn't make sense. This needs to be rewritten for clarity.

      The summary statement will be rewritten: "Our results show that stable cortical sub-ensembles preserve the emotional content of the fear memory over time, despite ongoing ensemble reorganization, and that population-level firing rate similarity in response to tones associated with threat correlates with the degree of behavioral threat generalization."

      (c) 'In CS⁺15 mice, positively modulated sound-responsive neurons exhibited graded tone activity reflecting the contingency learned valence as well as the inferred valence of novel tones across testing days...'.

      Can this be rewritten as 'In CS⁺15 mice, positively modulated sound-responsive neurons exhibited graded activity to the tone CS and its variants that were used to assess generalization.'? The overloading of the text with references to 'contingency learned valence' and 'inferred valence' is unnecessary and makes it much harder to understand what has been shown in the results.

      We will adopt the reviewer's suggested rewording: "In CS+15 mice, positively modulated sound-responsive neurons exhibited graded activity to the tone CS and its variants that were used to assess generalization."

      We will systematically review the entire manuscript to ensure consistency with this revised framing.

      (3) Re the same passage of text as in 2c:

      Is it the case that these neurons are simply tracking the expression of freezing to the various tones? The same question applies to the results obtained for the CS+3 mice. If this is the case, then why should the results be taken to support the banner statement that 'Sound-modulated PL population responses encode learned and inferred valence' - these analyses do not support that statement. And, as indicated, I don't believe that the language of learned and inferred valence is appropriate to such statements, given the nature of the protocol used and results obtained. It is a study looking at how populations of neurons in the PL respond during presentations of auditory stimuli that were subject to discriminative conditioning, and during tests of generalized freezing to other [intermediate] auditory stimuli.

      The reviewer is correct that the graded population responses observed in PL could reflect freezing behavior across tone frequencies rather than encoding an abstract threat-value representation. This important concern was also raised by other reviewers. To address it directly, we will follow Reviewer 3’s suggestion and implement a Generalized Linear Model (GLM) using inferred spiking activity derived from the Ca2+ signals, with both tone identity and freezing behavior included as predictors. This analysis will allow us to dissociate the respective contributions of tone frequency and freezing to the graded neural responses. Based on the outcome of this analysis, we will revise and appropriately adjust our conclusions.

      In addition, we will revise the section heading and surrounding text to remove the terminology of “learned and inferred valence.” Instead, the findings will be described more conservatively as: “PL population responses reflect behavioral generalization to auditory stimuli following discriminative fear conditioning.”

      (4) It is stated that:

      'In no-shock controls, although both positive and negative responses were present, population activity was not modulated by tone frequency or valence'.

      What does this mean? I can understand that population activity was not modulated by tone frequency. But what does it mean to say that it was not modulated by valence? Why should it have been when none of the tones were conditioned in this group and, hence, mice were responding to all the tones equally? And given that this is true, I don't understand the use of 'valence' here, or the subsequent statements in this paragraph that 'graded responses require associative learning' and that 'PL population responses encode graded sound-valence associations that reflect both learning and inference, closely matching behavioral generalization.' The latter statement is particularly unwarranted and, again, highlights a major issue with the paper. It could and should be rewritten as 'PL population responses reflect behavioral generalization.' There is nothing in the additional language that adds to the reader's understanding of what has been shown. The reference to 'graded sound-valence associations that reflect both learning and inference' is completely unwarranted, given the nature of this study. It is anathema to the vast literature on stimulus generalization. If the authors wished to make statements of this sort, they should have taken a different approach, perhaps using protocols like those featured in Gu and Johansen.

      The reviewer is correct that controls do not form threat associations; however, these animals still could respond differentially to distinct frequencies, something that is not reflected in the data. We will correct the section indicating that distinct neutral frequencies do not produce graded responses: "graded responses require associative learning" will be retained but reframed simply as: "graded frequency-dependent population responses were absent in animals that did not receive fear conditioning." The concluding statement of the paragraph will be rewritten as: "PL population responses reflect behavioral generalization to acoustically similar stimuli following discriminative conditioning," in line with the reviewer's suggestion.

      (5) The section titled, 'Consistently active neurons preserve valence representations as newly recruited neurons sharpen remote memory traces' ends with the following summary:

      'Together, these results indicate that consistently active neurons maintain stable representations of learned and inferred sound associations across time, whereas neurons recruited after conditioning progressively acquire graded tuning at later retrieval stages. This dynamic refinement suggests that cortical memory representations become increasingly selective during systems consolidation, while a stable neuronal subpopulation preserves the core emotional content of the memory.'

      Once again, the summary is not in keeping with the results obtained. The 'dynamic refinement' of representations is far more likely to reflect the repeated testing across days 1, 15, and 30 rather than anything to do with systems consolidation - at the very least, it is the simplest interpretation of the results. The impact of repeated testing is evident in the sharpening of generalization gradients over time, which is contrary to what is otherwise observed in the literature - the incredibly well -documented broadening of generalization gradients with time. Given this impact of repeated testing, surely the changes in the neuronal population that underlie performance are more likely to reflect the learning that occurs on days 1, 15, and 30, which is reflected in reduced freezing to the non-conditioned tones. If this is a reasonable take on the results, then I don't see the basis for invoking systems consolidation at all, and I don't see the basis for inferring a stable neuronal subpopulation that preserves the emotional content of the memory. Rather, non-reinforced presentations of 'never-reinforced' tones result in recruitment of additional neurons that result in suppression of freezing responses to those stimuli.

      We respectfully disagree with the reviewer’s interpretation. While repeated testing cannot be entirely excluded as a contributing factor, several lines of evidence suggest that it cannot fully account for our observations.

      Regarding extinction: discrimination ratios between CS+ and all other frequencies either remained stable or increased over time (new analysis included in resubmission), indicating that animals continued to discriminate threat value across the testing period rather than showing the progressive suppression expected under extinction — the opposite of what we observe.

      Regarding the recruitment of new neurons: repeated non-reinforced tone exposure would be expected to produce stimulus-specific adaptation — characterized by reduced, less discriminative neural responsiveness and flatter tuning profiles [2]— not the progressive sharpening we observe. The same would be expected if these neurons represent or are associated with new extinction learning.

      Finally, sharpening of generalization gradients during repeated within-subjects testing has been reported previously [3], suggesting that successive exposures may promote more precise discrimination in some cases. Consistent with this, discrimination learning has also been shown to narrow or sharpen fear generalization gradients rather than broaden them [4], supporting the idea that discriminative conditioning enhances stimulus specificity during testing. Although we cannot exclude the possibility that more extended training could eventually broaden the generalization gradient, under the training parameters and temporal window used in our study, the data support a progressive sharpening of the gradient over time. In the revised Discussion, we will present systems consolidation as the primary interpretive framework and further elaborate on why repeated testing is unlikely to account for the full pattern of behavioral and neural findings reported here.

      (6) In the section titled, 'Population vector similarity at stimulus onset determines degree of generalization', it is stated that:

      'Because population similarity peaked shortly after stimulus onset, we quantified similarity during the first 5 s after tone onset relative to the CS⁺. In CS⁺15 mice, population similarity was highest for 15/15 and 15/11 tone pairs with no differences between them.'

      Isn't this consistent with the view that the population response in the PL simply reflects the level of freezing? Freezing to the 15-15 and 15-11 tones is most likely to be similar on their first presentation prior to the effects of extinction on the 11 Hz tone; hence the results obtained. That is, these results appear to clearly indicate that neuronal responses in the PL reflect the degree of stimulus generalization, as evidenced in freezing behavior. Given all that we know about the involvement of the PL in expressing fear responses, it is not appropriate to claim that 'population vector similarity at stimulus onset *determines* the degree of generalization. The PL responses simply reflect the varying levels of performance displayed to the different types of tones. What have I missed that could be taken to support additional statements?

      The GLM analysis described in our response to reviewers 1 and 3 will directly address the contribution of freezing. We will report these results in the resubmission and revise the interpretive language in the manuscript accordingly.

      However, regarding the analysis of population vector similarity, we need to clarify a point of confusion. The reviewer states “Freezing to the 15-15 and 15-11 tones is most likely to be similar on their first presentation prior to the effects of extinction on the 11 Hz tone; hence the results obtained”. The similarity vectors were calculated by correlating activity across all tone presentations within each testing day, not only the first two presentations. In Fig. 4, “Early” and “Late” refer to the order of a tone within a trial, which we will clarify more explicitly in the resubmission. Notably, repeated-measures analyses did not reveal any effect of the time variable (Fig. 4e,f), indicating that similarity across tone presentations remained high for tones associated with high threat value. Importantly, our data showed no evidence that responses to 11 kHz or 15 kHz in the CS15 group, or to 3 kHz in the CS3 group, exhibited extinction-like patterns at either the behavioral or neural level. Therefore, the persistence of high population similarity across time provides additional evidence against extinction as the primary explanation for our findings.

      We will remove the word "determines" from the manuscript, as our data cannot conclusively establish a causal relationship.

      Later in the same section, it is stated that 'population-level similarity at stimulus onset scales with behavioral threat generalization and is maximal for tones associated with robust threat responses.' For simplicity and, therefore, clarity, this should be rewritten as 'population-level similarity at stimulus onset reflects behavioral threat generalization.'

      We will make this correction.

      (7) In the section titled, 'Different subnetworks encode acoustic versus learned properties of sound association', it is stated that:

      'Our previous analyses show that learned and inferred associations are represented at the population level. However, these results do not resolve whether graded responses arise from pooled activity of frequency-selective neurons or from subnetworks encoding integrated learned valence across tones.'

      What does it mean to say 'integrated learned valence across tones'? As it presently stands, the meaning of the phrase is unclear. It only makes sense if one supposes that generalized freezing responses to the 11 and 7 kHZ tones reflect separate associations between those tones and the aversive foot shock US. This supposition is inconsistent with the rich literature on generalization of Pavlovian conditioned fear responses. Specifically, it is inconsistent with the many theories of fear generalization, which attribute the reduction in fear as one moves away from the specific conditioned stimulus to a decrement in the ability of the test stimulus to activate the trained CS-US association. My strong impression is that the authors would do well to ground their findings in theories of stimulus/fear generalization, of which there are many. This would better serve the results obtained [and the reader's appreciation of them] - at present, the unnecessary invocation of concepts does very little to enhance the reader's appreciation or understanding of what has been found in the study.

      We thank the reviewer for raising this point. The phrase "integrated learned valence across tones" refers specifically to a subpopulation of neurons that respond to all four frequencies in a graded manner, with response magnitude scaling according to threat value. This is distinct from tone-selective neurons, which respond preferentially to a single frequency. The neurons responding to all tones in a graded manner are present only in conditioned animals and not in no-shock controls, demonstrating that their graded response profile is shaped by associative learning.

      We agree, however, that the phrase "integrated learned valence" is unnecessarily opaque and we will replace it with more precise language: these neurons will be described as showing graded frequency-dependent responses whose magnitude scales with threat value. We believe this subpopulation represents a genuinely novel finding that complements the behavioral generalization literature by identifying a specific neural substrate for the generalization gradient within PL.

      (8) Another example of what has been a common theme in this review:

      '...we hypothesized that the PL active ensemble segregates into functionally distinct subnetworks: one encoding tone-specific sensory features with dynamic characteristics, and another responding to all frequencies encoding stable core memory content and inferred emotional valence.'

      What does it mean to say 'all frequencies encoding stable core memory content and inferred emotional valence'? Do the authors mean to say '...and another that tracks freezing/defensive responses regardless of whether they were elicited by the trained CS or one of the generalization test stimuli'?

      As stated in our previous responses, in the resubmission we will determine the contribution of freezing. If we find that freezing predicts graded neural responses, we will adjust the language of the manuscript.

      (9) It is stated that - 'Graded clusters encode emotional valence but constitute only a fraction of the active population; yet valence coding at the population level remains accurate and precise. This indicates that neurons newly recruited into the population-likely frequency-selective and organized within learning-independent clusters-can be shaped by associative processes through modulation of firing activity.'

      What does this mean? Are the authors trying to say that - 'Some clusters of PL neurons track freezing responses. In spite of the fact that these are only a fraction of the total active neuronal population, the population-level response of PL neurons also tracks the levels of fear to the trained tone and its variants used in the test for generalization.' If this is what one wants to say, then the final statement in the reproduced section does not follow. That is, there is no indication that 'neurons newly recruited into the population-likely frequency-selective and organized within learning-independent clusters-can be shaped by associative processes through modulation of firing activity.' As noted, the characteristics of other ensembles that become active across the repeated tests on days 1, 15, and 30 are more likely to reflect learning from non-reinforcement that occurs within and across those sessions. Perhaps this is what is meant by the phrase, 'shaped by associative processes'? If so, it should be stated explicitly instead of left to the reader to work out.

      We thank the reviewer for highlighting the lack of clarity in this passage and agree that the original phrasing was insufficiently precise. What we intended to convey is that only a subset of PL neurons displays graded tuning that tracks behavioral generalization across tones. Nevertheless, despite constituting only a fraction of the total active population, this graded coding is also reflected at the population level. Therefore, we suggest that neurons recruited into the active population after conditioning — likely frequency-selective neurons — contribute to the graded population responses through changes in their firing-rate activity, which is modulated by threat value (Fig. S8). We will rewrite this passage in the resubmission to make this interpretation explicit rather than leaving it to the reader to infer.

      Regarding the reviewer's suggestion that the characteristics of newly recruited neurons more likely reflect learning from non-reinforced exposures during repeated test sessions, we respectfully maintain that this interpretation is difficult to reconcile with two aspects of our data. First, graded-response neurons are absent in no-shock controls that are exposed to nonreinforced repeated testing. Second, as detailed in our responses to previous points, the progressive sharpening of population responses over time is inconsistent with what would be expected from repeated non-reinforced exposure, which would more plausibly produce broader or flatter tuning profiles.

      We agree that the phrase "shaped by associative processes" was ambiguous and will replace it with explicit language clarifying that we refer to fear conditioning as the associative process driving the emergence of graded responses, rather than any learning occurring during the test sessions themselves.

      (10) The following points all relate to the Discussion and reiterate many of the points above. 

      (a) 'A subset of neurons remains consistently active across sessions, preserving core components of the memory trace and supporting inference of emotional valence for novel sounds, while neurons recruited after conditioning progressively acquire valence selectivity at remote time points.'

      'Inference of emotional valence' is unclear and unwarranted for all of the reasons provided above regarding the use of language.

      We will modify the language as stated in the prior points.

      (b) '...Our data reconcile these views by demonstrating that cortical representations of emotional valence emerge rapidly after learning and persist within stable subnetworks, even as the broader population undergoes substantial turnover. This architecture preserves core mnemonic content while allowing flexibility in the surrounding ensemble.'

      These statements assume that the PL neuronal responses reflect something more than the levels of freezing behavior to the different stimuli; what are the grounds for this assumption?

      We will incorporate new analysis (GLM) to better address this point and conclusions.

      (c) 'Importantly, these subnetworks encode both learned contingencies and the inferred valence of novel stimuli along a graded representational axis, suggesting that strong recurrent connectivity provides a stable scaffold for emotional memory representations.'

      What is a graded representational axis, and what part of the first statement suggests that 'strong recurrent connectivity provides a stable scaffold for emotional memory representations'? If the authors' goal was to make statements about emotional memory representations vis-à-vis emotional memory content, they should have used protocols that allowed them to probe such content. The auditory fear conditioning protocol used here [followed by tests for generalization to other auditory stimuli that differ in frequency from the conditioned tone] is not one that lends itself to analysis of emotional memory representations or content.

      We thank the reviewer for this comment and agree that both phrases require clarification or revision.

      By "graded representational axis" we intended to convey that PL population activity varies systematically as a function of stimulus similarity to the conditioned tone — that is, population responses are not categorical but scale continuously with spectral proximity to the CS+. We agree this was not clearly stated and will revise the manuscript accordingly.

      Regarding recurrent connectivity, we agree with the reviewer that nothing in our data directly measures or manipulates connectivity between neurons. This statement was intended as a speculative interpretive hypothesis in the Discussion, motivated by the established literature linking strong recurrent connectivity in prefrontal circuits to stable population-level representations [5]. However, we acknowledge that invoking it in this context, without direct evidence, risks overstating our conclusions. We will revise this sentence to make its speculative nature explicit and ground it more carefully in the cited literature rather than presenting it as an inference from our own data.

      In summary, we will ensure our conclusions will be restricted to population-level coding of learned threat value and its generalization across auditory frequencies. We will revise the relevant passages in the Discussion to ensure that speculative interpretations regarding emotional memory content are either removed or clearly flagged as speculative hypotheses.

      (d) 'Dynamic tone-selective responsive neurons emerge independently of learning, as they are present in both control and experimental mice, reflecting pre-existing PL sensory-driven properties (Hockley & Malmierca, 2024; Zikopoulos & Barbas, 2006).'

      Maybe. They are also likely to have developed as a consequence of the repeated testing on days 1, 15, and 30, which involved intermixed exposures to the tones of different frequencies. That is, rather than 'pre-existing PL sensory-driven properties', the responses of these neurons might reflect the emergence of discrimination between the various tones across testing, and greater suppression of freezing to the non-trained tones compared to the trained tone across the various test intervals.

      We thank the reviewer for this point. Our interpretation that these neurons reflect pre-existing PL sensory-driven properties was based on the observation that tone-selective responses were present in control animals that never received conditioning, consistent with prior reports of sensory responsiveness in PL cortex ([6, 7]. Because these responses emerge from the first time we expose mice to the intermediate frequencies, they cannot be explained by repeated exposure. Moreover, we did not observe progressive refinement, emergence of discrimination-like changes, or suppression of responding to non-reinforced tones in control mice. This difference between conditioned and control animals indicates that repeated tone exposure alone is not sufficient to produce the observed dynamics — associative learning is necessary. We therefore maintain that the tone-selective responses of these neurons reflect pre-existing sensory-driven properties of PL cortex that are present independently of conditioning history.

      In summary, we thank the reviewer for suggesting clarifications to our interpretation, for raising the possibility that freezing behavior may contribute to graded neural responses, and for raising the question of whether repeated tone exposure may contribute to the properties of neurons recruited after conditioning. In the revised manuscript, we will include additional analyses to better dissociate the contributions of freezing behavior and tone identity, clarify passages that were insufficiently precise, and include a paragraph in the Discussion addressing potential alternative explanations alongside our own interpretation of the data.

      Reviewer #3 (Public review):

      Summary:

      Normandin et al. explore the coding of stimuli predicting an aversive event in the prelimbic cortex. Stimuli could either be explicitly paired, explicitly unpaired, or novel but with an inferred association with the aversive event (generalization). Long-term tracking of GCaMP-positive neurons allowed them to examine how coding evolves out to a month following training. In general, they found two types of ensemble codes. One was ensembles coding for each stimulus independently, but with enhanced responding to the one eliciting a freezing response. The other was ensembles that responded to all stimuli in proportion to their similarity to the stimulus paired with the aversive event, either increasing or decreasing their activation with the degree of freezing elicited by a stimulus. Importantly, this second set of ensembles was more stable across days, potentially providing a memory trace.

      Strengths:

      (1) The authors track ensembles in prelimbic cortex over long time scales, providing valuable information on the consolidation of neural codes.

      (2) Neural coding of generalization is examined, which is under-examined in the field.

      We thank the reviewer for appreciating our design to track ensembles over time and the relevance of studying the neural substrates of generalization.

      Weaknesses:

      (1) Difficult to determine if responses treated as encoding stimulus valence are driven instead by the behavior that the stimulus elicits, freezing.

      We thank the reviewer for this thoughtful and constructive comment. We agree that an alternative interpretation is that the graded-response ensembles may partially reflect freezing-related activity rather than mnemonic or salience-related representations of the conditioned stimuli themselves. In the revision, we will acknowledge that prior work has identified PL neurons that encode freezing independently of stimulus identity or associative content. Furthermore, we will implement the reviewer’s suggested generalized linear model (GLM) approach using inferred spiking activity derived from the Ca2+ signals. Specifically, we will include both stimulus identity and freezing behavior as predictors. Because freezing varies across trials whereas stimulus presentation is fixed, this analysis will allow us to dissociate the relative contributions of stimulus-related versus freezing-related activity to the graded neuronal responses. We thank the reviewer for this excellent suggestion.

      If graded stimulus coding remains significant after accounting for freezing behavior, this would strengthen the interpretation that these ensembles encode learned salience or associative properties of the stimuli rather than behavioral output alone. Conversely, if freezing explains a substantial proportion of the variance, we will revise our interpretation accordingly.

      (2) The study implies that the identified ensembles are causally related to valence memory, but no experimental interventions are performed to justify this.

      We appreciate the reviewer's point. We agree that our data are correlational in nature and that establishing a causal relationship between identified ensembles and valence memory would require experimental interventions such holographic two-photon manipulations, which are beyond the scope of the present study but represent an important direction for future work.

      To provide an indirect link between ensemble organization and behavior within the constraints of the current dataset, we will examine inter-individual variability in the revised manuscript. Specifically, we will test whether the proportion of neurons participating in stable graded-response ensembles versus dynamic stimulus-specific ensembles predicts individual differences in freezing behavior and fear generalization across retrieval sessions. If animals with a higher proportion of stable graded-response neurons show stronger discrimination and less generalization to non-conditioned tones, this would strengthen the association between ensemble organization and behavioral outcome, while remaining correlational in interpretation.

      We will modify the manuscript terminology accordingly, replacing causal language with phrasing that accurately reflects the associative nature of our conclusions.

      References

      (1) Aschauer, D.F., et al., Learning-induced biases in the ongoing dynamics of sensory representations predict stimulus generalization. Cell Rep, 2022. 38(6): p. 110340.

      (2) Kato, H.K., S.N. Gillet, and J.S. Isaacson, Flexible Sensory Representations in Auditory Cortex Driven by Behavioral Relevance. Neuron, 2015. 88(5): p. 1027–1039.

      (3) Vervliet, B., et al., Generalization gradients in human predictive learning: Effects of discrimination training and within-subjects testing. Learning and Motivation, 2011. 42(3): p. 210–220.

      (4) Dunsmoor, J.E. and K.S. LaBar, Effects of discrimination training on fear generalization gradients and perceptual classification in humans. Behav Neurosci, 2013. 127(3): p. 350–6.

      (5) Mante, V., et al., Context-dependent computation by recurrent dynamics in prefrontal cortex. Nature, 2013. 503(7474): p. 78–84.

      (6) Hockley, A. and M.S. Malmierca, Auditory processing control by the medial prefrontal cortex: A review of the rodent functional organisation. Hear Res, 2024. 443: p. 108954.

      (7) Zikopoulos, B. and H. Barbas, Prefrontal projections to the thalamic reticular nucleus form a unique circuit for attentional mechanisms. J Neurosci, 2006. 26(28): p. 7348–61.

    1. eLife Assessment

      This work provides a fundamental advance through a detailed, integrative analysis of how the tsetse fly feeds on blood, demonstrating that successful penetration depends on subtle structural adaptations rather than extreme forces or unusual anatomy. By combining high-resolution imaging, innovative biomechanical measurements, and experiments on artificial skin, the study offers complementary and compelling evidence, with clear data supporting a robust mechanistic interpretation. These findings have broad significance, as they clarify the biomechanics of vector feeding and have implications for the transmission of diseases such as African trypanosomiasis across diverse hosts.

    2. Reviewer #3 (Public review):

      Summary:

      Human and animal trypanosomiasis are fatal illnesses caused by African trypanosomes transmitted by tsetse flies during a bloodmeal. Thus, tsetse fly feeding is the key physical step in disease transmission to mammals. Tsetse fly feeding is not a new story, but it is revisited here through the application of sophisticated imaging techniques and novel biomechanical methods of analysis. The author's aim is to provide a high-resolution picture of the structures and forces involved in feeding to provide mechanistic insights into the process of feeding, from attachment, penetration, drinking and retraction of the feeding parts.

      Largely the authors have achieved their aims. They (i) examine the structures and forces involved in attachment; (ii) they provide detailed multi image analysis of the proboscis providing insights into its probing ability and physical mechanism of penetration; (iii) they conduct a controlled analysis of the physical forces involved in penetration and report that they are in the low nM range, not especially strong but much higher that the mosquito bite and finally they provide a first analysis of blood uptake during feeding.

      Strengths:

      The study images the tsetse fly feeding structures in unprecedented detail, with resolution to the uM scale, in 3-D, and during feeding. The resulting images are dramatic and insightful (and beautiful and frightening!) that researchers interested in trypanosomes, tsetse flies or blood feeding by flies in general will want to see.

      They conclude that flies attach strongly to smooth surfaces, because of interactions possible via the array of acanthae of the pulvillus pad at the ends of the tarsi. The estimated attachment forces are similar in male & female flies, in the low mM range (they look impressively strong in video 1). They provide a very striking analysis of the proboscis and labellum and associated tooth structures (Figs 4 & 5). I recall many years ago observing that tsetse flies are messy feeders, and these structures, especially the rasping teeth structures on the reverse folded labial tips explain why! This seems more like a chainsaw than a jigsaw in action, but the authors are probably correct that these structures and probing/retraction mechanism explain many features of tsetse fly feeding and their ability to feed on a wide range of hosts with very different skin types.

      The impressive aspect of this paper is the range of imaging techniques, (CLSM, SEM, uCT, FIB SEM), the quality of the images which attests to the obvious care taken with sample preparation. The biomechanically analysis, especially the penetration analysis is impressive. Finally, the paper is clearly written and presented, it was a very easy read and overall, a very engaging study.

      Weaknesses:

      I suppose it could be said that the paper is a descriptive study; it doesn't really test a hypothesis but that is not a prerequisite for publication. Perhaps the least convincing prats are the imaging of the flexible v rigid parts of the structures, which is based on amount of resilin (flexible) and chitin-protein (stiff) based on their autofluorescence. In seems odd that the joints would be less blue (stiffer) in Fig 1i, or what the blue structures correspond to in Fig. 6B-D.

      Comments on revised version.

      In revised version these issues have been satisfactorily addressed

    3. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This manuscript provides a comprehensive and mechanistic analysis of how tsetse flies feed on blood across a wide range of host skin types. The authors combine detailed anatomical characterization of the feeding apparatus with quantitative measurements of mechanical properties, probing forces, and blood uptake, complemented by experiments using artificial skin. They show that tsetse flies do not rely on extreme forces or uniquely specialized structures, but instead on subtle and highly efficient structural and mechanical adaptations (such as the toothed labellum and coordinated proboscis movements) to achieve effective blood pool feeding. The study successfully moves beyond descriptive anatomy to a quantitative, functional analysis that explains how feeding is accomplished across diverse substrates.

      Strengths:

      A major strength of the work is the impressive integration of multiple complementary approaches. Advanced imaging tools provide a convincing three-dimensional view of the proboscis, labellum, and associated structures, while direct force measurements and blood intake quantification place these observations on a solid quantitative footing. The use of artificial skin with different mechanical properties is particularly powerful, as it allows structure-function relationships to be tested under controlled and reproducible conditions. Together, these datasets provide strong and coherent support for the authors' central conclusions. The quantitative treatment of feeding mechanics represents a significant advance over largely descriptive prior work by others (e.g., Gibson W et al 2017) and establishes a valuable mechanistic insight for studying blood feeding in insect vectors more broadly.

      Weaknesses:

      The study focuses almost entirely on uninfected flies and does not address how infection might alter feeding mechanics or performance. Previous work has shown that trypanosome infection can affect salivary gland function and feeding time (Van Den Abbeele et al 2010), and even cause damage to mouthparts, all of which can influence feeding behavior and efficiency. While this does not detract from the technical quality or the core findings of the study, a more explicit discussion of these biological variables would help place the results in a broader transmissionrelevant context and clarify how generalizable the conclusions are to natural infection settings.

      We thank the reviewer for this important comment. While our study focused on uninfected flies, we agree that parasite infection may influence feeding performance and should therefore be considered when assessing the broader relevance of our findings. Previous studies have shown that trypanosome infections can alter salivary gland physiology and saliva composition (Van Den Abbeele et al., 2010; Matetovici et al., 2016). In addition, transcriptomic analyses suggest that infection with T. congolense may affect the molecular and physiological state of the proboscis (Awuoche et al., 2017). However, there is currently no direct evidence that these changes translate into fundamental alterations of the mechanical properties or function of the mouthparts themselves, which were the primary focus of our study. We have now expanded our manuscript to discuss this (lines 508-520):

      "It is also important to note that our experiments were conducted using uninfected flies. Previous studies have shown that trypanosome infection can alter feeding behaviour, leading to increased probing activity and prolonged feeding times (Jenni et al., 1980; Van den Abbeele et al., 2010). These effects have primarily been attributed to infection-induced changes in saliva composition and the resulting interactions with host blood (Van den Abbeele et al., 2010). However, a different study found no significant effects of infection with either salivary gland-resident T. brucei or with proboscis-colonizing species such as T. congolense and T. vivax on Glossina feeding behaviour (Moloo, 1983). Furthermore, although infection-associated transcriptional changes in the salivary glands and proboscis have been reported (Awuoche et al., 2017; Matetovici et al., 2016), there is currently no direct evidence that trypanosome infection alters the mechanical properties or function of the mouthparts themselves."

      Overall, this is an outstanding and carefully executed study that will have a significant impact on the fields of vector biology and parasite transmission.

      Reviewer #2 (Public review):

      Summary:

      This manuscript presents an impressively detailed, multidisciplinary analysis of the mechanics of blood feeding in Glossina spp. Combining SEM, CLSM, µCT, FIB-SEM, macro-videography, and quantitative force measurements, the authors characterize the structures and biomechanics of attachment, proboscis deployment, tissue penetration, and blood uptake. They also examine interactions with diverse host-type substrates, from human skin equivalents to cow, deer, and lizard skin, and integrate these with force measurements to quantify penetration and retraction dynamics.

      The work's key conclusion is that the tsetse fly does not rely on any single exceptional morphological innovation, but rather uses a suite of subtle structural features and retractive forces to feed efficiently across diverse hosts. This result is novel, insightful, and evolutionarily compelling. Overall, this is a strong manuscript that combines methodological sophistication with biological relevance. It should be of high interest to researchers studying vector biology, biomechanics, parasite transmission, and vector-host interactions.

      Strengths:

      (1) The combination of SEM, CLSM, µCT, and FIB-SEM provides an unusually comprehensive anatomical characterization of the tsetse feeding apparatus.

      (2) The direct measurement of proboscis penetration and retraction forces across diverse substrates is highly original and fills a major knowledge gap in vector-host interaction mechanics.

      (3) The study bridges morphology, mechanics, behavior, and host tissue properties, which strengthens the overall conclusions.

      (4) Imaging of trypanosomes within the hypopharynx and surrounding tissue during feeding provides new information about parasite delivery mechanisms.

      Main Comments:

      (1) The authors conclude that feeding versatility arises from the sum of subtle adaptations. This interpretation is reasonable, but it would help to sharpen which findings most robustly support this statement. For example, the relative similarity of proboscis forces across skin types is compelling evidence that the proboscis is broadly tuned rather than specialized. The observation that tsetse targets softer interscale regions on lizard skin suggests behavioural selectivity, not morphological specialisation. It would strengthen the discussion to highlight which data most directly refute the hypothesis of a unique specialization.

      We thank the reviewer for this comment. To address this point more explicitly and to sharpen the interpretation of our findings, we have expanded the final conclusion in the Discussion (lines 528544):

      "Ultimately, the objective of this study was to investigate how tsetse flies can feed on a seemingly random selection of animals with highly diverse skin structures. In our detailed anatomical studies and force measurements, we did not identify a single dominant trait that explains the fly's feeding versatility.

      Instead, our results indicate that this capability emerges from the combined effect of multiple, more subtle traits. In particular, the proboscis generates broadly similar penetration forces across a wide range of skin types, suggesting a generalised mechanical mechanism rather than hostspecific optimisation. The intricate architecture of the labellum and the strong retractile forces during probing likely contribute to efficient penetration and blood pool formation across heterogeneous substrates. Behaviourally, tsetse flies further increase feeding success by flexibly targeting mechanically favourable sites, such as the softer interscale regions on lizard skin, rather than relying on specialised morphological adaptations.

      This composite strategy likely reflects evolutionary fine-tuning that enables the broad host range of tsetse flies. By allowing efficient blood feeding across diverse vertebrate hosts, this versatility may also have facilitated the ecological success and transmission opportunities of African trypanosomes.”

      (2) A central finding is that retraction forces exceed penetration forces across substrates, implying that backward pulling is a key component of wound creation. However, the biological interpretation could be deepened. Specifically, do the authors believe retraction serves primarily to enlarge the pool-feeding site? How does this compare mechanically to mosquito fascicle oscillation or other blood-feeding arthropods (especially other flies such as those in the tabanidae family)? Could retraction forces contribute to anchoring or resisting host grooming behaviors?

      The stronger retraction forces observed during probing indeed suggest that backward pulling is not a passive withdrawal, but likely an active component of tissue disruption. As discussed in the manuscript (lines 475–482), we interpret these repeated pullback movements, together with the outward-facing prestomal teeth of the everted labellum, primarily as a mechanism to enlarge the feeding lesion and improve access to blood, consistent with the blood pool feeding strategy of tsetse flies. To make this more clear, we have added a half sentence to line 482 "..., thereby creating a larger blood pool for feeding."

      We also already compare this mechanism to mosquito feeding mechanics in the discussion (starting from line 487). In mosquitoes, high-frequency fascicle oscillations are thought to reduce insertion resistance and facilitate minimally invasive capillary feeding. Although we also observed oscillatory movements during tsetse feeding (Video 4), the underlying mechanical strategy appears fundamentally different. In contrast to the mosquito’s system optimized for delicate penetration, the tsetse proboscis appears adapted for forceful tissue disruption during pool feeding. Notably, the oscillations observed in tsetse flies seem to occur during active blood uptake rather than initial tissue penetration. Consequently, the functional role of these oscillations in tsetse flies remains unclear. We have now addressed this more specifically in the discussion (lines 490-495):

      "Oscillatory movements were also observed during tsetse probing (Video 4). Notably, these oscillations appeared predominantly during active blood uptake rather than during the initial penetration phase, suggesting that they are associated with ingestion rather than insertion. Whether they facilitate blood flow, prevent occlusion of the feeding canal, or simply reflect pump activity remains unknown."

      When looking at other species, stable flies (Stomoxys) may represent a particularly relevant comparison because they employ a similar penetration mechanism and are pool feeders with prominent prestomal teeth (Krenn and Aspöck. Function and evolution of the mouthparts of blood-feeding Arthropoda. Arthropod structure and development, 2012). In contrast, tabanids employ a different mouthpart architecture with rasping/cutting structures but without comparable prestomal teeth. Whereas mosquito mouthparts have been described as functioning like a syringe, and we compare the tsetse proboscis to a saw, tabanid mouthparts have been likened to scissors (Krenn and Aspöck. Function and evolution of the mouthparts of blood-feeding Arthropoda. Arthropod structure and development, 2012). Although tabanids are also known to inflict substantial tissue damage, it remains unclear whether their feeding movements produce retraction-dominated force patterns comparable to those we observed in tsetse flies.

      Lastly, we agree that the elevated resistance generated during retraction may contribute to withstanding host defensive behaviour such as shake-off responses. Structurally, the orientation of the prestomal teeth and the architecture of the everted labellum could provide temporary anchoring during feeding, as we have already briefly discussed in the manuscript (lines 458– 461). However, while stronger anchoring may increase feeding stability, it could also increase the risk of injury to the fly if detected by the host. Compared to other pool-feeding flies such as stable flies, tsetse flies have been reported to respond more readily to host defensive behaviour (Schofield and Torr. A comparison of the feeding behaviour of tsetse and stable flies. Medical and Veterinary Entomology, 2002). We therefore currently consider anchoring to be a possible secondary function but lack direct experimental evidence to assess its practical importance.

      (3) The study analyzes a diverse set of substrates, which is a strength. However, some caveats deserve explicit discussion. Human skin equivalents and dermal equivalents lack the full mechanical complexity of real skin (e.g., innervation, perfusion, tension). Frozen or ethanol-stored samples, particularly reptile skin, may also exhibit altered mechanical properties compared to live tissues. These limitations do not undermine the findings but should be explicitly acknowledged as they influence the interpretation of absolute force magnitudes.

      The reviewer raises a valid point regarding the interpretation of absolute force magnitudes across the measured substrates. We have therefore added a clarifying statement to the discussion (lines 469-476):

      "When interpreting absolute force magnitudes, it is important to bear in mind that our samples do not fully recapitulate physiological conditions. Skin explants and skin equivalents may behave differently to skin under active perfusion and native tissue tension, as may our fixed and frozen animal skin samples. Nevertheless, comparative force measurements revealed consistent biomechanical signatures across substrates, suggesting that the observed force patterns reflect fundamental aspects of the feeding mechanism that are likely relevant in vivo.”

      (4) The SEM and FIB-SEM images showing trypanosomes in the hypopharynx and surrounding tissue during penetration are visually striking and suggest rapid dispersal. It would be helpful to connect these observations more clearly to the kinetics of parasite deposition and whether mechanical tissue laceration is likely to increase inoculation efficiency. Without conducting additional experiments, the authors could discuss whether these findings support or modify existing models of salivary-gland-derived parasite release.

      We have now expanded the Discussion to clarify that our observations of trypanosomes in the hypopharynx are consistent with the established model of salivary-gland-derived parasite release during probing and feeding, in which infective metacyclic trypanosomes are delivered with saliva into the host tissue. Furthermore, the presence of trypanosomes beyond the immediate feeding canal supports rapid parasite dispersal following inoculation, as described in previous work (Reuter et al., 2023). In this context, the tissue laceration generated by the tsetse proboscis may facilitate local parasite distribution by creating a larger, mechanically disrupted feeding lesion. However, our data provide high-resolution structural snapshots and were not designed to quantify deposition kinetics or inoculation efficiency. We therefore refrain from concluding that mechanical laceration increases transmission efficiency and instead view this as a plausible consequence that should be tested directly in future work. Specifically, we have added this paragraph to the discussion (521-527):

      "Overall, our observations of trypanosomes within the fly's hypopharynx, labial gutter, and host tissue are consistent with the established model of salivary-gland-derived parasite release during probing and feeding. Their presence beyond the immediate feeding canal is consistent with rapid local dispersal following inoculation, as described previously (Reuter et al., 2023). This process may be facilitated by the extensive tissue disruption caused by the tsetse mouthparts, although this hypothesis will require direct experimental testing."

      (5) The authors demonstrate that tsetse attachment abilities fall within the range of generalist insects and are far lower than those of obligate ectoparasites. However, the manuscript could discuss how attachment forces relate to the tsetse's ecological context, e.g., whether their attachment is generally brief, whether host shaking strongly selects for grip strength, etc. Is there evidence that other Glossina species or tabanids with different host preferences show variation in attachment performance? This would broaden the relevance of the findings.

      Tsetse flies are obligate blood feeders, but host contact is typically brief and frequently interrupted by host defensive behaviour. As a result, selection may favour rapid and efficient feeding rather than exceptionally strong attachment. This interpretation is supported by Schofield and Torr (A comparison of the feeding behaviour of tsetse and stable flies. Medical and Veterinary Entomology, 2002), showing that tsetse flies experience more feeding interruptions than the stable fly Stomoxys calcitrans, despite completing successful blood meals in less time. These differences are consistent with life-history theory (Anderson and Roitberg. Modelling trade-offs between mortality and fitness associated with persistent blood feeding by mosquitoes. Ecology Letters, 1999), which predicts that long-lived species with low reproductive rates, such as tsetse flies, should be less willing to risk injury by persisting on a host than shorter-lived, more fecund species. Against this background, our finding that tsetse attachment forces fall within the range reported for generalist insects, appears biologically plausible. Their attachment performance needs to be functionally sufficient for brief feeding events rather than maximized for prolonged host retention.

      We are not aware of comparative biomechanical data on attachment performance across different Glossina species or tabanids. We agree that such comparative studies would be valuable to test whether differences in host preference and feeding ecology correlate with variation in attachment capacity.

      (6) In video 4, could the authors clarify whether the observed maxillary vibrations are hypothesized to reduce penetration resistance or serve another function?

      The vibrations of the maxilla specifically appear during active blood uptake rather than during initial tissue penetration, suggesting they are linked to the ingestion phase. Whether they serve a mechanical function, such as facilitating blood flow or preventing canal occlusion, or represent a passive consequence of pump activity, remains unclear. We consider this an open and interesting question that warrants dedicated investigation.

      We have therefore clarified that the functional significance of these oscillations remains unresolved to date (lines 490-495). This reads: “Oscillatory movements were also observed during tsetse probing (Video 4). Notably, these oscillations appeared predominantly during active blood uptake rather than during the initial penetration phase, suggesting that they are associated with ingestion rather than insertion. Whether they facilitate blood flow, prevent occlusion of the feeding canal, or simply reflect pump activity remains unknown.”

      Reviewer #3 (Public review):

      Summary:

      Human and animal trypanosomiasis are fatal illnesses caused by African trypanosomes transmitted by tsetse flies during a bloodmeal. Thus, tsetse fly feeding is the key physical step in disease transmission to mammals. Tsetse fly feeding is not a new story, but it is revisited here through the application of sophisticated imaging techniques and novel biomechanical methods of analysis. The authors aim to provide a high-resolution picture of the structures and forces involved in feeding to provide mechanistic insights into the process of feeding, from attachment, penetration, drinking and retraction of the feeding parts.

      Largely, the authors have achieved their aims. They (i) examine the structures and forces involved in attachment; (ii) they provide detailed multi image analysis of the proboscis providing insights into its probing ability and physical mechanism of penetration; (iii) they conduct a controlled analysis of the physical forces involved in penetration and report that they are in the low nM range, not especially strong but much higher that the mosquito bite and finally they provide a first analysis of blood uptake during feeding.

      Strengths:

      The study images the tsetse fly feeding structures in unprecedented detail, with resolution to the uM scale, in 3-D, and during feeding. The resulting images are dramatic and insightful (and beautiful and frightening!), so researchers interested in trypanosomes, tsetse flies, or blood feeding by flies in general will want to see.

      They conclude that flies attach strongly to smooth surfaces because of interactions possible via the array of acanthae of the pulvillus pad at the ends of the tarsi. The estimated attachment forces are similar in male & female flies, in the low mM range (they look impressively strong in video 1). They provide a very striking analysis of the proboscis and labellum and associated tooth structures (Figures 4 & 5). I recall many years ago observing that tsetse flies are messy feeders, and these structures, especially the rasping teeth structures on the reverse folded labial tips, explain why! This seems more like a chainsaw than a jigsaw in action, but the authors are probably correct that these structures and the probing/retraction mechanism explain many features of tsetse fly feeding and their ability to feed on a wide range of hosts with very different skin types.

      We agree that “jigsaw” may be too specific and not fully appropriate in this context. We have therefore replaced it in the manuscript with the more general term “saw.”

      The impressive aspect of this paper is the range of imaging techniques (CLSM, SEM, uCT, FIB SEM), the quality of the images, which attests to the obvious care taken with sample preparation. The biomechanical analysis, especially the penetration analysis, is impressive. Finally, the paper is clearly written and presented; it was a very easy read and, overall, a very engaging study.

      Weaknesses:

      I suppose it could be said that the paper is a descriptive study; it doesn't really test a hypothesis, but that is not a prerequisite for sharing it. Perhaps the least convincing parts are the imaging of the flexible versus rigid parts of the structures, which is based on the amount of resilin (flexible) and chitin-protein (stiff), based on their autofluorescence. It seems odd that the joints would be less blue (stiffer) in Figure 1i, or what the blue structures correspond to in Figure 6B-D.

      Our analysis is based on established CLSM approaches that use exoskeleton autofluorescence as a proxy for relative differences in cuticular composition and material properties (Michels & Gorb, 2012; Michels et al., 2016). In the tarsus, the observed differences in inferred stiffness are relatively subtle, with most regions exhibiting broadly comparable material properties. This becomes particularly evident when compared with the proboscis, where the contrasts in cuticular composition are much more pronounced (Figure 6). We also note that locally stiffer regions at joints are not unexpected, as stiffness gradients in arthropod joints can provide mechanical support and help constrain the direction of movement. Importantly, our images show a flexible, ring-like blue region directly at the articulation, surrounded by slightly stiffer material. We therefore interpret this pattern as a combination of a flexible hinge region and adjacent supporting structures that together enable controlled joint motion.

      The blue structures in Figure 6B–D correspond to flexible regions of the furca (f). Because this spring-like cuticular element undergoes substantial configuration changes during labellar eversion, the presence of highly flexible regions is consistent with its proposed mechanical function.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      No further experiments or analyses are suggested. However, the Discussion would benefit from briefly acknowledging how trypanosome infection can alter feeding behavior and mouthpart function, based on prior work, to place the mechanical findings in a more biologically relevant transmission context.

      We thank the reviewer for this suggestion. Previous studies have indeed shown that trypanosome infection can alter tsetse feeding behavior, primarily through changes in saliva composition. Van den Abbeele et al. (2010) demonstrated that infection with T. brucei significantly impairs the anti-haemostatic activity of tsetse saliva, resulting in prolonged prefeeding probing and therefore extended feeding times. These findings are consistent with earlier observations by Jenni (1980), who reported increased probing frequency in infected flies.

      Jenni (1980) also proposed that these behavioral changes might be linked to altered mechanoreceptor function. However, Van den Abbeele et al. (2010) argued against this interpretation for T. brucei, noting that this parasite does not colonize the mouthparts where these mechanoreceptors are located. Taken together, the available evidence suggests that the observed changes in feeding behavior are mediated primarily through altered interactions with host blood rather than through direct effects on the mouthparts themselves.

      It should be noted that this conclusion is specific to T. brucei. Other tsetse-transmitted trypanosome species, such as T. congolense, do colonize the proboscis. However, a comparative study examining flies infected with T. brucei, T. congolense, or T. vivax found no significant effects of infection on feeding behaviour relative to uninfected controls (Moloo, 1983). To our knowledge, there is currently also no direct evidence that any trypanosome species alters the physical properties or mechanical function of the mouthparts, or causes damage that would directly affect feeding performance. We have added this paragraph to the Discussion (lines 508520):

      "It is also important to note that our experiments were conducted using uninfected flies. Previous studies have shown that trypanosome infection can alter feeding behaviour, including increased probing activity and prolonged feeding times (Jenni et al., 1980; Van den Abbeele et al., 2010). These effects have primarily been attributed to infection-induced changes in saliva composition and the resulting interactions with host blood (Van den Abbeele et al., 2010). However, a different study found no significant effects of infection with either salivary gland-resident T. brucei or with proboscis-colonizing species such as T. congolense and T. vivax on Glossina feeding behaviour (Moloo, 1983). Furthermore, although infection-associated transcriptional changes in the salivary glands and proboscis have been reported (Awuoche et al., 2017; Matetovici et al., 2016), there is currently no direct evidence that trypanosome infection alters the mechanical properties or function of the mouthparts themselves."

      Reviewer #2 (Recommendations for the authors):

      Several figures (particularly SEM-based ones) contain very dense labeling. Consider providing simplified overviews or annotated "orientation guides" in figure supplements to improve navigability for readers unfamiliar with proboscis anatomy.

      We thank the reviewer for this helpful suggestion. While we agree that orientation aids can be valuable, we have decided not to include additional simplified overview figures, as we consider that introducing separate schematic summaries could potentially complicate rather than improve navigation of the structural detail. We therefore rely on consistent labelling within the existing figures and detailed captions to guide interpretation.

      The manuscript uses appropriate non-parametric tests, but could benefit from reporting effect sizes and indicating sample sizes on all plots.

      Sample sizes are reported in the figure legends, Methods section, and Supplementary material for all experiments. We agree that reporting effect sizes can be informative and will consider this in future studies. However, because the primary objective of the statistical analyses in the present work was to support comparisons between experimental conditions rather than to estimate effect magnitudes, and because the figures are already information-dense, we therefore decided not to further modify the graphical presentation in this revision.

      Reviewer #3 (Recommendations for the authors):

      (1) P5 L111. Perhaps indicate these knobs on the image Figure 1S). I assume these are the structures visible under the pointer labelled spa? Maybe highlight some of the worn areas in Figure 1G.

      The knob-like structures in Figure S1 are highlighted in green and we have now revised the figure description from:

      “…showing fine crests on the underside and surface modifications (green) on the upper side.”

      to:

      “…showing fine crests on the underside and knob-like surface modifications (green) on the upper side.”

      Regarding Figure 1G, the purpose of the panel is to illustrate the contrast between deformed spatulae (Figure 1G) and intact spatulae (Figure 1H). We therefore chose to retain the original presentation, as we feel that additional markings would not substantially improve interpretation and could obscure structural details. We hope that the direct comparison between the two panels provides sufficient visual guidance.

      (2) P9. The frictional force (and P38/39) has the units of N (kg.m/Sexp2). The safety factor is this force divided by the weight of the fly? So are there units (Kg/sexp2) or are these not shown? Perhaps this is a convention.

      The safety factor is defined as the ratio of the total frictional force to the fly’s weight force (m·g), where m is body mass and g is gravitational acceleration. Since both quantities are express in Newtons (kg·m·s<sup>-2</sup>), the safety factor is dimensionless.

      We agree that the terminology in the original manuscript may have been ambiguous, as “body weight” is sometimes used colloquially to refer to body mass. To avoid confusion, we have revised the text to explicitly refer to weight force and now define the safety factor as the total friction force divided by weight force (mg, where m is body mass and g is gravitational acceleration). We have clarified this in the main text, the Figure 2 legend, and the description of Supplementary Material 1.

      (3) P10 Figure 2G & H. It is not very clear...are these the data, the average of all readings across all surfaces in E and F? If so, why is this value useful...how does it add to what is already shown?

      The figures 2G and 2H summarize the friction forces (G) and safety factors (H) across all tested substrates, based on the values from the male (B, E) and female (C, F) datasets. The purpose of these panels is to provide an overall comparison between sexes independent of substrate type. While this information can also be inferred from the substrate-specific plots, the sex-separated presentation does not make the absence of an overall sex difference immediately obvious. Figures 2G and 2H therefore serve as concise summary plots highlighting this result.

      (4) P12. For the nonspecialist, it might be useful to draw a cartoon showing the organisation of the labium, labrum and the hypopharynx...this is visible in Figure 4i but not in the dissected proboscis and labellum ....only the labium as the labrum doesn't extend this far?

      To clarify the anatomical arrangement in the dissected specimen, we have added the following statement to the Figure 4 legend (lines 224–226):

      “In an intact fly, the labrum would be positioned within the empty groove of the labium visible in J; however, it is absent in this dissected preparation.”

      (5) P17 legend to Figure 5. Include the Lm abbreviation in the legend, and maybe a close-up of the rsp teeth?

      We have added “lm, labellum” to the Figure 5 legend (line 250), as this abbreviation was previously missing. Panel J is a close-up of the rasping teeth.

      (6) F3S and Video 3. Are the images in B and C taken from the FIB SEM video images? It is not clear. A small legend descriptor for video 3 would be helpful.

      The images in Supplementary Figure 3B and C are reconstructed from the same FIB-SEM dataset shown in Video 3, but they are displayed in a different orientation. This is indicated schematically in Supplementary Figure 3A, which illustrates the viewing plane used for the reconstruction.

      We already included the following legend for Video 3 (lines 1146–1149):

      "Video 3: FIB-SEM of the tsetse labellum. Sequential cross sections reveal internal ultrastructure progressing from near the tip of the labellum downward. Data were acquired on a Crossbeam 540 (Zeiss) with the EsB detector in continuous milling mode."

      To improve clarity, we have now added a sentence to the video legend linking the figures to the video: (lines 1149-1150)

      “Reconstructed images from this dataset are shown in Figure 5A and Supplementary Figure 3B and C.”

      In addition, we have now explicitly cross-referenced Video 3 in the legends of Figures 5 and S3 to make the connection clearer for the reader.

      (7) Figure 7. These are amazing images, especially G-I.

      Thank you for this positive feedback, we appreciate it.

      (8) P24. It is really good to see that there is a difference in force penetration for full skin v dermal...this deserves a comment.

      We agree and have revised the text accordingly. We replaced:

      "Human skin substrates required the lowest penetration forces, with 0.97 mN for full-thickness skin equivalents, 0.67 mN for dermal equivalents, and 0.85 mN for native skin explants (Figure 8C, D)."

      With this (lines 363-367):

      "Human skin substrates showed the lowest penetration forces, with dermal equivalents requiring less force (0.67 mN) than full-thickness skin equivalents (0.97 mN), reflecting the additional mechanical resistance of the epidermal layer absent in dermal-only constructs. Native skin explants fell intermediate at 0.85 mN (Figure 8C, D)."

      (9) P26 Figure S5. Panel c, there seems to be a big scatter in the drinking time. Was there an outlier?

      Indeed, the observed scatter is due to a single fly with an unusually long drinking time of 184.44 seconds, which is approximately six times the median duration. We have verified the underlying data and found no indication of a measurement error; the value therefore remains included in the analysis. The data for the plots in Supplementary Figure 5 are also available in Supplementary Material 3.

    1. eLife Assessment

      This important work addresses a very relevant biological question: what is the cellular basis of wound healing? Using the Drosophila pupal notum as a model, the paper provides an elegant, thorough, descriptive characterization of syncytia-driven wound closure using state-of-the-art confocal live imaging of the pupal notum. The authors meticulously characterize the cell-cell fusion events during wound healing and inhibit cell fusion to show to that it is necessary to speed wound closure. This study provides convincing evidence that cell fusion allows actin resources at be partitioned to the leading edge.

    2. Reviewer #1 (Public Review):

      [Editors' note: this version has been assessed by the Reviewing Editor without further input from the original reviewers.]

      Summary:

      This study aims to understand how cell fusion contributes to wound healing using a laser-induced injury in the notum epithelium of a developing fruit fly. The authors meticulously characterize the epithelial fusion events using a live imaging approach and report that syncytia arise by 'border breakdown' and 'cell shrinking'. The syncytial epithelial cells also appear to outcompete mononucleated cells and preferentially dissolve their tangential borders, which correlates with the accumulation of actin at the leading edge.

      Strengths:

      The strength of this study is the authors' live imaging approach to capture these dynamic fusion events that are a fundamental yet poorly understood biological process.

    3. Reviewer #2 (Public Review):

      Summary:

      Overall, this study provides a thorough description of the formation of syncytia following wounding of the proliferation-competent diploid epithelium of the pupal notum. While this phenomenon has already been described briefly for this particular tissue by the Galko lab in Wang et al 2015, the authors provide a much more detailed description and characterisation of the process providing some novel insights (radial versus tangential border breakdown, cell shrinkage, timings, syncytia outcompeting mononucleated cells, etc.).

      Strengths:

      This paper provides an elegant, thorough, descriptive characterisation of syncytia-driven wound closure using state-of-the-art confocal live imaging of the pupal notum. The authors show that laser-induced wounding of this diploid, proliferation-competent epithelium results in the formation of syncytia of various sizes in the first few cell rows around the wound edge, which progressively become bigger as healing proceeds. This results in ~50% of cells becoming part of these syncytia. The cell fusion events were convincingly demonstrated by showing the disappearance of p120ctnRFP and E-Cadherin-GFP from cell-cell borders as well as cytoplasmic GFP mixing of GFP-positive cells with a GFP-negative cell.

      Apart from cell-cell fusion by border breakdown that mostly happens in the first 2h following wounding, the authors also found that at later stages of wound healing cell shrinkage following cytoplasmic mixing contributed to syncytia formation.

      Next, the authors provided some convincing evidence that syncytia outcompete mononuclear cells for being positioned in the first cell row around the wound.

      The authors then show that radial border breakdown occurs much less frequently than tangential border breakdown. They suggest that radial border breakdown reduces the requirement for cell-cell intercalations. They also hypothesise that tangential border breakdown might allow fused cells to share resources and provide more resources to be used near the wound edge, e.g. for actomyosin cable formation. To test this, the authors generate single-cell clones that overexpress Actin-GFP. They then show convincingly how a single Actin-GFP-positive cell in the second cell row fuses with one GFP-negative cell in the first cell row. The Actin-GFP signal then spreads in the fused cell and labels some previously unlabelled actin-rich structure near the wound edge which most likely is the actomyosin cable. This provides some evidence for resource sharing by cytoplasmic mixing following fusion.

    4. Reviewer #3 (Public Review):

      In this revised manuscript, White et al. aimed to understand the wound-induced syncytia formation behavior in wound repair of Drosophila melanogaster pupal notum. For this purpose, the authors characterized two different types of adherens junctions' outcomes during syncytia formation around the wound region - border breakdown versus apical shrinking which appear to happen in different time points and for different time durations. The authors characterized cell-cell fusion events using cytoplasmic, junctional and nuclear markers. They determined that about half of the cells within 70 um radii from the wound undergo cell-cell fusion. They studied wound induction on the border between control epithelia and pnr domain suggesting that Atg1 is required for post-wound syncytia formation and wound closure. They showed that during wound closure syncytia gradually invade the wound leading edge mostly by radial fusion events. The data suggests that intercalation of cells from the leading edge slows down the wound closure process. They propose that cell fluidity of syncytial cells plays a role in wound closure speed. Finally, the authors showed that actin is concentrated to the front edge of syncytia located in the wound leading edge. The authors described some aspects of syncytia formation during wound closure using different approaches.

    5. Author response:

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #1 (Public Review):

      Summary:

      This study aims to understand how cell fusion contributes to wound healing using a laser-induced injury in the notum epithelium of a developing fruit fly. The authors meticulously characterize the epithelial fusion events using a live imaging approach and report that syncytia arise by 'border breakdown' and 'cell shrinking'. The syncytial epithelial cells also appear to outcompete mononucleated cells and preferentially dissolve their tangential borders, which correlates with the accumulation of actin at the leading edge.

      Strengths:

      The strength of this study is the authors' live imaging approach to capture these dynamic fusion events that are a fundamental yet poorly understood biological process.

      Comments on revised version.

      The manuscript overall is significantly improved and authors addressed majority of my concerns. The addition of the computational vertex model (Figure 7) as well as Atg1 RNAi (Figure 4) to inhibit cell fusion provide more mechanistic insight to their study. However, the analysis of Atg1 RNAi wound assay falls short as it does directly measure changes in syncytium frequency nor size to confirm that cell fusion is reduced. The authors should quantify the number of nuclei per syncytium over the 2hr wound healing period as performed for WT in Figure 1C. It would have been ideal if they could have also performed the Act-GFP spreading assay in WT and Atg1 RNAi strains to determine if Act-GFP movement is dependent on cell fusion as purposed. At the least, further quantification of Atg1 RNAi phenotype is warranted to support their conclusions.

      In response to the reviewer's comment, we have repeated the analysis of Fig 1C and generated a new panel, Fig. 4C, which is directly comparable to the control and shows that syncytial size is dramatically reduced in the Atg1 knockdown area. Unfortunately, we cannot perform the second analysis of actin-GFP spreading in the Atg1 knockdown cells because we need Gal4 for labeling individual cells and for knocking down Atg1, and we can't do both at the same time.

      Reviewer #2 (Public Review):

      Summary:

      Overall, this study provides a thorough description of the formation of syncytia following wounding of the proliferation-competent diploid epithelium of the pupal notum. While this phenomenon has already been described briefly for this particular tissue by the Galko lab in Wang et al 2015, the authors provide a much more detailed description and characterisation of the process providing some novel insights (radial versus tangential border breakdown, cell shrinkage, timings, syncytia outcompeting mononucleated cells, etc.).

      Strengths:

      This paper provides an elegant, thorough, descriptive characterisation of syncytia-driven wound closure using state-of-the-art confocal live imaging of the pupal notum. The authors show that laser-induced wounding of this diploid, proliferation-competent epithelium results in the formation of syncytia of various sizes in the first few cell rows around the wound edge, which progressively become bigger as healing proceeds. This results in ~50% of cells becoming part of these syncytia. The cell fusion events were convincingly demonstrated by showing the disappearance of p120ctnRFP and E-Cadherin-GFP from cell-cell borders as well as cytoplasmic GFP mixing of GFP-positive cells with a GFP-negative cell.

      Apart from cell-cell fusion by border breakdown that mostly happens in the first 2h following wounding, the authors also found that at later stages of wound healing cell shrinkage following cytoplasmic mixing contributed to syncytia formation.

      Next, the authors provided some convincing evidence that syncytia outcompete mononuclear cells for being positioned in the first cell row around the wound.

      The authors then show that radial border breakdown occurs much less frequently than tangential border breakdown. They suggest that radial border breakdown reduces the requirement for cell-cell intercalations. They also hypothesise that tangential border breakdown might allow fused cells to share resources and provide more resources to be used near the wound edge, e.g. for actomyosin cable formation. To test this, the authors generate single-cell clones that overexpress Actin-GFP. They then show convincingly how a single Actin-GFP-positive cell in the second cell row fuses with one GFP-negative cell in the first cell row. The Actin-GFP signal then spreads in the fused cell and labels some previously unlabelled actin-rich structure near the wound edge which most likely is the actomyosin cable. This provides some evidence for resource sharing by cytoplasmic mixing following fusion.

      Comments on revised version:

      The authors have extended their original manuscript by adding two key parts. First, they show a role of Atg1 in mediating cell fusion (Figure 4). Second, they provide additional evidence for a contribution of radial border fusions to wound closure through its effect on tissue fluidity and through computational modelling (Figure 7).

      This new version of the manuscript is greatly improved and provides significant new insights into the role of syncytia in aiding wound repair. There are just a few minor, yet important, additions needed to back up Figure 4 which should not require new experiments.

      Minor but important points:

      The authors show a role of Atg1 in mediating syncytia formation in Figure 4. However, since the Pnr>+ side of the wound closes slower than the non-Pnr side (control side), a few additions to this figure would be important and should not require additional experiments.

      (1) The authors should show, similar to the data shown in Figure 4D of the wound radius over time for control versus Pnr>Atg1RNAi, also the same type of data for control versus Pnr>+.

      The data the reviewer requests is available in our bioRxiv manuscript, in Fig. 6B (Hua, Krystofiak, Pumford, Page-McCaw, and Hutson, https://doi.org/10.64898/2026.05.31.728998). These experiments were all done at the same time. As you can see, the difference in closure rate is quite subtle in control wounds.

      (2) Since Pnr>+ also slows down wound healing, albeit to a lesser extent than Pnr>Atg1, the authors should also show an extra graph that provides evidence that Pnr>Atg1RNAi reduces syncytia formation more than Pnr>+ does. E.g. Two graphs could be added that show individual cell size at 4 or 5h post wounding for control versus Pnr>Atg1RNAi as well as for control versus Pnr>+ and also another graph with the same data but comparing cell size between Pnr>+ and Pnr>Atg1RNAi. Otherwise, if the expected minimum cell size for a syncytium is easy to estimate, a graph could be added that shows the percentage of cells that are above this threshold (e.g. above 100 square micron) for control versus Pnr>Atg1RNAi and control versus Pnr>+ and Pnr>+ versus Pnr>Atg1RNAi.

      In response to this comment and the comment from reviewer 1, we have now added new Fig. 4C, which addresses the reviewer's question about the comparative frequency of fusion in pnr>Atg1RNAi and pnr>+. These graphs show that Atg1 knockdown significantly reduces the size of syncytia.

      Reviewer #3 (Public Review):

      In this revised manuscript, White et al. aimed to understand the wound-induced syncytia formation behavior in wound repair of Drosophila melanogaster pupal notum. For this purpose, the authors characterized two different types of adherens junctions' outcomes during syncytia formation around the wound region - border breakdown versus apical shrinking which appear to happen in different time points and for different time durations. The authors characterized cell-cell fusion events using cytoplasmic, junctional and nuclear markers. They determined that about half of the cells within 70 um radii from the wound undergo cell-cell fusion. They studied wound induction on the border between control epithelia and pnr domain suggesting that Atg1 is required for post-wound syncytia formation and wound closure. They showed that during wound closure syncytia gradually invade the wound leading edge mostly by radial fusion events. The data suggests that intercalation of cells from the leading edge slows down the wound closure process. They propose that cell fluidity of syncytial cells plays a role in wound closure speed. Finally, the authors showed that actin is concentrated to the front edge of syncytia located in the wound leading edge. The authors described some aspects of syncytia formation during wound closure using different approaches. Some clarifications are needed as described below.

      Major suggestions:

      (1) Introduction, page 4. The examples of developmental syncytia formation of invertebrates and vertebrates are confusing. The authors may want to make the examples clear and add additional examples. Currently, readers may assume that C. elegans cell fusions occur only in the hypodermis - other structures can be mentioned like the vulva, pharyngeal muscles, glia, tail. In addition, the authors may want to add injury-induced fusions like the C. elegans' PLM and PVD neurons (Ghosh-Roy et al., 2010; Newman et al., 2015; Oren-Suissa et al., 2017).

      We appreciate the suggestions and have included the additional examples of C. elegans vulva and PLM and PVD neurons. We are limiting ourselves to those because we don't want to focus too heavily on C. elegans examples, as that's not the direction this paper is heading.

      (2) In cases where it is not clear whether fusion has occurred or whether mononucleated cells were ejected from the leading edge, membrane markers can be used. Page 6. Lines 96-99. The authors may want to use a membrane marker like RFP-PH driven by the epithelial cell promoter.

      At this point in the manuscript, we are introducing syncytia and are not concerned yet with their origin. 

      (3) Pages 8-10. The authors may want to clearly explain that apical junctions shrinking is a post fusion event. That the apical shrinking is caused by the expansion of fusion pores and the migration of apical junctions towards the basolateral domain. This is something that was clearly shown during physiological epidermal cell-cell fusion in C. elegans by Mohler et al., 1998 and 2002. A cartoon showing the process of cell-cell fusion, pore expansion and apical junction dynamics would make the manuscript much clearer.

      Apical shrinking cannot be caused by the "migration of apical junctions towards the basolateral domain" because that is not what we observed -- rather, we observed labeled adherens junctions remaining at the apical surface while the area they enclose becomes smaller (shrinks). Further, despite close reading of the Mohler papers, it is not clear how similar the apical shrinking events of this manuscript are to the fusion events described there. Finally, we do not want to include a schematic describing this process because that would suggest certainty that we do not have. Unlike in C. elegans development, wound-induced cell fusion is stochastic, not stereotyped; with cells that display apical shrinking, the fusion partner of a labeled cell is difficult to identify because it is often not a neighboring cell. These factors make it difficult to describe this process in detail, but we have sufficient data to conclude that these are indeed cell fusion events.

      (4) Page 9. Line 170. "...as these cells represent fusion initiation events (fusion pore) but were unable to productively stabilize and expand the site of fusion and so returned to the diploid state." The authors may want to make clear that this is an assumption that needs to be tested. Live imaging using a membrane marker may resolve whether a reversible fusion pore was generated.

      Thank you for the suggestion; we have updated this text to make it clear that this is an interpretation.

      (5) Page 11. It is not clear whether Atg1 is directly required for cell fusion, or that autophagy is required for efficient cell fusion or both Atg1 and autophagy participate in the fusion process.

      Our data show that Atg1 is required for cell fusion. The work that inspired this experiment, Kakanj et al 2022, concluded from their more comprehensive studies that the process of autophagy was required. We have clarified the text.

      (6) Page 12. Line 235. "Indeed, we observed that several hours after wounding, the entire leading edge was occupied by syncytia." This observation is based only on the adherens junction marker. Can they test basal cell membrane marker? Is it possible that the mononucleate cell in the leading edge is under the two syncytia?

      Unfortunately, there are not good basal markers -- the recently reported basal spot markers also label adherens junctions. Nonetheless, we are confident that the mononuclear cell is not under the syncytia because we image Z-stacks and thus can detect cell overlap.

      Recommendations for the authors:

      Reviewer #3 (Recommendations For The Authors):

      Minor suggestions:

      (1) Figure 1. The authors may want to add an image immediately after laser ablation of the actual wound and the area around the wound. Add an arrow to mark the wound.

      With this wounding modality, the extent of the wound is unclear for ~30 min. As we reported in O'Connor et al, PLoS One, 2021, there is a gradient of damage emanating out from the center of the wound, and cells with greater amounts of damage die while those with less damage repair and survive. Immediately after laser ablation, very little visible damage is evident by 120ctn-RFP and Histone-GFP (the markers in Fig. 1) until the cells die and the surrounding cells respond.

      (2) Page 6. Line 86. "A mitotic tissue utilizes cell-cell fusions during wound repair." replace "during wound repair" with "after wound induction" since in this section the authors do not show that this process is part of wound repair.

      Thank you for the suggestion - we reworded this heading to remove "wound repair".

      (3) Page 6. Line 92. The authors may want to be consistent with the terms used in the text and in the figure - His2GFP in the text versus Histone GFP in the figures.

      Thank you for the suggestion, we have revised for consistency.

      (4) Figure 1 - supplement figure 1D. The "v" of Div panel moved below D.

      Thank you, we have corrected it.

      (5) Figure 1 - supplement figure 1G. add "i" to second Gii to make it Giii.

      Thank you, we have corrected it.

      (6) Page 24. Figure 1H legend. 3 or 4 wounds?

      Thank you for catching this error - 4 wounds.

      (7) Page 7. Line 124. "GFP mixing always preceded border breakdowns (n=11)" instead of "always" use "in all observed cases".

      We have made this change.

      (8) Figure 2. Switch the writing "Apical Shrinking: Nuclear Transfer" since apical shrinking represented in panel 2A and Nuclear Transfer in panel 2B. If this description applies only to panel 2B, make it clear.

      We consider this heading to apply to panels A and B together (as they show the same sample, just different channels).

      (9) Figure 2C. Is ActinGFP a cytoplasmic GFP driven by actin promoter or Actin-bound GFP? Cytoplasmic GFP versus membrane-cortex GFP?

      It is a transgene expressing an actin-GFP fusion protein, as noted in the key reagents table and discussed in Fig. 8. We corrected the manuscript to ensure it is always referred to now as Actin-GFP in the text, figures, and legends.

      (10) Video 3 - Impressive movie!

      Thank you!

      (11) Page 9. Line 155. "In both these cells, as the cell lost its basal volume, cytoplasm moved laterally to join the neighboring syncytia." It seems that the apical shrinking cells' cytoplasm joined the neighboring syncytia even before.

      Because both indicated cells (yellow and white arrows) and the neighboring syncytium are all labeled with GFP, it is not possible to determine precisely when the cells' cytoplasm joined the syncytium.

      (12) Page 9. Line 158. "...but fusions associated with apical shrinking occurred later and were more numerous." Did the fusion occur later or the apical shrinking itself as was mentioned before and shown in Figure 2F?

      We have changed the wording, as for many apical shrinking events we cannot tell exactly when the fusions were initiated.

      (13) Page 25. Figure 3A legend. What is the meaning of morphological fusion? Border breakdown and apical shrinking? The authors may want to define it.

      We have defined it now in the legend.

      (14) Page 26. Figure 3B-C legend. "Panel C shows that apical shrinking fusion and border-breakdown fusion occur at similar distances from the wound." It seems that fusion by apical shrinking mostly occurs within 60-70 um from wound center and fusion by breakdown occurs equally at all distances up to 80 um.

      We don't disagree with your comment, but we feel the dataset is too small to make such a statement. The data is presented so the interested reader can make their own conclusion.

      (15) Page 9. Line 165. "...but infrequently (n=3) with GFP mixing and no subsequent cell fusion..." Does this mean that there were GFP mixing without border breakdown or apical shrinking?

      Yes, that is correct. We assume that in this case a fusion pore opened and then closed again. We have added a phrase to clarify.

      (16) Page 9. Line 175. "...the spatial distribution of fusing cells that shrank vs. lost borders was similar (compare Figures 1G and 2E)." Even though the visual comparison suggests similar spatial distribution, the carefully quantified distribution in figure 3C suggests more fusion by shrinkage at 60-70 um from wound center of the 5 tested wounds.

      As we noted to comment 14, we feel the data set is too small to make such a statement. The data is presented so the interested reader can make their own conclusion.

      (17) Figure 3. The shown pies sum the results from 5 wounds. It would be interesting to add a graph comparing the percentage of fused and persisted cells per wound to see the variability, if exists.

      Unfortunately, the number of fused/persisting cells in each wound is greatly affected by the heat-shock conditions that generate the labeled clones; even the ratio of these fates would be heavily influenced by noise because the numbers are small in each animal. Further, the frequency of fusion is determined by the wound size as shown in Fig. 1. Because of these variables, such data could be easily misinterpreted.

      (18) Figure 3 - figure supplement 1D. Even though it was mentioned that the duration of some border breakdown is finished within minutes it is worth comparing it with shrinking duration on one graph.

      Unlike apical shrinking, it is difficult to identify exactly when border breakdown concludes, so this data is difficult to compare. We have provided several examples of border breakdown in the manuscript that give an overview of the process.

      (19) Video 1 is not mentioned in the main text.

      Thank you for catching that omission. We now refer to it in the first paragraph of the results.

      (20) Figure 4B. The difference between the treated group and the control group is unclear. Add arrows.

      We have added some arrows to Fig. 4B.

      (21) Figure 4C. For consistency use percentage for both border breakdown and shrinking cells.

      In response to the reviewer's comment, we now provide the consistent metric of number of lost borders and number of shrinking cells.

      (22) Page 11. Did the authors try other wound types (e.g. mechanical/chemical wounds)? May other wound causes besides laser ablation result in different response? This may help to answer whether there is a causation between syncytia formation and speed wound closure.

      There are reports of puncture and pinch wounds inducing cell fusion. Perhaps the reviewer is suggesting that we might be able to identify a wounding method that does not induce cell fusion and then compare the rate of wound closure. However, another type of wound would probably inflict different amounts of cell damage and so would be hard to compare. Overall, we think the half-and-half system of comparing responses on the two sides of the wound is the best, most controlled comparison.

      (23) Figure 4F-G. It was mentioned that there is less syncytia formation in Atg KD cells, however the difference in cell area between control, WT and Atg KD is not obvious. The authors may want to mark the dots that represent syncytia to distinguish them from mononucleated cells.

      The point we are trying to make (now Fig. 4G-H) is that cell area is related to distance moved, regardless of how cell area is determined. We do not have the ability to count nuclei in the control sides (nuclei are labeled only on the pnr side), and further, we have reported separately (White et al, 2024) that there is a limited amount of endocycling in these cells, which should also increase area.

      (24) Figure 5G. y axis. The authors may want to change "small cells" to "mononucleate cells".

      We changed it to "unfused cells" which is the term we used in the legend. In these wounds we were unable to visualize nuclei.

      (25) Page 12. Line 243. (Figure 5D,G) instead (Figure 5D).

      We changed it to read (Figure 5D,G).

      (26) Page 12-13. Lines 241-246. The description of "mononuclear cells removed" and "syncytia outcompete unfused cells" may be clearer if explained here as mononuclear cells joining the syncytium by cell-cell fusion.

      Here we are describing a different phenomenon - not that fusion is removing all the smaller cells but rather that the syncytia are faster/better/more effective at wound closure than the smaller cells. This is illustrated in Fig. 5Cii-Ciii.

      (27) Page 13. Line 261. "Thus, there were about five-fold more tangential borders lost to fusion than radial" Is this conclusion also true when analyzing each wound individually?

      This is a reproducible finding, that there is more fusion across tangential borders than across radial borders. The ratio of tangential-border loss: radial-border loss for each wound is as follows:

      wound 1, 63:12

      wound 2, 39: 11

      wound 3, 44:8

      wound 4, 50:8

      (28) Page 31. Figure 6 - figure supplement 1 legend, Line 652. Make "B)" bold.

      Done.

      (29) Page 14. Lines 270-275. If there is an advantage to radial fusion versus cell intercalation for wound closure speed, how do the authors explain that the percentage of radial fusion is lower than the percentage of intercalation? (Figure 6D) How does the wound affect the molecular level (fusogen expression?) of the surrounding cells? 

      We expect that radial fusion specifically reduces the need for intercalation at the leading edge, as shown in Fig. 6C. Both would speed closure, however, as any increase in cell area will allow more efficient redistribution of resources such as actin and will also reduce the total number of junctions needing to be remodeled as the wound closes. Since we don't know the fusogen, we can't say how the wound affects its distribution.

      (30) Page 14. It is not clear where the experimental data ends and the model starts. For example, in line 276, it would be clearer to describe the "tissue fluidity as measured" or is it more precise to write instead "as estimated/calculated". The fusion between observations and model is confusing and maybe this should be unfused.

      This text, referring to the analysis in Fig. 7A, B, is not a computational model but rather a quantitative analysis of tissue fluidity as measured by a pre-existing metric, the shape index. This is experimental data. The computational model begins in the next paragraph, accompanying Fig. 7C, D. We edited the language slightly in this paragraph to clarify.

      (31) Figure 8 versus Figure 2C. Actin-bound GFP versus cytoplasmic GFP? Both mentioned as Actin GFP. Make it clear.

      They are indeed the same thing, actin protein fused to GFP, as described in the text and legend, and they are labeled identically.

      (32) Figure 8Biii, Div. Nice presentation of signal distribution between the cells.

      Thank you!

      (33) Page 30. Figure 8G legend. Lines 623-628. Is the shown mean profile plot based on specific images shown in Fi and Fv or just the cells represented there? Since Fi is a single z slice and Fv is maximum intensity projection which are not comparable.

      In response to the reviewer's question, we reanalyzed the image. Fig. 8G compares Z-projections.

      (34) Page 15. Line 303-304. "Tangential border fusions allow resources from distant cells to be mobilized to the wound edge." Does not this leading-edge actin localization happen in radial fusions close to the region of the wound?

      Fusions along radial borders, as shown in the top panel of Fig. 6A, would not offer the opportunity to move actin from distant cells to cells nearer to the wound.

      (35) Page 15. Did the authors test any predictions from the simulations of the model experimentally?

      This isn’t so much a predictive model as an exploratory model that addresses one question: is it plausible that the presence of syncytia can speed closure by reducing the need for intercalations, even if the syncytia have no other special properties. The only prediction would be that inhibiting fusion would slow wound closure.

      (36) Page 17. It would be interesting to discuss the following questions: (A) Is autophagy required for fusion. (B) Is Atg1 required for epithelial cell fusion? (C) Is autophagy required for wound repair? Are any of the combinations correct (A&B, A&C, B&C, A&B&C)

      The role of autophagy in wound-induced cell fusion was thoroughly explored in the 2022 EMBO J paper from Maria Leptin's lab, "Autophagy-mediated plasma membrane removal promotes the formation of epithelial syncytia" by Kakanj et al. We merely knockdown a gene they discovered to be important for wound-induced epithelial fusion, Atg1, as one means of investigating how syncytia contribute to wound closure. Our results don't add to their findings, and the role of autophagy is not what we want to focus on in our Discussion.

      (37) Page 19. Line 381-384. "If N represents the number of cells that fused, our results suggests that syncytia can apply up to N times more actin to the leading edge; considering that we observed syncytia with dozens of nuclei, this could represent a significant enhancement of actin at the leading edge. Increased actin might explain the ability of syncytia to outcompete diploid cells at the leading edge." To enhance this suggestion, the authors may want to compare actin signal in the leading edge of different size syncytia.

      We thought a lot about this experiment because reviewer 2 asked for it in the previous round of review, but as we said then, we can imagine too many caveats to the interpretation to make it worthwhile.

      (38) Page 22. Line 447-448. "...fusion would act the fastest after wounding because there is no need for DNA replication." There may be a potential need for protein (fusogen) synthesis.

      The timing of fusion, which we report here begins within 10 minutes after wounding, suggests that if there is a fusogen, it is already present in the cells before wounding.

      (39) Page 34. Line 716. Add "C" to "29{degree sign}".

      Done.

      (40) Page 38. "Wound closure analysis" part. Can the wound closure be visualized using brightfield?

      The scar also impedes imaging through bright-field microscopy.

    1. eLife Assessment

      This valuable study addresses the effects of selection for aggression on fitness and life-history trade-offs in Drosophila melanogaster. The evidence presented is overall solid, however, the data as they are do not completely support the claims of increased survival of highly aggressive males at the expense of reproductive success. The main limitation is the choice to use males from only one aggressive Drosophila line, that do not allow disambiguation between nonaggression-related factors and aggression-related factors influencing lifespan.

    2. Reviewer #1 (Public review):

      Summary:

      This study asks how selection for male aggressiveness affects life-history and reproductive fitness traits in Drosophila melanogaster males.

      Strengths:

      Multiple comprehensive assays are used to address the question.

      Weaknesses:

      (1) The flies used for comparisons are inadequate. Behavioral assays compare Bully males mated top non-coevolved Cs females with Cs males mated to coevolved Cs females.

      (2) Lifespan analysis is done on male progeny of Cs females mated to either genetically more distant Bully or co-evolved Cs males, the longer lifespan and performance on the former is interpreted as trade-off with aggressiveness, rather than a simple explanation of hybrid vigor.

      (3) Differences in CHCs between Bully and Cs males and Cs females mated to those males are not shown to cause difference in measured behavioral outcomes.

      Comments on revised version.

      I appreciate authors responding to reviewer's comments. The inclusion of additional Bully lines in behavioral analysis, and Bully homozygous male progeny in lifespan analysis gives more strength to the authors' conclusions. It does not exclude other possible explanations for the observed results, but now authors note genetic drift as an alternative explanation for some of their results.

      I do want to point to a potential misunderstanding of male-female co-evolution by authors. The authors state that "The Bully lines used in our work were derived from Canton-S flies and thus did co-evolve with Cs". This statement is incorrect if the process of selection and line maintenance in this study was the following:

      In my understanding to create Bully lines the most aggressive males were first chosen from an ancestral Cs line and their most aggressive male progeny were mated to their sibling females, repeating the process for 37 generations. Therefore, Bully females were co-evolving with Bully males during selection process of over 37 generations, while Cs females were staying co-evolved with their own males, since they mated within the line. Moreover, after aggressive lines were created, they were kept separate from each other, and from Cs line since about the year 2010, until the experiments described in the paper were performed (which must over 10 years?). Over 10 years, a significant genetic drift can happen, that changes allele frequencies, and may results in differences in male-female co-evolved traits and in lifespan that are unrelated to selection for aggression.

      Also, decapitating females does not completely prevent female influence over mating process, but just removes central brain control over it. In Drosophila, however, the main control over copulation process for males and female is not central. Therefore, you do not completely remove the effect of coevolved or non-coevolved female traits over copulatory and post-copulatory processes.

    3. Reviewer #2 (Public review):

      Summary:

      The authors compare "Bully" lines, selected for male aggression, to Canton-S controls and find that Bully males have lower mating success, shorter mating durations, and remate sooner. Chemical analyses show Bully males have distinct cuticular hydrocarbon (CHC) signatures and transfer markedly less cVA to females, offering a plausible mechanistic link to weaker mate-guarding. Paradoxically, Bully males live longer and remain fertile at older ages when Cs males no longer mate, indicating a shift in the reproduction-survival trade-off in aggression-selected populations. Importantly, the work sheds light on proximate mechanisms, demonstrating that shifts in CHCs and pheromone transfer co-occur with changes in fitness traits.

      Strengths:

      The manuscript's strengths lie in its comprehensive and integrative approach framed within an evolutionary context. By combining behavioral assays, chemical profiling, and lifespan measurements, the authors reveal a coherent pattern linking aggression selection to life-history trade-offs. The direct quantification of cVA in the female reproductive tract after mating provides a particularly compelling mechanistic correlate, strengthening the link between behavior and chemical signaling. Findings on altered 5-T and 5-P levels further highlight how chemical communication shapes mating and mate-guarding strategies. Analytical approaches are largely rigorous, and the results provide valuable insights into the pleiotropic effects of selection on socially relevant traits.

      The revision responds directly to the main concerns raised previously. The addition of a third, independently selected line (Bully C), together with the Bully × Bully data, considerably reduces the concern that the behavioral phenotypes reflect line-specific drift or founder effects rather than a correlated response to selection. The reorganized survival figure (Figure 5) is a clear improvement over the previous version, with isolated and group-housed males in separate panels and a heterozygous Bully condition added, so the longevity claim can be evaluated more directly. The isolated-male data are especially useful here, since those flies never mate, and a longevity difference under that condition argues that the effect is not simply a consequence of Bully males mating less often. The behavioral schematic, corrected symbols, and reported sample sizes also help, as does the reinterpretation of the post-mating courtship data in terms of courtship motivation rather than a refractory-period effect once no latency difference was found.

      Weaknesses:

      Most of the remaining weaknesses are ones I raised in the first round, and the revision has narrowed them. The links between the altered CHC profiles, the reduced cVA transfer, and the behavioral outcomes remain correlative. The causal experiments that would establish them (for example, perfuming or cVA-equalization) are acknowledged by the authors as future directions, which is reasonable, but it means the mechanistic claims should be read as candidate explanations rather than demonstrated ones. It is also worth noting that the CHC differences and the behavioral differences may both be downstream of a common selection target (for instance, genes affecting oenocyte function or CHC biosynthesis) rather than one causing the other; the Discussion would be more balanced if this alternative were stated explicitly.

      My main remaining concern is with the lifespan data. The behavioral phenotypes are replicated across Bully A, B, and C, but the survival assays were done on Bully A only, so a line-specific contribution to the longevity result, including drift, cannot be excluded, even though this has been addressed for the behavioral traits. This matters because the title and abstract present the survival-reproduction trade-off as a general consequence of selection for aggression, whereas the survival evidence rests on a single line. The authors can either run the lifespan assays on a second line, or calibrate the text, title, and abstract so that the strength of the survival claim matches the single-line evidence behind it, with second-line lifespan data noted as a future step.

      The Bully C line is currently underused. Its intermediate aggression, together with the absence of a significant reduction in mating duration, points to a graded rather than binary relationship between aggression intensity and mating duration. This is one of the more interesting features of the expanded dataset, and it deserves more than its present role as a justification for focusing on Bully A.

      The authors have appropriately softened causal language in the title, subheadings, and much of the Discussion. A few residual passages still imply causation or directional transfer and would benefit from the same treatment.

    4. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment:

      This valuable study addresses the effects of selection on aggression on fitness and life-history trade-offs in Drosophila melanogaster. However, the evidence presented is incomplete and does not support the claims proposed in the study of increased survival of highly aggressive males at the expense of reproductive success and shorter mating duration. The main limitation of the study is the choice to use males from only one aggressive Drosophila line in combination with Canton-S females, that do not allow disambiguation between nonaggression-related factors, such as hybrid vigor and aggression-related factors influencing mating and lifespan.

      We would like to clarify the points raised in the eLife assessment.

      The report states that we relied on a single line of hyper-aggressive males tested with Canton-S females, and implies that Bully and Cs have not co-evolved. This is a misunderstanding: Bully flies were derived from Cs population. Thus, Bully and Cs have co-evolved. In addition to the Bully A line presented in the main figures of the manuscript, we replicated several of our findings with a second independent selected line, Bully B. Results from courtship assays involving both Bully A and Bully B couples males and females were presented in Figure Supp1. We apologies for not having made this more explicit in the original manuscript, which we will correct. These experiments should alleviate the concerns from the reviewers; they demonstrate that our conclusions are supported by two independent hyper-aggressive lines, and these include assays with selected male and female flies.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This study asks how selection for male aggressiveness affects life-history and reproductive fitness traits in Drosophila melanogaster males.

      Strengths:

      Multiple comprehensive assays are used to address the question.

      We thank the reviewer for recognizing these strengths.

      Weaknesses:

      (1) The flies used for comparisons are inadequate. Behavioral assays compare Bully males mated to non-coevolved Cs females with Cs males mated to coevolved Cs females.

      We thank the reviewer for this comment, which made us realize that we had not sufficiently highlighted some of our experiments. The Bully lines used in our work were derived from Canton-S flies and thus did co-evolve with Cs. As originally described by Penn et al. (2010), highly aggressive “Bully” lines were generated through selective breeding from Canton-S males that consistently won aggressive encounters. After 34–37 generations, stable Bully lines were established. Thus, 1) Bully and Cs flies have co-evolved and 2) the selection applied was male-specific. Independent selection replicates produced distinct lines, including Bully A and Bully B. Previous studies only characterized Bully A (Penn et al., 2010; Chowdhury et al., 2017), but our work includes both Bully A and Bully B (Fig. S1).

      The rationale for pairing Bully or Cs males with Cs females (with which both male types co-evolved) follows the approach used by Dierick et al. (2006), who investigated how the male-specific selection for aggression affected courtship and mating behaviors by testing them with standard Canton-S females. This design allows to isolate the effects of male genotype and behavior on courtship and mating outcomes, avoiding confounding effects from female behavioral changes.

      We initially compared selected Bully pairs (Bully males × Bully females) (Fig. S1) with Cs pairs and observed similarly shortened mating durations in both Bully × Bully and Bully × Cs matings (Fig. S1, Fig. 1F and G). Thus, the reduction in mating duration arises specifically from Bully males. We therefore chose to use Cs females as a standard background to assess the consequences of male-specific selection for aggression on reproductive behaviors.

      (2) Lifespan analysis is done on male progeny of Cs females mated to either genetically more distant Bully or co-evolved Cs males; the longer lifespan and performance on the former is interpreted as a trade-off with aggressiveness, rather than a simple explanation of hybrid vigor.

      We appreciate this comment, which again stems from a poor explanation from our part about the origin of the Bully line in the original manuscript. The Bully flies were derived from the same original population as the Cs line. Hybrid vigor typically arises when crossing individuals from distinct populations, which is not the case here as both Bully and CS come from the same population.

      To further support our conclusions, we conducted additional experiments using progeny from within-line crosses (Bully males × Bully females) and results revealed the same phenotype: the progeny of these flies also exhibited significantly longer lifespans than Cs males x Cs females progeny. This finding argues against hybrid vigor as the main explanation for the observed phenotype, since both the Bully and Cs crosses result in inbreeding, yet give longer lifespan in Bully. We will include these additional longevity data (currently not included in the manuscript) to strengthen our results and reinforce our interpretation.

      (3) Differences in CHCs between Bully and Cs males and Cs females mated to those males are not shown to cause differences in measured behavioral outcomes.

      We thank the reviewer for raising this important point regarding causality. One way to establish a causal link between differences in CHCs observed in Bully and Cs flies and the corresponding behavioral outcomes would be to experimentally manipulate CHC profiles. For instance, one could perfume oenocyte-less males with the compounds found in higher abundance in Bully flies, then perform behavioral assays to assess causality. We agree that such experiments would be highly informative in determining the functional roles of specific CHCs elevated in Bully males. However, this approach is technically challenging, as the perfuming technique must be optimized to transfer precise amounts of each compound. For example, this method can be used to gradually perfume flies to assess dose–response behavioral effects, whereas matching exactly the natural concentrations found in individuals, especially given inter-individual variability, remains difficult.

      We considered conducting such experiments during our study but did not pursue them for these technical reasons. Nevertheless, we can include a statement in the Discussion acknowledging this as an important future direction to test the causal relationship between CHC variation and behavior.

      Reviewer #2 (Public review):

      Summary:

      The authors compare "Bully" lines, selected for male aggression, to Canton-S controls and find that Bully males have lower mating success, shorter mating durations, and remate sooner. Chemical analyses show Bully males have distinct cuticular hydrocarbons (CHC) signatures and transfer markedly less cVA to females, offering a plausible mechanistic link to weaker mate-guarding.

      Paradoxically, Bully males live longer and remain fertile at older ages when CS males no longer mate, indicating a shift in the reproduction-survival trade-off in aggression-selected populations.

      Importantly, the work sheds light on proximate mechanisms, demonstrating that shifts in CHCs and pheromone transfer co-occur with changes in fitness traits, thus offering new entry points for understanding life-history evolution.

      We thank the reviewer for this positive summary of our work.

      Strengths:

      The manuscript's strengths lie in its comprehensive and integrative approach framed within an evolutionary context. By combining behavioral assays, chemical profiling, and lifespan measurements, the authors reveal a coherent pattern linking aggression selection to life-history trade-offs. The direct quantification of cVA in female reproductive tracts after mating provides a particularly compelling mechanistic correlate, strengthening the link between behavior and chemical signaling. Findings on altered 5-T and 5-P levels further highlight how chemical communication shapes mating and mate-guarding strategies. Analytical approaches are largely rigorous, and the results provide valuable insights into the pleiotropic effects of selection on socially relevant traits. The study will be of interest to Drosophila biologists working on sexual selection, behavioral evolution, and aging.

      We thank the reviewer for recognizing the integrative design and mechanistic contributions of our study.

      Weaknesses:

      The weaknesses are primarily conceptual rather than procedural. The generality of the findings is uncertain, as selection appears to be represented by only one (and a second closely related) Bully line, limiting conclusions about selection responses versus line-specific drift or founder effects. The causal link between aggression selection and increased longevity is not established: the data show a correlated shift but do not identify mechanisms underlying lifespan extension. In several places, the manuscript uses causal language (e.g., that selection 'influences' longevity or mating strategy) where association would be more accurate; this should be toned down to avoid overstatement. Ecological relevance is also not addressed, since laboratory conditions may bias the balance between costs and benefits of aggression compared with variable natural environments. Addressing these points would strengthen both the impact and clarity of the study.

      (1) Generality of findings and potential line effects

      We agree that our results presented in the main figures of the manuscript relied mainly on one Bully line (Bully A). To address potential line-specific effects, we replicated key courtship experiments with another independent line, Bully B, selected in parallel from the same Canton-S stock but through distinct selection replicates. The results obtained from Bully B closely matched those from Bully A, suggesting that the observed phenotypes are consistent consequences of aggression selection rather than random drift or founder effects.

      (2) Causality versus correlation

      We concur that some sentences in the manuscript could overstate causal interpretations. We will revise the text to clearly distinguish correlation from causation and to avoid implying direct causal relationships where data only support association.

      (3) Ecological relevance

      We appreciate this point. Our experiments were performed under controlled laboratory conditions, which may not fully capture the ecological contexts shaping the costs and benefits of aggression. We will acknowledge this limitation and expand the Discussion to consider how environmental variability could modulate the fitness trade-offs associated with aggression in natural populations.

      We thank both reviewers for their constructive feedback, which will help us strengthen the rigor and clarity of the manuscript. We believe that the additional results and revisions will satisfactorily address their concerns.

      Recommendations for the authors:

      Reviewing Editor Comments:

      The major weaknesses raised by the reviewers, namely the flies used in the study (CsxBully compared to CsxCs) where the effect on lifespan could be explained by hybrid vigor and the use of only one Bully and one Cs line that does not allow to link unambiguously the observed effect to the selection for aggression, should be addressed by a different experimental design and additional lines to exclude the effect of non-aggression related factors.

      We thank the Reviewing Editor for these comments.

      (i) Experimental design and hybrid vigor:

      Hybrid vigor typically arises from crosses between genetically divergent populations. In our study, Bully lines were derived from Canton-S background and are not therefore not genetically distant from controls. To directly address this concern, we included new data from Bully × Bully pairs (Figure 1), using independently selected Bully lines. These experiments reproduce the key aggression and courtship phenotypes observed in Cs × Bully assays, indicating that the effects are not attributable to hybrid vigor.

      (ii) Use of additional selected lines:

      We now include data from two independently selected lines (Bully A and Bully B), both derived from Cs, which show consistent behavioral phenotypes. This supports the conclusion that the observed effects are associated with selection for aggression rather than line-specific artifacts. We note that generating such lines is time- and labor-intensive, and only a few laboratories have established aggression-selected lines in Drosophila melanogaster (e.g., Penn et al., 2010; Dierick et al., 2006; Edwards et al., 2006). Accordingly, we have revised the manuscript to explicitly acknowledge this limitation and to frame our conclusions in terms of association rather than causation.

      Reviewer #1 (Recommendations for the authors):

      I can't see any way to interpret the data using CsxCs vs CsxBully comparisons.

      We thank the reviewer for this important point. This concern appears to arise from the assumption that Bully and Cs represent genetically distinct or non-coevolved populations. However, Bully lines were directly derived from Cs and therefore share a common genetic background. We have clarified this point in the Introduction (lines 104-106) and Results (lines 132-137).

      Importantly, we now include additional data showing that key phenotypes, including reduced mating duration, are also observed in Bully × Bully pairings and across independently selected Bully lines (new Figure 1). These results demonstrate that the observed effects are driven by the male genotype and do not depend on the female background.

      Because selection for aggression was applied specifically to males, we used Cs females as a standardized background to isolate male-specific effects while minimizing variability arising from female genotype or behavior. This rationale is now explicitly stated in the Results (lines 163-166) and at the beginning of the Discussion (lines 326-330). This experimental design allows interpretation of male-specific effects, and the observed differences cannot be attributed to cross design artifacts or hybrid vigor.

      Reviewer #2 (Recommendations for the authors):

      Major comments:

      (1) Several passages currently imply causality, whereas the data support correlations between selection and trait differences rather than direct causation. This overstatement also appears in section subheadings within the Results, such as "Hyper-aggressive males display reduced mate-guarding efficiency, without compromising female fertility." Please consider toning down the wording by replacing active causal verbs with more neutral phrasing. Additionally, it would be important to include a clear, explicit sentence in the Discussion acknowledging this caveat, as the existing phrase "is associated with changes in reproductive traits" does not fully convey this nuance.

      We thank the reviewer for this important comment. We have revised the manuscript throughout, including Results subheadings, to replace causal language with association-based phrasing. We also rephrase the first sentence of the Discussion to clarify that our conclusions are correlational (see line 320).

      (2) Figure 1 would benefit from a simple schematic of the behavioral paradigm and the arena, since the authors' arena design minimizes manual handling; a cartoon would help readers quickly grasp the assay flow and the conditions under which interactions occur.

      We thank the reviewer for this helpful suggestion. We have added a schematic to Figure 1 illustrating the behavioral paradigm and arena design. Additional details are provided in the Materials and Methods (Trannoy et al., 2015). This improves clarity and accessibility of the experimental design.

      (3) In Figures 2A-B and A'-B', the higher post-mating UWE in Bully males is intriguing, but these panels do not actually measure the refractory period. It would be helpful to include 'latency' in the first UWE after mating in these swapped-female conditions. This could also be repeated with pheromone-standardized (cVA/CHC-equalized) decapitated females to disentangle effects of female pheromone load from male sensory perception. In addition, a baseline courtship control (naive males with decapitated virgins) is necessary to test whether Bully males simply have a lower threshold for initiating courtship.

      We thank the reviewer for this suggestion. The referenced panels are now shown in Figure 3. We quantified post-mating courtship latency; however, latencies were very short across conditions, and no differences were observed between genotypes. We therefore revised the text to interpret these results in terms of post-mating courtship motivation rather than refractory period. The baseline courtship control with decapitated virgins is provided in Fig 2G. These changes clarify the interpretation of post-mating behavior and address the reviewer’s concerns.

      (4) Related to my above point, the results in Figure 2B-B′ raise the possibility that Bully males have reduced perception or neural sensitivity to anti-aphrodisiac pheromones deposited by CS males, which could account for their elevated post-mating courtship; the authors might consider experiments that directly test male sensory responsiveness to these cues or mention this possibility in the Discussion.

      We thank the reviewer for this point. We performed additional assays to test males’ sensory responsiveness using binary choice assays and measured the time spent performing UWE towards decapitated females versus males. These results were added in Figure 3-Figure Supp 1, and indicate that both Cs and Bully males displayed courtship preferentially towards females, providing a control for sensory perception.

      (5) In multiple figure panels, virgin and mated females are depicted with the same symbols, which makes interpretation confusing.

      Thank you for pointing this out. We have updated the figure panels to use distinct symbols for virgin and mated females to improve clarity.

      (6) For Figure 4, it would be helpful to provide standalone KM curves for Bully versus CS males, including a separate panel for isolated (never-mated) males, and present mating counts in a separate panel while reporting survival models that incorporate mating frequency (or use it as a time-dependent covariate). Although Figures 4C-D report median survivals, full KM plots and an isolated-male curve are important since mating itself elevates mortality and can otherwise confound intrinsic lifespan differences.

      Thank you for this important point to improve clarity of this figure. We have reorganized this figure (now Figure 5) to now, present survival curves first (isolated and group-housed males), followed by lifetime mating counts. This reorganization separates survival from mating activity and addresses the potential confounding effect of mating on lifespan. For the lifetime mating data, we used bar plots rather than curves to better visualize individual mating events.

      (7) The following sentences overstate the results and imply causality; consider toning down: "These findings suggest that 5-P and 5-T might contribute to promoting remating in females that have previously mated with Bully males (Figure 2F' and G'). Given that Bully males also showed higher levels of both 5-P and 5-T compared to naïve Cs males (Figure 3B), it is likely that the elevated levels of these compounds observed in females result from their transfer during mating."

      We thank the reviewer for this important point. We have rephrased these sentences to remove causal language and instead describe associations between CHC profiles and behavioral outcomes. In particular, statements implying that 5-P and 5-T promote remating or are directly transferred during mating have been revised to reflect correlational evidence only (see lines 253-256).

      (8) It is not entirely clear how aggression was quantified in each generation, what proportion of males were selected to breed, and whether the findings generalize beyond a single Bully line (Figure Supplement 1 shows data from a closely related Bully line). Without independent replicate lines or sham-selected controls, it remains difficult to rule out drift or line-specific artifacts, and this limitation should be explicitly acknowledged.

      We thank the reviewer for this important point. We have clarified the aggression selection procedure by adding methodological details from Penn et al., including how aggression was quantified and how breeders were selected (lines 104-106 and 131-137). Briefly, independent selection replicates were initiated from the same Canton-S population, generating three lines (Bully A, B, and C), which were maintained separately.

      To address generality, we now include data from multiple lines. In particular, a new Figure 1 presents aggression and courtship phenotypes across Bully A, B, and C, and key behavioral results are consistent across independent lines.

      We acknowledge that additional independent lines would further strengthen generality; this limitation is now explicitly stated in the Discussion (lines 325-326).

      These additions clarify the selection procedure and support that the observed phenotypes are associated with aggression selection rather than line-specific artifacts.

      Minor Comments:

      (1) Exact sample sizes for every experiment should be included in the main figure legends.

      Thank you. We have added the number of replicates in each figure legends.

      (2) In Figure 1-Supplement 1, the orientation for depicting mating success is reversed compared to Figure 1, which is a bit jarring; it would be clearer to keep the orientation consistent with the main figure.

      Thank you. We have incorporated the results initially presented in Figure 1-Sup 1 into a new Figure 1 with additional results, and have taken into account reviewers’ comment.

      (3) For multivariate analyses, I suggest including important details such as group sample sizes, p-value, and the percent variance, etc., in the figure legend rather than keeping this only in Supplementary Table S1.

      We have inserted these details directly into the figure legends for clarity.

      (4) Why was the food cup used for arenas where decapitated virgins were used in mating assays?

      Thank you for pointing this. We now have clarified the experimental procedure in the M&M of the revised manuscript (lines 465-469).

      (5) For cartoons in Figure 2, the current yellow background makes it very difficult to distinguish flies drawn in yellow or green. Please adjust to a higher-contrast background or add darker outlines so that the cartoons are clearly legible.

      Thank you. We have increased the contrast of the female bodies to ensure the cartoons are clearly distinguishable (now figure 3).

      (6) Addition of line numbers in the manuscript would be helpful during the review process.

      Line numbers have been added throughout the manuscript.

      (7) I noticed a few typos in the manuscript. For example, in the Introduction, "seminal fuids" should be corrected to "seminal fluid." In the Discussion, the phrase "CHCs profiles compared those" requires a "to" before "those." Please carefully review the manuscript for similar errors.

      Thank you for pointing this out. We carefully reviewed the manuscript for typos and corrected all identified errors.

    1. eLife Assessment

      This potentially valuable study investigates the anti-senescence effects of red light exposure, proposing that reduced SIRT4 levels enhance fatty acid metabolism and H3K9ac, thereby attenuating ageing-related phenotypes. The authors use multiple approaches, including cultured cells, animal models, and molecular analyses, to support their conclusions. Following revision, many of the concerns previously raised by the reviewers have been addressed. The evidence is solid, whereas additional controls and stronger mechanistic data are still needed to fully substantiate the proposed pathway, particularly the mechanism by which red light exposure leads to reduced SIRT4 levels.

    2. Reviewer #1 (Public review):

      Summary:

      Deng and colleagues pursue the possibility that red light exposure can provide some benefits and anti-senescence effects in aged mouse models. In addition, they show how red light influence metabolism in cultured keratinocytes. The authors provide a long dissection of the potential paths involved in the changes promoted by red light exposure, identifying CytC oxidase, SIRT4, PPARa and MCD as key players.

      Strengths:

      The authors did a thorough exploration of the multiple potential avenues by which red light exposure influence metabolism. The in vitro and in vivo evidence nicely complement each other.

      Weaknesses:

      This is a challenging hypothesis that would require some additional experimental controls. The pathway dissection, while extensive, sometimes is approach in unconvincing ways and the results are not always evident to judge or interpret. Technically, the western blots and transcriptomic analyses require notable improvements.

      Comments on revised version.

      The revised version of the manuscript provides some improvements. However, I feel that many aspects remain poorly addressed. In the authors' favour, many of these limitations are now acknowledged in their rebuttal, as well as in the discussion section.

    3. Reviewer #2 (Public review):

      Summary:

      This work identifies a previously unknown way that red light can slow ageing. The authors show that red light lowers the level of a protein called SIRT4 in skin cells. Reducing SIRT4 boosts fatty acid use and increases a type of histone modification that keeps genes active. These changes help cells clear away signs of ageing, reduce inflammation, and restore normal metabolism. The findings open the possibility of developing new treatments that target SIRT4 to reverse age‑related decline.

      Strengths:

      The evidence is solid because the authors use several complementary methods. They test red light in both cultured cells and naturally aged mice, and they confirm the key role of SIRT4 by silencing its gene. Measurements of metabolism, protein changes, and ageing markers all point in the same direction. However, the exact way red light lowers SIRT4 levels is not fully explained, which leaves a minor gap. Overall, the conclusions are well supported and convincing.

      Weaknesses:

      The paper does not evolve to use the mechanistic discoveries of the manuscript to help our community to identify the mechanism of photobiomodulation, which is not known so far.

      I would like to draw your attention to a recently published paper by Herrera et al. (FEBS Letters 2025, doi:10.1002/1873-3468.70195), which shows that red light (660 nm) stimulates mitochondrial fatty acid oxidation in keratinocytes via AMPK‑dependent phosphorylation of ACC, without altering expression of electron transport chain complexes. I believe this paper is highly complementary to current study.

      Herrera et al. demonstrate that red light increases basal, ATP‑linked, and maximal oxygen consumption rates in keratinocytes specifically through enhanced fatty acid oxidation (inhibited by etomoxir). This independently validates the central finding of the current manuscript ,i.e., red light boosts lipid metabolism, strengthening the robustness of this concept.

      While the current manuscript focusses on the SIRT4‑MCD axis, Herrera et al. identify AMPK phosphorylation and ACC inhibition as key effectors. Authors can integrate and expand their discussion, since SIRT4 downregulation may converge on AMPK activation, or they may represent parallel, reinforcing mechanisms. This would enrich the mechanistic model and open new hypotheses.

      The mechanism of photobiomodulation: Herrera et al. explicitly challenge the prevailing paradigm that red light acts solely via cytochrome c oxidase (by showing long‑lasting effects, unchanged OXPHOS protein levels, and no difference in permeabilized cells). The current finding (red light acts through SIRT4 downregulation, i.e., not direct enzymatic activation, aligns perfectly with Herrera´s critique.

      Long‑term metabolic effects - Herrera et al. show that a single red light exposure elevates oxygen consumption for up to 2 days. The current study focuses on changes at 12‑24 h. Their data extend the time window and suggest that the metabolic reprogramming you describe may persist longer than currently discussed, which is clinically relevant.

      Discussing Herrera et al. results would not only acknowledge independent, corroborating evidence but also allow the authors to position your SIRT4‑centric mechanism within a broader, emerging understanding of red‑light photobiomodulation.

      Comments on the latest version:

      The authors have made a terrific work in answering the reviewers and modifying the manuscript.

    4. Author response:

      The following is the authors’ response to the original reviews.

      We sincerely thank the editors and reviewers for your careful evaluation of our manuscript and for the constructive recommendations that have helped us improve the rigor, clarity, and balance of the study. We are pleased that the reviewers recognized the potential value of linking red light exposure to SIRT4 downregulation, fatty acid metabolism, H3K9 acetylation, and attenuation of ageing-related phenotypes. We have revised the manuscript extensively in response to the reviewers’ comments.

      In particular, we have clarified the wavelength specificity of the red-light response, reanalyzed and more cautiously interpreted the omics data, improved the presentation and quantification of semi-quantitative experiments, revised statistical reporting, corrected gene/pathway annotations, toned down mechanistic claims where direct evidence was insufficient, and expanded the Discussion to integrate recent evidence on red-light-induced fatty acid oxidation and AMPK/ACC signaling. We also added a dedicated limitations paragraph addressing the use of female mice, the absence of a complete in vivo wavelength-control and source-blocked sham cohort, and the need for future direct metabolic flux and isolated mitochondria studies.

      Public Reviews:

      Reviewer #1 (Public review):

      Weaknesses:

      This is a challenging hypothesis that would require some additional experimental controls. The pathway dissection, while extensive, is sometimes approached in unconvincing ways, and the results are not always evident to judge or interpret. Technically, the western blots and transcriptomic analyses require notable improvements.

      We would like to thank the reviewer for the careful and patient examination of the issues identified in our manuscript. The poor quality of some of the Western blot bands in Figure 4 may have been caused by inappropriate electrophoresis conditions during the Western blot experiments. In the revised manuscript, we will optimize the electrophoresis conditions to obtain higher-quality protein bands and update the quantitative data. Regarding the quantification format, we believe that heatmaps provide a more intuitive representation of trends in protein expression across different treatment groups. This approach more accurately reflects the results of our biological replicates than simply analyzing the significance of differences in the grayscale values of protein bands. For the analysis of transcriptomic data, we will conduct a more detailed analysis of signal pathway enrichment and the identified differentially expressed genes to ensure that predicted genes are excluded from our current results and redundant data presentation is removed.

      Regarding additional experimental controls, such as incorporating experimental data under blue light treatment conditions as a control for red light. While exploring the optimal red light irradiation dose at the cellular level, we simultaneously conducted experiments on the effects of blue light irradiation at the same dose on keratinocyte activity. The results indicated that as the blue light irradiation dose increased (0–160 J/cm<sup>2</sup>), the keratinocyte activity exhibited a dose-dependent decline. This indicates that blue light is phototoxic to keratinocytes. The relevant experimental results have already been published in our previous study (Communications Biology 2024, doi: 10.1038/s42003-024-06973-1). Taken together with the data from our study, this demonstrates that the anti-ageing effects of red light reported in the current manuscript are indeed driven by red light.

      Reviewer #2 (Public review):

      Weaknesses:

      The paper does not evolve to use the mechanistic discoveries of the manuscript to help our community to identify the mechanism of photobiomodulation, which is not known so far.

      I would like to draw attention to a recently published paper by Herrera et al. (FEBS Letters 2025, doi:10.1002/1873-3468.70195), which shows that red light (660 nm) stimulates mitochondrial fatty acid oxidation in keratinocytes via AMPK‑dependent phosphorylation of ACC, without altering expression of electron transport chain complexes. I believe this paper is highly complementary to the current study.

      Herrera et al. demonstrate that red light increases basal, ATP-linked, and maximal oxygen consumption rates in keratinocytes specifically through enhanced fatty acid oxidation (inhibited by etomoxir). This independently validates the central finding of the current manuscript, i.e., red light boosts lipid metabolism, strengthening the robustness of this concept.

      While the current manuscript focuses on the SIRT4-MCD axis, Herrera et al. identify AMPK phosphorylation and ACC inhibition as key effectors. The authors can integrate and expand their discussion, since SIRT4 downregulation may converge on AMPK activation, or they may represent parallel, reinforcing mechanisms. This would enrich the mechanistic model and open new hypotheses.

      The mechanism of photobiomodulation: Herrera et al. explicitly challenge the prevailing paradigm that red light acts solely via cytochrome c oxidase (by showing long-lasting effects, unchanged OXPHOS protein levels, and no difference in permeabilised cells). The current finding (red light acts through SIRT4 downregulation, i.e., not direct enzymatic activation) aligns perfectly with Herrera´s critique.

      Long-term metabolic effects-Herrera et al. show that a single red light exposure elevates oxygen consumption for up to 2 days. The current study focuses on changes at 12-24 h. Their data extend the time window and suggest that the metabolic reprogramming you describe may persist longer than currently discussed, which is clinically relevant.

      Discussing Herrera et al.'s results would not only acknowledge independent, corroborating evidence but would also allow the authors to position their SIRT4-centric mechanism within a broader, emerging understanding of red-light photobiomodulation.

      We would like to thank the reviewer for providing us with constructive suggestions for discussion. Our results showed that under red light conditions, both glycolipid and lipid metabolism were activated in keratinocytes, and cellular metabolic flux increased. The activation of lipid metabolism directly led to an increase in metabolism-associated H3K9ac and drove the upregulation of anti-ageing-related genes; we believe this is key to the anti-ageing effects of red light. Mechanistic analysis combining proteomics and acetylation proteomics revealed that red light significantly downregulated SIRT4 expression and increased the acetylation of MCD, a protein regulated by SIRT4 that governs cellular fatty acid oxidation rates. Through validation using cell-level knockdown and inhibitors, we confirmed that SIRT4 inhibition exerts anti-ageing effects in vitro and that inhibiting MCD function under red light conditions suppresses H3K9ac. These results establish the role of the SIRT4-MCD signalling axis in mediating the anti-ageing effects of red light.

      The study by Herrera et al. included a substantial body of validation data confirming the role of red light in promoting fatty acid oxidation, providing robust empirical support for our research. Furthermore, Herrera et al. revealed that red light-induced fatty acid oxidation depends on AMPK and ACC phosphorylation. This mechanism of red-light photobiomodulation may refute the notion that its bio-regulatory effects rely solely on the action of mitochondrial cytochrome c oxidase. Furthermore, together with our study revealing that red light exerts anti-ageing photobiomodulatory effects via the SIRT4-MCD signalling axis, these findings independently confirm that red light regulates cellular fatty acid oxidation, thereby demonstrating the pivotal role of activated fatty acid oxidation in the bio-regulatory effects of red light. In the revised manuscript, we will include a discussion on the potential link between the red light-driven downregulation of SIRT4 and the phosphorylation of AMPK/ACC. This will be of positive value in elucidating how SIRT4 exerts its anti-ageing effects by regulating lipid metabolism, as well as in explaining the possible mechanisms by which red light downregulates SIRT4.

      Recommendations for the authors:

      Summary of Major Revisions

      Changes made in the revised manuscript:

      (1) Added a clearer explanation of why the 625-635 nm red-light regimen was considered the active intervention and how the available blue-light data from our previous work support wavelength-dependent effects on keratinocytes.

      (2) Revised the language describing inflammatory regulation. We now avoid presenting red light as producing a uniform anti-inflammatory effect and instead describe selective remodeling of ageing-associated inflammatory and SASP signatures.

      (3) Improved figure presentation and quantification for immunofluorescence, metabolite, and western blot assays; clarified image-analysis regions, replicate numbers, and normalization procedures.

      (4) Reanalyzed transcriptomic, proteomic, and acetyl-proteomic datasets with appropriate multiple-testing correction and corrected erroneous pathway/gene annotations in metabolic gene panels.

      (5) Replaced overly strong causal wording with more conservative language, especially regarding PI3K/Akt/mTOR, cytochrome c oxidase, SIRT4 localization, PPARα immunofluorescence, and direct fatty acid oxidation flux.

      (6) Expanded the Discussion to incorporate Herrera et al. (FEBS Letters 2025, doi:10.1002/1873-3468.70195), highlighting convergence between the SIRT4-MCD model and AMPK/ACC-dependent fatty acid oxidation.

      (7) Corrected typographical, nomenclature, and figure-legend inconsistencies throughout the manuscript.

      Reviewer #1 (Recommendations for the authors):

      (1) Wavelength specificity and need for a non-red-light control

      As a reader, one is left wondering whether the effects are due to red light specifically. An important control would have been to irradiate mice and cells with another light wavelength, such as blue light.

      We agree that wavelength specificity is a critical issue for interpreting photobiomodulation studies. In the revised manuscript, we have clarified that the anti-ageing and metabolic effects described here apply specifically to our 625-635 nm red-light regimen, rather than to visible light in general. We have also added a discussion of our previously published blue-light experiments, in which keratinocyte viability decreased in a dose-dependent manner across the same 0-160 J/cm<sup>2</sup> dose range (Communications Biology 2024, doi: 10.1038/s42003-024-06973-1). These data indicate that blue light and red light produce distinct biological outcomes in keratinocytes. Because high-dose blue light was cytotoxic under comparable cellular conditions and because the present study was designed to investigate the long-term effects of red light in aged mice, we did not perform prolonged in vivo blue-light irradiation as an ageing intervention.

      Changes made in the revised manuscript:

      Clarified in the revised Introduction and Discussion that the conclusions are specific to 625-635 nm red light under the irradiation parameters used in this study. (Lines 92 to 94, Lines 1146-1149)

      Added text summarizing the published blue-light comparison data from our previous study (Communications Biology 2024, doi: 10.1038/s42003-024-06973-1), including the dose-dependent decline in keratinocyte activity after blue-light irradiation. (Lines 90 to 92, Lines 1146-1149)

      Added data on the wavelength range of the red light used in this study. (Lines 529 to 531, Fig S1a)

      (2) Complexity of inflammatory effects

      The manuscript repeatedly emphasizes anti-inflammatory effects, yet some cytokines such as IL-18, Ccl2, TNF-α, Ccl2, and IL-8 appear increased. This suggests that the effects may be more complex than presented and may require additional readouts or stronger statistical power.

      We unanimously agree that the inflammatory response to red light should not be described as a simple, uniform suppression of all cytokines. We have demonstrated that changes in the levels of the senescence-associated secretory phenotype (SASP) at the cellular level and in skin tissue following red light treatment not only indicate that red light-induced metabolic activation can reduce the age-related inflammatory baseline, but also reveal red light-driven short-term reparative effects or stress-related cytokine responses. We consider this to be consistent with the findings, and the downregulation of NF-κB-related signalling observed in skin tissue following periodic red light irradiation of aged mice further supports the conclusion that red light alleviates the age-related inflammatory baseline. In fact, in our previous study, we did observe that red light treatment promoted increased levels of the cytokine Ccl2, which plays an important positive role in rapid wound healing (Communications Biology 2024, doi: 10.1038/s42003-024-06973-1). In the revised manuscript, we have reworded the relevant Results and Discussion sections to indicate that red light remodels ageing-associated inflammatory signalling rather than globally reducing every inflammatory mediator. We have also toned down statements suggesting that red light ‘reverses’ or ‘suppresses’ inflammation where the underlying data support a more selective effect.

      Changes made in the revised manuscript:

      Replaced broad terms such as “anti-inflammatory effects” with more precise wording such as “remodeling of ageing-associated inflammatory signaling” where appropriate. (Lines 541 to 543, Lines 573 to 574, Lines 893 to 895)

      Expanded the Discussion to explain that red-light-induced metabolic activation may simultaneously reduce senescence-associated inflammatory tone while allowing transient reparative or stress-related cytokine responses. (Lines 1135 to 1142)

      (3) Figure clarity, semi-quantitative methods, western blot quality, and inconsistent band patterns

      Many differences are difficult to see or require orthogonal validation. Some tissue-specific signals and western blots are difficult to judge. Several western blots are of poor quality, and multiple markers show inconsistent band profiles across experiments, including SIRT4 in Figure 5.

      We thank the reviewer for highlighting these technical and presentation issues. We have reviewed the semi-quantitative data and revised the presentation of the figures to improve their interpretability. For Western blot experiments, we optimised the electrophoresis and transfer conditions, replaced low-quality representative images where possible, and updated the semi-quantitative results. Furthermore, regarding the lack of clarity in the SIRT4 protein band, we have conducted repeat experiments and updated the main text to include a clearer image of the band. The issue with the annotation of the protein location was in fact due to an oversight during the data analysis process; we have carried out a detailed review and provided the uncropped full-length Western blot images for all experiments in the Supplementary Materials for the reviewers’ scrutiny. Finally, we would also like to point out that factors such as sample origin, protein extraction, electrophoresis conditions, antibody exposure time, and potential non-specific detection may all contribute to differences in band patterns. At present, the core conclusions regarding SIRT4 are supported by multiple lines of evidence, including mRNA analysis, immunofluorescence, Western blotting of bands at the expected sizes, and SIRT4 knockdown experiments, rather than being based solely on any single semi-quantitative Western blot result. We therefore believe that the conclusions drawn from the data presented in the revised manuscript are equally convincing.

      Changes made in the revised manuscript:

      Replaced or improved low-quality Western blot panels and updated quantitative analyses in revised Figures 3-5 and associated supplementary material. (Fig 3m, Fig 4, Fig 5f)

      Clarified the normalization approach for H3K9ac/H3 and target/loading-control comparisons, and the use of Actin or H3 as appropriate loading controls. (Lines 283 to 293)

      Bands with nonspecific profiles were excluded from quantitative conclusions and the manuscript conclusions no longer depend on those ambiguous signals.

      Revised the Results (Repeat the experiment to update the low-quality Bands) to avoid overstating changes that are not clearly visible or not supported by statistical analysis. (Fig 4n and r)

      (4) Choice of pharmacological agents and need for genetic strategies

      The choice of drugs in Figure 4 is puzzling. More specific and widely used inhibitors could be used to block PI3K/Akt or mTOR, and natural agonists such as insulin or EGF could be used. Genetic strategies should complement these observations.

      We agree that pharmacological perturbation experiments should be interpreted with caution. In this study, our criteria for selecting inhibitors were based on transcriptomic and proteomic analyses; we sought to determine how the most direct inhibition of red light-activated signalling pathways would affect H3K9ac levels. In the revised manuscript, we have clarified the rationale for the compounds used and have reduced the causal weight assigned to these inhibitor/agonist experiments. These data are now presented as supportive evidence that red light is associated with metabolism-related signalling changes, rather than as definitive proof that PI3K/Akt/mTOR is the primary upstream mechanism. We have also emphasised the genetic SIRT4 knockdown experiments as a more direct mechanistic test for the SIRT4-centred part of the model. We acknowledge that additional experiments using more selective inhibitors, physiological agonists such as insulin or EGF, and genetic perturbation of PI3K/Akt/mTOR components would be valuable for future studies.

      Changes made in the revised manuscript:

      Revised the text describing pharmacological experiments to distinguish supportive pathway modulation from direct causal evidence. (Lines 787 to 789)

      Added a limitation and future direction noting that genetic perturbation of PI3K/Akt/mTOR and physiological pathway activation with insulin or EGF would strengthen the model. (Lines 1190 to 1194)

      (5) Serum NADH measurement

      In Figure 1t, the authors measure serum NADH. NADH is poorly detectable in serum or plasma, and changes may reflect blood-cell lysis during collection rather than circulating NADH.

      We appreciate this technical concern. We have revised the manuscript so that serum NADH is no longer used as a central mechanistic readout. We now treat this measurement only as an exploratory indicator of systemic redox-related changes and explicitly acknowledge that serum or plasma NADH is vulnerable to artifacts from blood-cell disruption during sampling. The mechanistic interpretation has been shifted toward cellular and tissue measurements, including intracellular NADH/NADPH/GSH, ATP, acetyl-CoA, fatty acid uptake, and H3K9ac, which are more directly relevant to keratinocyte metabolic remodeling.

      Changes made in the revised manuscript:

      Removed serum NADH from the main causal argument linking red light to metabolic flux and H3K9ac.

      Placed greater emphasis on cell-based metabolite assays, tissue acetyl-CoA, and H3K9ac measurements as the main metabolic-epigenetic evidence. (Lines 582 to 587, Fig 1s)

      (6) Direct assessment of glycolysis and fatty acid oxidation

      The authors propose that red light increases glycolysis and fatty acid oxidation, but this could be assessed directly rather than through surrogate measures.

      We agree. Our current data include multiple metabolic readouts, including glucose and fatty acid uptake, ATP, NADH/NADPH/GSH, triglycerides, fatty acids, pyruvate, lactate, acetyl-CoA, and MCD-dependent changes; however, these assays are not equivalent to direct flux measurements such as Seahorse extracellular flux analysis, isotope tracing, or etomoxir-sensitive respiration. We have therefore revised the wording throughout the manuscript to distinguish metabolic remodeling and fatty-acid-oxidation-related signatures from direct measurements of fatty acid oxidation flux. We also incorporated the recent independent work by Herrera et al., which directly measured oxygen consumption and demonstrated red-light-induced fatty acid oxidation in keratinocytes. This external evidence supports the biological plausibility of our SIRT4-MCD model while making clear which aspects are directly measured in our study and which are inferred.

      Changes made in the revised manuscript:

      Replaced overstrong language such as “red light increases fatty acid oxidation” with “red light promotes PPAR-α-related fatty acid metabolism pathway” where direct flux data were not measured in our experiments. (Lines 882 to 883)

      Expanded the Discussion to integrate direct FAO evidence from Herrera et al. and to place the SIRT4-MCD axis within a broader red-light metabolic framework. (Lines 1169 to 1189)

      (7) Incorrect annotation of metabolic genes in Figure 3e

      Acss2, Aldh3b1 and Aldh3a1 are not glycolytic enzymes, Aldh3a3 does not appear to exist, and several enzymes classified as FAO are fatty acid synthesis enzymes. This questions the interpretation of the data.

      We thank the reviewer for identifying these annotation errors. We have rechecked the gene names and pathway assignments in the transcriptomic analysis and corrected the metabolic gene panels. We have confirmed that Acss2 is an acetyl-CoA synthase involved in the metabolism of acetate to acetyl-CoA. The Aldh family genes, meanwhile, are associated with aldehyde metabolism and detoxification. We have made the corresponding adjustments in the manuscript. The incorrectly listed Aldh3a3 entry has been removed. Furthermore, we have categorised genes involved in fatty acid metabolism as ‘fatty acid metabolism-related genes’, rather than grouping them all under the FAO category. These revisions have significantly improved the accuracy of the metabolic interpretation.

      Changes made in the revised manuscript:

      Reannotated Figure 3e and the corresponding Results text to correct glycolysis, TCA cycle, pentose phosphate pathway, and fatty acid metabolism categories. (Fig 3d)

      Removed the erroneous Acss2 and Aldh family genes. (Fig 3d)

      Revised the metabolic model to avoid using incorrectly grouped genes as evidence for direct fatty acid oxidation. (Lines 707 to 709)

      (8) Transcriptomic analysis and implausible volcano-plot p-values

      The transcriptomic analysis raises concerns. For example, the volcano plot in Figure 3d appears incorrect, with -log<sub>10</sub>(P-value) around 300 despite n=3 biological replicates.

      We thank the reviewer for pointing out this important issue. We have reopened the transcriptomic data and found that the extremely high -log<sub>10</sub>(P-value) in the original volcano plot were caused by the automatic replacement of very small P-values—generated during the differential expression analysis—with zero in the tabular data. To avoid misleading visualisations, we have regenerated the volcano plot using Q-values in place of the original P-values. Differentially expressed genes were defined as those with a Q-value < 0.05 and |log<sub>2</sub> fold change| > 1. Furthermore, for visualisation purposes only, the upper limit for q-values was set to 1 × 10<sup>-50</sup> for values below 1 × 10<sup>-50</sup>. This adjustment does not affect the statistical classification of differentially expressed genes but prevents over-interpretation of extremely small values. The revised volcano plots and legends have been updated accordingly. To avoid any potential misinterpretation arising from these updates, the updated volcano plots are presented in the supplementary materials.

      Changes made in the revised manuscript:

      Reanalyzed transcriptomic data using appropriate multiple-testing correction and revised the volcano plot. (Fig S3a)

      Corrected the y-axis transformation and removed implausible -log<sub>10</sub>(P-value) presentation. (Fig S3a)

      Updated Methods to specify the statistical workflow for transcriptomic differential expression and pathway enrichment. (Supplementary materials Lines 46 to 52)

      Moved the analysis of metabolic pathways based on transcriptomic data to the supplementary material, thereby reducing the reliance of the conclusions on transcriptomic data (Fig S3b).

      (9) Need to tone down mechanistic claims regarding PI3K/Akt/mTOR, cytochrome c oxidase, SIRT4, and PPARα

      The mechanisms proposed must be toned down. PI3K/Akt/mTOR should not be called glycolytic pathways, the link to red light or cytochrome c oxidase is vague, SIRT4 reduction requires mitochondrial counterstaining, and PPARα appears cytosolic after SIRT4 knockdown.

      We agree and have substantially revised the mechanistic language. PI3K/Akt/mTOR is no longer referred to as a ‘glycolytic pathway’; instead, it is described as a metabolism-related signalling axis that may influence glucose uptake, growth and nutrient-responsive metabolism. We have also toned down statements attributing red-light effects directly to cytochrome c oxidase, as our study primarily examines downstream metabolic and epigenetic remodelling rather than direct photoreceptor activation. With regard to SIRT4, we have revised the text to avoid interpreting changes in SIRT4 immunofluorescence alone as evidence of altered mitochondrial abundance or mitochondrial localisation. The conclusion is now based on a combination of SIRT4 mRNA levels, western blot bands of the expected size, immunofluorescence trends, and SIRT4 knockdown phenotypes. With regard to PPARα, we have re-examined the PPARα antibody used for the cellular immunofluorescence experiments. In the original Figure 5p, we mistakenly used a PPARα antibody (PPARα, Abclonal, A25296) that is only suitable for Western blot (WB) experiments; we believe this was the cause of the mislocalisation of the fluorescent signal; Consequently, we conducted new experiments using a PPARα antibody (PPARα, Abclonal, A22887) specifically designed for cellular immunofluorescence. The relevant experimental data have been corrected in the manuscript.

      Changes made in the revised manuscript:

      Replaced “PI3K/Akt/mTOR glycolytic pathway” with “The PI3K-AKT signalling pathway is involved in the regulation of glucose metabolism” throughout the revised manuscript. (Lines 701 to 705, Lines 809 to 810, Lines 813, Lines 1193)

      Reduced mechanistic certainty around cytochrome c oxidase and framed it as a possible upstream photoreceptor rather than an experimentally proven mechanism in this study. (Lines 827 to 830, Lines 842 to 848)

      Repeat the PPARα immunofluorescence staining experiment. (Fig 5p)

      (10) Need for isolated mitochondria experiments and red/blue light comparison of mitochondrial respiration

      If the effect of red light relies on mitochondrial cytochromes, additional proof would be needed, potentially using isolated mitochondria and comparing how red and blue light influence respiration capacity.

      We agree that isolated mitochondria experiments would be an important way to test direct mitochondrial photoreception. Because the present study was designed around cellular and in vivo metabolic-epigenetic remodeling, we did not perform isolated mitochondria irradiation experiments. To address this concern, we have toned down statements implying direct cytochrome activation and revised the Discussion to distinguish between direct mitochondrial photoreceptor models and downstream metabolic reprogramming. We also added a future direction proposing isolated mitochondria or permeabilized-cell experiments comparing red and blue light effects on respiration, ATP-linked OCR, maximal respiration, and FAO-dependent respiration. The revised manuscript now emphasizes that our data support a downstream SIRT4-MCD-H3K9ac mechanism after red-light exposure, while the proximal photophysical event remains to be fully defined.

      Changes made in the revised manuscript:

      Added discussion of the need for isolated mitochondria, permeabilized-cell, and wavelength-comparison respiration experiments. (Lines 1194 to 1197)

      Reviewer #2 (Recommendations for the authors):

      (1) Statistical reporting, post-hoc tests, normality/equal-variance tests, exact p-values, and FDR control

      The manuscript states that one-way ANOVA followed by Tukey or Dunnett tests was used, but it does not consistently specify the post-hoc correction for each figure. Normality and equal-variance tests are not reported, p-values are shown only as asterisks, and FDR control is not mentioned for transcriptomics and proteomics.

      We agree that the statistical reporting needed to be more complete. We have revised the Statistics and reproducibility section and the figure legends to specify the statistical test used for each experiment, the post-hoc correction applied after ANOVA, the number of independent biological replicates, and the definition of error bars. Where multiple comparisons were performed, we now state whether Tukey’s or Dunnett’s correction was used. Regarding P-value presentation, we have retained the use of asterisks in the figures as visual indicators of statistical significance to maintain figure readability. For transcriptomic, proteomic, and acetyl-proteomic analyses, we have revised the Methods section to state that multiple-testing correction was performed using the Benjamini–Hochberg false-discovery-rate procedure. Adjusted P values or Q values were used for differential-expression and pathway-enrichment analyses. These revisions clarify the statistical workflow and strengthen the reproducibility of the study.

      Changes made in the revised manuscript:

      Revised the Statistics and reproducibility section to define statistical tests, post-hoc corrections, assumption checks, and multiple-testing correction. (Lines 509 to 522)

      Updated relevant figure legends to include n values, statistical tests, post-hoc corrections, and definitions of significance symbols. (Lines 515 to 516)

      Added FDR control details for RNA-seq, proteomics, acetyl-proteomics, and pathway-enrichment analyses. (Lines 450 to 455)

      (2) Figure clarity and quantitative analysis of fluorescence, JC-1, metabolite, and western blot data

      Several figures lack clarity or appropriate quantification. Figure 1i-j H3K9ac quantification should be based on whole-image or multiple fields; Figure 2e JC-1 should include red/green ratio quantification; Figure 2k-p metabolite data should include absolute concentrations; Figure 3j needs appropriate loading controls.

      We appreciate these specific suggestions and have revised the figure presentation accordingly. For H3K9ac immunofluorescence in skin sections, we have clarified the anatomical region quantified and performed a more objective quantification using multiple fields/regions per section rather than relying on a visually selected dashed area. The dashed regions in the representative images were made clearer and the quantification criteria were added to the Methods and legend. For JC-1 staining, the bar chart on the right-hand side of the mitochondrial membrane potential fluorescence image in Figure 2e shows the quantitative data for the red/green fluorescence ratio obtained from independent experiments; compared with providing only a representative image, these data offer a more easily interpretable quantitative measure of mitochondrial membrane potential. To avoid any potential misunderstanding, we have corrected the vertical axis. For metabolite assays, we clarified normalization to cell number or protein content and revised the data presentation to include absolute or normalized concentrations where available, rather than relying solely on fold changes with variable y-axis scaling.

      Changes made in the revised manuscript:

      Revised Figure 1i-j quantification using multiple fields/regions per mouse section and improved dashed-region visibility. (Lines 431 to 440, Fig 1i and j)

      Corrected the vertical axis of the quantitative data for the JC-1 red/green fluorescence ratio. (Fig 2e and Fig S2a)

      Updated metabolite panels and/or source data to include absolute or protein-normalized values where available, and standardized y-axis interpretation. (Fig 2k-p and Fig 4s and v, Given the diversity of intracellular fatty acid and triglyceride species, absolute quantification based solely on absorbance measurements would be technically challenging and may not accurately reflect the content of each molecular component. Therefore, we presented the changes in fatty acid and glycerol levels as percentage-normalized relative absorbance values, which allowed consistent comparison among the experimental groups.)

      (3) Experimental design limitations: sex of mice and sham control

      Only female C57BL/6 mice were used, although aging and metabolic responses can be sex-dependent. The thermal-control argument lacks a true sham control in which mice are placed in the same apparatus with the light blocked at the source.

      We agree with these points. We have added a section on limitations stating that all aged mice used in this study were female, and that sex-dependent responses to red light, SIRT4 regulation, metabolism and skin ageing should be investigated in future studies using both male and female cohorts. Furthermore, regarding the design of the non-irradiated control group: although the control mice underwent the same depilation and routine procedures, they did not receive red light irradiation. However, we also acknowledge that establishing a sham-irradiated control group with light shielding would allow for stricter control of factors such as restraint, contact with equipment and procedural stress. However, given that this experiment involved a continuous cyclic photoperiodic treatment lasting two years, we were unable to supplement the study with a control experiment involving only red light shielding. Nevertheless, based on the fact that we observed only minimal changes in the mice’s skin temperature following red light irradiation, we believe that the primary factor driving the alleviation of the skin ageing phenotype in the mice remains red light-induced.

      Changes made in the revised manuscript:

      Added a limitation noting that the study used female C57BL/6 mice only and that sex as a biological variable should be addressed in future studies. (Lines 1198 to 1201)

      (4) Textual errors, nomenclature inconsistencies, and ChIP-qPCR normalization

      Several textual errors and inconsistencies should be corrected, including Pparg1a/Ppargc1a, Ricotr/Rictor, Pi3k/PI3K, Sirt4/SIRT4 protein nomenclature, and the use of RPL30 normalization in ChIP-qPCR without showing that RPL30 is unchanged.

      We thank the reviewer for their careful reading. We have corrected the typographical errors and standardised gene and protein nomenclature throughout the manuscript and figure legends. Specifically, Ppargc1α has been corrected to Ppargc1a, Ricotr to Rictor, and the capitalisation of PI3K has been standardised. We now use Sirt4 for the mouse gene and SIRT4 for the protein, applying the same convention to other genes and proteins. For ChIP-qPCR, we have revised the Methods and Results sections to describe normalisation against input and IgG controls more clearly, and to specify the role of the RPL30 locus as an internal control. We have also included data in the Supplementary Materials showing relative enrichment of H3K9ac in the RPL30 promoter region in PAM212 cells before and after red light irradiation; the results indicate that H3K9ac enrichment at the RPL30 locus remained stable across treatment groups after normalization to input DNA and correction against IgG background. This result indicates that the use of RPL30 as an internal control in ChIP-qPCR experiments is feasible.

      Changes made in the revised manuscript:

      Corrected Ppargc1a, Rictor, PI3K, Sirt4/SIRT4, and related nomenclature throughout the manuscript.

      Supplement the experimental results on the effect of red-light irradiation on the level of H3K9ac enrichment at the RPL30 locus in keratinocytes. (Lines 339-352, Lines 539 to 541, Fig S1d)

      (5) Additional Revision Addressing the Public Review and Herrera et al.

      The reviewer suggested integrating the recent study by Herrera et al. showing that 660 nm red light stimulates mitochondrial fatty acid oxidation in keratinocytes through AMPK-dependent phosphorylation of ACC, without changing electron transport chain complex expression. The reviewer also noted that these findings may complement the SIRT4-MCD axis and challenge a cytochrome-c-oxidase-only model of photobiomodulation.

      We are grateful for this constructive suggestion. We have expanded the Discussion to incorporate Herrera et al. and to place our SIRT4-MCD-centered mechanism within the broader emerging model of red-light-driven metabolic remodeling. Herrera et al. provide direct oxygen-consumption evidence that red light enhances fatty acid oxidation in keratinocytes and that this effect involves AMPK/ACC signaling. This is highly complementary to our data, in which red light decreases SIRT4, increases acetylation of MCD, promotes fatty-acid-metabolism-related signatures, elevates acetyl-CoA, and increases H3K9ac. In the revised Discussion, we propose two nonexclusive models: red light-induced SIRT4 downregulation may converge with AMPK/ACC-dependent relief of fatty acid oxidation, or the two pathways may represent parallel reinforcing mechanisms that together enhance lipid metabolic flux.

      Changes made in the revised manuscript:

      Added a paragraph discussing Herrera et al. in the revised Discussion. (Lines A1169 to 1189, Lines 1206 and 1208)

      Revised the conceptual model of red-light photobiomodulation to emphasize downstream metabolic reprogramming rather than direct cytochrome c oxidase activation alone. (Lines 827 to 830, Lines 843 to 844)

      Added future directions to test whether red-light-induced SIRT4 downregulation causally affects AMPK/ACC phosphorylation and FAO-dependent respiration. (Lines 1169 to 1189)

      We again thank the editors and reviewers for their thoughtful and constructive comments. The revised manuscript now provides a more rigorous and balanced presentation of the evidence, distinguishes direct measurements from inferred metabolic flux, corrects pathway annotations, improves figure quantification and statistical transparency, and places the SIRT4-MCD-H3K9ac mechanism within a broader framework of red-light-induced fatty acid metabolic remodeling. We believe these revisions substantially strengthen the manuscript and clarify both the significance and the limitations of our findings.

    1. eLife Assessment

      This study presents valuable findings on the role of specific dopamine neurons for aversive learning and modulation of innate behavior in Drosophila larvae. The authors present convincing evidence backed up by detailed behavioral quantification and rigorous testing. Their data confirms previous findings and will be of interest to the learning and memory community.

    2. Reviewer #1 (Public review):

      Summary:

      The authors investigate the role of different specific dopaminergic neurons in the mushroom body of Drosophila larvae for learning and innate behavior. All the tested neurons are thought to be involved in punishment learning. The authors discover that artificial activation of single DANs in training leads to safety learning, but not punishment learning. Furthermore, activation of single DANs can lead to changes in locomotion behavior, which can affect light preference. The authors provide a deeper understanding of the functional diversity of single dopamine neurons; however, it is unclear how translatable these findings are to learning experiments with real punishment stimuli.

      The authors provide a detailed behavioral analysis of locomotion in response to activation of various dopamine neurons. This analysis allows them to exclude that the locomotion defects affect memory recall behavior.

      Strengths:

      The authors disentangle which kind of memories are formed with the activation of different dopamine neurons - safety learning and/or punishment learning. They further investigate whether the US is required in the test for recall. They do indeed find differences, and the results will be of interest to the learning and memory community.

      Interestingly, optogenetic activation of a single DAN during training leads to safety memory, but not punishment memory. Furthermore, DAN activation also affects innate locomotion, and the authors show that optogenetic activation of different DANs affects locomotion differently.

      Weaknesses:

      All experiments in the manuscript use optogenetic activation of DANs, thus it is not clear what kind of memories are formed. Several stimuli can be used as punishment, such as electric shock, salt, bitter, and light - it is not clear what kind of memory the authors investigate here. The findings could be discussed in the context of what DANs respond to. Furthermore, studies in adults and larvae showed that most DANs can code for both valences - etc., aversive DANs can be activated by punishment, and inhibited by reward. Thus, safety learning might be a result of a decrease in activity in DANs during odor presentation. The authors also do not discuss possible feedback loops from MBONs to DANs across compartments. Could such connections allow for safety learning in larvae?

      The authors show that artificial activation with different light intensities can form different memories and that increasing the light intensity sometimes leads to no memories. Also, using different optogenetic tools reveals different results. This again raises the question of how applicable the results will be for learning with real stimuli. Is there a natural stimulus that only induces safety learning, but no punishment learning? The authors discuss these limitations.

    3. Reviewer #2 (Public review):

      Summary:

      This study provides valuable context for ongoing research on the role of dopamine in memory and locomotion. DANs have been a fascinating area of study due to their complexity, and this work dissects specific DANs, exploring their roles in different memory-related behaviors while offering some explanations. The discussions provided by the authors effectively situates the study in the broader field of learning, memory, DAN circuitry and behavioral computation in insect brains. The study achieves what it sets out to and it does so unequivocally. The experiments were elegantly designed, leaving little room for doubt in the study's claims. However, the study lacks context regarding the molecular pathways underlying these results. While it strengthens current knowledge by providing robust evidence, it does little to explore the molecular mechanisms behind these effects.

      Strengths:

      (1) Experiment design is one of the strengths of this study. The experiments are thorough and cover the length and breadth of the core findings of the study. Although a lot of work has already been done in studying the role of dopamine in memory and locomotion, the dissection of the functions of distinct DANs in larvae has been done meticulously with well-structured experiments.

      (2) This study fits quite nicely into the puzzle of memory, especially in the context of Dopamine. Previous studies in *Drosophila* adults have shown the opposing roles of DANs in locomotion depending on the context of DAN activation. This study drives that point home for larvae, providing conclusive evidence in that regard.

      (3) The use of clear figures and simple language is one of the strengths of this paper. The figures are comprehensive, complete and manage to narrate the story by themselves. The flow of information is smooth. The simple and effective language used maintains scientific rigor while remaining accessible to those new to the field. A pleasant read.

      Weaknesses:

      (1) The authors have done a great job at structuring the figures. But some main figures would benefit from including the controls instead of placing them in supplementary.

      (2) The paper would benefit from a deeper discussion regarding molecular mechanisms underlying their results. It would be interesting to see what the authors think about different Dopamine receptors and how they relate to the findings of this paper.

      (3) Throughout the paper, the authors have been clear and comprehensive, but in some cases, further explanation of their choices were missing. For example, the choice to compare bending and tail velocity over other parameters within the same clusters is unclear.

      Comments on revised version.

      Most of the comments have been addressed.

    4. Reviewer #3 (Public review):

      Summary

      Across species, dopamine release serves seemingly diverse functions, such as reinforcing memories and regulating locomotion and flight. However, whether distinct dopaminergic neurons (DANs) are allocated to each function is unclear. In this study, Toshima et al. have used the numerically simple organization of the Drosophila larval brain to answer this question. They use optogenetic activation to systematically stimulate a small set of DANs, individually and collectively, and study the effect on diverse functions such as memory formation, retrieval, and locomotion. The reproducibility of optogenetic activation is a strength of this approach. At the same time, this is a caveat, as optogenetic activation may not recapitulate natural modes of activation and may lead to outcomes not observed under natural conditions. They find that singly or collectively, DL1 DANs can induce punishment and/or safety memory formation and retrieval. DANs can even gate the expression of memory. Finally, the same DANs also modulate locomotion in the larvae. The authors speculate that dopaminergic neurons in other species may also share such overlapping functions. Their findings are nicely summarised in Figure 9.

      Strengths

      The study systematically activates neurons in the DL1 cluster. Individual and collective stimulation of the Dl1 DANs has been conducted to assess the induction and gating of aversive punishment memory, safety memory, and acute locomotion.

      Specific adult Drosophila DANs are known to induce dual behaviors and functions. The same MP1/y1pedc DANs are recognized for gating appetitive memory expression and representing aversive teaching signals downstream of sensory stimuli such as electric shocks, bitter tastes, and heat. Neurons in the PPL1 cluster regulate adult flight and food-seeking behavior. The authors deserve credit for conducting an organized examination of dopaminergic neuronal functions in larvae, thereby making their findings more comparable and facilitating the proposal of a holistic model.

      They have provided substantial evidence for their findings and have frequently presented replicated behavioral datasets. They have been transparent about the results that were difficult to explain. Additionally, they have provided an impressive body of supporting data to strengthen their main findings.

      Weaknesses

      As mentioned above, optogenetic activation may not recreate natural neuronal activation in response to external stimuli. This could have led to outcomes that will not occur under other natural circumstances.

      Comments on revised version.

      I appreciate the author's responses, and I do not have any comments or suggestions at this point.

    5. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Weaknesses:

      All experiments in the manuscript use optogenetic activation of DANs, thus it is not clear what kind of memories are formed. Several stimuli can be used as punishment, such as electric shock, salt, bitter, and light - it is not clear what kind of memory the authors investigate here. The findings could be discussed in the context of what DANs respond to.

      This is indeed a caveat of our study and we discuss this issue now in lines 557-566. We refrained from testing necessity to specific US on purpose as we knew that another research group was focussing on this question in parallel (Weber et al., 2023, also published in eLife) and therefore rather focussed on complementary experiments. That study also includes a rather deep discussion about the inputs to individual DANs. We briefly refer to this discussion (lines 442-444) but decided to not go into detail to avoid too much overlap.

      Furthermore, studies in adults and larvae showed that most DANs can code for both valences - etc., aversive DANs can be activated by punishment, and inhibited by reward. Thus, safety learning might be a result of a decrease in activity in DANs during odor presentation. The authors also do not discuss possible feedback loops from MBONs to DANs across compartments. Could such connections allow for safety learning in larvae?

      We thank the reviewer for raising these points and included a brief discussion of both scenarios in lines 469-474.

      The authors show that artificial activation with different light intensities can form different memories and that increasing the light intensity sometimes leads to no memories. Also, using different optogenetic tools reveals different results. This again raises the question of how applicable the results will be for learning with real stimuli. Is there a natural stimulus that only induces safety learning, but no punishment learning?

      We do not know of such a stimulus. Based on our data, a US that only activates a single DAN should only make safety memory – however, the available data of which US activates which DAN is very limited in larvae and currently no such US is known. We discuss this point briefly in lines 557-563 and 572-574.

      The authors provide a detailed behavioral analysis of locomotion behavior; however, the detailed analysis seems unnecessary for that dataset. Modulation of speed and bending rate has been described before with simpler methods (specifically for MBONs). The revealed locomotion phenotypes probably affect larval locomotion during memory recall with light activation, thus the authors should show that larvae are potentially able to move during light-on memory tests.

      We expanded our locomotion analysis of the innate and learned odor preference experiments (new Fig. 6) and show that in these experiments, even with TH-DANs being activated, larvae indeed can move relatively normal and the existing locomotion phenotypes are not correlated to their olfactory choices.

      We do not agree that the locomotion analysis is unnecessary. Modulations of speed and bending have been described for MBONs but to our knowledge not for DANs. It is not trivial at all that DANs and MBONs cause the same behavioral modulations (see, for example, this adult study: Mohammad et al., 2024 Plos Biol). There is extremely limited knowledge about the motoric effects of dopaminergic neurons in larvae - we therefore find it important to describe our results in detail. We added some further rationale of why we think it is crucial to explore the functions of DANs for learning and movements together (lines 81-86).

      Reviewer #2 (Public review):

      Weaknesses:

      (1) The authors have done a great job at structuring the figures. But some main figures would benefit from including the controls instead of placing them in supplementary.

      We had decided to put the controls into the supplement in some cases to prevent the main figures to be overcrowded. We revised this decision upon the reviewer’s comment for Fig. 8 (previously Fig. 7) but decided to keep other figures unchanged as we feel that the current design best fits the purpose of each figure. We provide a figure-for-figure rationale in our response to the recommendations for the authors.

      (2) The paper would benefit from a deeper discussion regarding molecular mechanisms underlying their results. It would be interesting to see what the authors think about different Dopamine receptors and how they relate to the findings of this paper.

      We thank the reviewer for the suggestion. Although we agree that such a discussion would be interesting, we hesitate to expand on this topic, as the discussion is already quite long and our study does not contribute any new data to clarify the molecular dopaminergic mechanism.

      (3) Throughout the paper, the authors have been clear and comprehensive, but in some cases, further explanation of their choices were missing. For example, the choice to compare bending and tail velocity over other parameters within the same clusters is unclear.

      We understand that this choice was not clearly explained and expanded on our rationale in lines 244-251.

      Reviewer #3 (Public review):

      Weaknesses:

      The larvae exhibit directed locomotory action to express punishment or safety memory. If the larvae did not move, we would not be able to assess memory function. Hence, functional activation of DANs could result in one action, which seems like two different functions of memory expression and locomotion. It can also be argued that activation of DANs represents a teaching signal to the KCs, and then eventually, downstream of the MBONs, it results in locomotion modulation. Hence, the seeming functional diversity could be a function of different downstream neuronal pathways and not molecular context-dependent diversity inside dopaminergic neurons. The authors should address this possibility or point out the fallacy in the above argument.

      We thank the reviewer for raising this issue. To the first point, we expanded our locomotion analysis of the innate and learned odour preference experiments (new Fig. 6) and show that in these experiments, even with TH-DANs being activated, larvae indeed can move relatively normal. In addition, the existing locomotion phenotypes in these experiments were not correlated with the animals’ olfactory choice. This makes it unlikely that the changed locomotion directly determines our observation during the olfactory experiments.

      We do agree that it is possible that both the preference after learning and the changed locomotion could come through the same dopaminergic mechanism via diverse downstream pathways. We cover this hypothesis in Fig. 10G and address this question briefly in lines 580-582.

      The finding that activation of TH-GAL4 conveys aversive valence and R58E02-GAL4 conveys appetitive valence seems redundant (Figure 6). I understand they say this in the context of locomotion. However, they may not have mentioned similar findings in adults. In adults, artificial activation of DANs covered by the same GAL4 lines acts as aversive and appetitive teaching signals for memory formation. These references should be cited appropriately in the results and discussion if not currently included.

      We thank the reviewer for this comment and tried to include the relevant adult literature (see e.g. lines 351-356 and 583-604). In particular, we added a quite detailed discussion about a paper published after our initial submission that performed similar experiments for the adult PAM-DANs Lozada-Perdomo et al., 2025, iScience).

      We do not agree, however, that the experiments in Fig. 7 (previously Fig. 6) are redundant. Recent studies in adults found no correlation between the rewarding/punishing effects and the innate valence a given dopaminergic neuron induces (Rohrsen et al., 2021, bioRxiv; Mohammad et al., 2024, PLOS Biol; Lozada-Perdomo et al., 2025, iScience). To our knowledge, no such studies have been carried out in larvae so far. Therefore, we think that it is not only important to test it but that the respective results compared to the results in adults are of relevance for the readership.

      The evidence for the role of dopamine (Figure 7) can be bolstered by using other available RNAi lines against TH. A valium20 vector-based shRNA line is recommended. The current evidence is based mainly on non-specific pharmacological intervention with 3IY.

      We agree to this caveat and made it transparent now in lines 387-389 and 401-403. We nevertheless chose, for the time being, to not include further experiments to address this point in the current study.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Activation of specific or multiple DANs seems to increase naïve odor preference (f1 or TH). Is this due to locomotion defects in TH - how does the odor preference develop over time? Can this increased odor preference explain the safety learning - where they also approach the odor stimulus?

      The reviewer is right that in presence of light, we see increased odor preference both innate and after unpaired training – theoretically, that could be the same effect. However, this would not explain why we see the same increased preferences after unpaired learning with all driver strains but increased innate preferences only for some of them. Moreover, when activating TH, we also see increased odor preference in absence of light after unpaired training but not innately. We therefore think that these are independent effects.

      We also include a new Fig. 6 providing additional information, including the development of preference over time, and addressing the question whether the modulations in locomotion can explain differences in odor preference.

      Locomotion behavior was assessed in 30s light-on periods - was the behavior different from the memory test or naïve preference test which had light on for 3 (or 2.5) minutes - which light intensity was used for the data shown in Figure 5.2? How did the larvae move in the high light concentration/ or with ATR - did they not show memory due to impaired locomotion?

      Light intensity for all odour preference and learning experiments with ChR2-XXL was 100 µW/cm<sup>2</sup> (except Fig. 2 – S2C), i.e. equivalent to the experiments in Fig. 5 – S1 and Fig. 9 – S2 (weak light). We replaced Fig. 5 – S2 with a new expanded Fig. 6 analysing the locomotion of our experiments shown in Fig. 2F, 3F and Fig. 2 – S1F. We show in this figure that the larvae can move relatively normal and that the locomotion effects are weaker than in our experiments with 30s light periods. Unfortunately, we do not have videos available for all experiments and therefore cannot make a similar analysis for Fig. 2 – S2C and D when we used strong light or ATR feeding. The experimenters did not notice impaired locomotion during the experiment and the animals did show normal odour preferences similar to those shown in Fig. 3 – but the preferences were the same after paired and unpaired training, resulting in zero Memory Scores. Therefore, we do not think that the locomotion prevented the memory expression.

      In several experiments, even genetic controls seem to show learning with blue light activation. Thus, the light itself seems to activate DANs. The authors should discuss these effects and explain what this could mean for the findings. The light stimulus might not just activate the specific DAN that expresses the optogenetics, but also additionally other DANs which respond to light.

      We thank the reviewer to point this out and point out this caveat in lines 159-165.

      The authors speculate about the function of potential MB circuits - the DAN-MBON circuit is not well described so far and might be required for the US in test memory recall. A straightforward experiment to investigate the involvement of this circuit in punishment or safety memory recall would be to block dopamine receptors in the MBON.

      We very much agree to this suggestion, but believe these experiments are beyond the scope of the current study. We therefore decided to not perform these experiments for the current paper.

      Reviewer #2 (Recommendations for the authors):

      (1) As self-explanatory as the figures are, it would be interesting to also see controls in some of them. For example, in Figure 5, the effect size graph (Figure 5C) clarifies to an extent the difference between control genotypes and the experimental genotypes. It would be nice to see the results of genetic controls in Figures 5A and 5B instead of in Figure 4 - supplement 2.

      We originally decided to put the controls into the supplement to prevent the main figures to be overcrowded. We revised this decision upon the reviewer’s comment for Fig. 8 (originally 7). For Fig. 4 and 5 specifically, we decided to keep the current layout because each serves a different purpose: Fig. 4D-L, Fig. 5 – S1 and S2 present the actual data with all genotypes that were made in parallel and therefore can be compared directly. Fig. 4 – S2 aims to visualize the effect of the light by comparing all controls across all experiments, normalized to the same starting value. Fig. 5A and B aim to compare the shape and effect size of activating DANs on top of the effect of the light - therefore, we subtracted the controls in each experiment from the experimental group. We think that adding the controls’ behaviour to Fig. 5A and B would undermine the aim of this figure.

      (2) It is a bit unclear why bending and tail velocities were the parameters chosen to compare between groups while in most cases they were of lower relative importance according to Figure 4 - supplement 1. Elaborating on this would strengthen the differences in behavior and also the claims of this study.

      We thank the reviewer for the suggestion and tried to make our choice clearer. Please see our answer to the respective part of the public review.

      (3) In adults, it has been shown that the same DAN can encode opposing valence depending on whether it was activated before or after odor presentation. Discussing the importance of temporal order of stimulus processing would bolster the results regarding paired and unpaired training in Figure 3.

      We thank the reviewer for this very good suggestion – also in larvae, this temporal function has been described. We discuss these observations in relation to our results in lines in 481-494.

      Reviewer #3 (Recommendations for the authors):

      Toshima et al., as stated in the public reviews, have done an admirable job with this manuscript. Below are specific suggestions that could improve the manuscript. It is, of course, up to the authors to decide which ones to attend to.

      Treat controls consistently. In Figure 2 and others, parental controls are not pooled, but in Figure 3, for odor preference, controls are pooled.

      We agree that the same things should be treated in the same way throughout a study and normally adhere to this principle. We nevertheless made an exception for Fig. 3 only because its goal is to provide a post-hoc analysis across several replications of experiments, some of which included genetic controls, others not (from Fig. 2, Fig. 2-S2 and S3). Due to relatively small effect sizes and high variability in odor preferences, to answer the question of paired and unpaired learning, we need higher sample sizes than each individual experiment provided. We therefore decided to pool all “equivalent” data across all these experiments. We do agree that this is a suboptimal approach but hope the reviewer can understand the rationale behind it. We explained our rationale clearer now (lines 180-183).

      I prefer to see all data points in a graph. It is more transparent than the box plots. Also, could you note why the data median is preferable to show over the mean?

      Although we in principle agree to the notion that presenting all data points is more transparent, we opted against it as it makes some graphs harder to read in particular with high sample sizes – in some of our figures, we have hundreds of data points per group. We explain our choice, including for using the median, in the method section (lines 835-839).

      Please undertake another round of language editing to handle spelling errors, etc. Use consistent British/American English.

      We thank the reviewer for their suggestion and tried our best to fix any spelling and grammar errors.

      I urge the authors to move beyond the false dichotomy of 'p' value statistics to using the statistical framework of estimation statistics for data analysis. I understand switching from familiar statistical analysis in such a late manuscript stage is very difficult. However, the authors can consider the estimation statistics framework in subsequent studies. https://www.estimationstats.com is a good starting point for biologists to get to know a framework that has been extensively worked on and is arguably a more 'honest' way of analyzing data. Disclaimer: I am not associated with the above website.

      We agree that the p-value has problems and are aware of the estimation statistics framework. We had considered applying it here, but we decided against switching to a completely different statistical framework for a research project that was ongoing since several years. However, we are sincerely considering it for our current research projects.

    1. eLife Assessment

      In this solid work, Fukui et al. re-examined the ATP hydrolysis mechanism in GHKL ATPases, revealing a cooperative role for two conserved acidic residues rather than a single one. This valuable study used a range of biochemical and structural techniques on various mutants from different members of the GHKL ATPase family to test and validate their proposed mechanism. An updated and extended mechanistic model of ATP hydrolysis by this class of enzymes is proposed.

    2. Reviewer #1 (Public review):

      Summary:

      In this manuscript the applicants study two residues in the GHKL ATPase active site of Aq MutL and GyrB, and argue that the catalytic base function is shared between two conserved acidic residues that are 3 residues apart.

      In the manuscript, they generated mutant versions in MutL and GyrB (both ala and the appropriate Asn/Gln version) and performed ATPase analysis. They also generated high resolution crystal structures of the GyrB NTD with AMPPnP for WT and mutants of the two acidic residues. The data show that mutation in either of these residues does not fully kill activity (with the exception of the Alanine mutation of the first of the two, that interferes with ATP (or AMPPnP) binding). When the acidic residues are mutated to Asn/Gln, the catalytic water can still be positioned, and hence these mutants are more active than the Ala mutants. In both cases the double mutation is catalytic dead.<br /> The authors then perform phylogenetic analysis and ancestral gene reconstruction and based on this they argue that HSP90 forms a different class of GHKL ATPases, and lost rather than gained this separate status.

      Strengths:

      The biochemical analysis seems solid.

      Weaknesses:

      - A major question that remains, is why the mutations have so much more detrimental effect in MutL (100-fold lower kcat/KM) than they do in GyrB (3-fold lower). Can the authors explain this? Doesn't this argue against the proposed catalytic conservation?

      The authors need to discuss this issue explicitly to make it clear that conservation of the mechanism is not complete and that other interpretations are possible.

      - The structure figures all have omit maps for just the AMPPnP and the water, whereas the density for the the acidic residues and their mutants are not shown.

      This has been addressed.

      There are some issues with figure S2B and S5.

    3. Reviewer #2 (Public review):

      Summary:

      In this manuscript, Fukui et al. re-examined the ATP hydrolysis mechanism in GHKL ATPases, revealing a cooperative role of two conserved acidic residues rather than one. The authors have used a range of biochemical and structural techniques on various mutants from different members of the GHKL ATPase family to test and validate their proposed mechanism.

      Through a detailed re-analysis of their previously published structure of the aqMutL NTD (ATPase domain) in complex with AMPPCP, they identified Glu29 and Glu32 as interacting with nucleophilic water for the catalysis. The authors carefully dissected the respective roles of these two acidic residues with a series of site-directed mutations. Mutations at Glu29 impaired ATPase activity without affecting protein secondary structure or ATP binding in the case of the E29Q mutant. Moreover, mutations at Glu32 did not affect secondary structure (except for E32G) but reduce ATPase activity. Activity was abolished when both residues (E29Q/E32Q) are mutated.

      The authors extended their study to another GHKL ATPase, aqGyrB. Their findings further supported the cooperative function of the corresponding acidic residues in aqGyrB (Glu48 and Asp51) during ATP hydrolysis. Mutation of these residues partially impaired ATP hydrolysis without affecting protein secondary structure. ATPase activity was completely lost in the double mutant E48Q/D51M. While the E48Q mutant retained the ability to bind ATP, the E48A mutant did not. High-resolution structures of the WT and E48A, E48Q, D51A and D51N mutants of the aqGyrB NTD demonstrated that nucleophilic water positioning depended on these residues. E48 played a dominant role in water positioning and is critical for stabilising ATP lid formation and associated conformational changes, whereas D51 contributed cooperatively to catalysis.

      The authors investigated the functional impact of mutating the corresponding residues in the human MutL homologs PMS2 and MLH1. Clinical variants consistently exhibited reduced or abolished ATPase activity, providing a potential molecular basis for Lynch syndrome, through impaired DNA mismatch repair.

      Lastly, through evolutionary analysis, the authors inferred that the second acidic residue was likely present in the common ancestor of MutL, GyrB, and MORC proteins, but was lost in the case of Hsp90.

      Strengths:

      (1) This study contains a detailed structural and biochemical analysis of a biologically important set of GHKL ATPases. The authors identify a second acidic residue that is conserved and contributes to catalysis in a large subset of GHKL ATPases. An updated and extended mechanistic model of ATP hydrolysis by this class of enzymes is proposed, which involves cooperative and partially overlapping roles for the catalytic residue pair. This revised mechanistic model is invaluable for the interpretation of clinical variants of GHKL ATPases such as PMS2 and MLH1.

      (2) The work described was performed to an excellent and rigorous technical standard. The structural and biochemical data are sound. The evidence supporting the claims is compelling.

      Weaknesses:

      (1) The identification in this study of a second acidic residue contributing to catalysis but not absolutely essential for catalysis is a useful finding. However, given that many structures of GHLK ATPases have been determined with different nucleotide analogs bound and that the essential role of the first acidic residue is well established, the importance and scope of the advances described here remain focused within the field of study of GHKL ATPases.

      (2) The authors assessed the consequences of variants in the human MutL homologs PMS2 and MLH1, but various other human GHKL ATPases contain clinically relevant variants, some of which have stronger disease associations than the mutations examined in this study. A broader analysis of any effect of disease-linked mutations in GHKL ATPases would have strengthened this study.

      (3) The effect of other aqMutL NTD E32 mutants, particularly, the E32K mutant on ATP binding remains unclear, although experimental assessment of nucleotide binding would be challenging due to the high protein concentrations required for the equilibrium dialysis assay.

    4. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      (1) A major question that remains is why the mutations have so much more detrimental effect in MutL (100-fold lower k<sub>cat</sub>/K<sub>M</sub>) than they do in GyrB (3-fold lower). Can the authors explain this? Doesn't this argue against the proposed catalytic conservation?

      We agree that the quantitative effects of the mutations differ between MutL and GyrB. However, we do not think that this difference argues against conservation of the catalytic mechanism. The trends of the mutational effects are highly consistent between the two enzymes. In both proteins, replacement of the conserved catalytic glutamate with Ala (E29A in aqMutL and E48A in aqGyrB) abolished ATPase activity and ATP binding, whereas replacement with the isosteric amide residue (E29Q and E48Q), which preserves hydrogen-bonding capability but lacks proton-accepting capacity, retained ATP binding and measurable ATPase activity. Likewise, substitutions of the second acidic residue (E32Q in aqMutL and D51N in aqGyrB) also retained substantial ATPase activity despite the loss of proton-accepting capability. Most importantly, simultaneous substitution of both acidic residues (E29Q/E32Q in aqMutL and E48Q/D51N in aqGyrB) completely abolished ATPase activity in both enzymes. Therefore, although the magnitude of the activity reduction caused by the individual mutations differs between MutL and GyrB, the qualitative pattern is essentially identical. We therefore think that the proposed catalytic mechanism is conserved, while the quantitative differences likely reflect differences in the local catalytic environment rather than differences in the underlying mechanism.

      (2) The structure figures all have omit maps for just the AMPPnP and the water, whereas the density for the acidic residues and their mutants is not shown.

      We have added Supplementary Fig. S2, which shows the 2F<sub>o</sub>–F<sub>c</sub> electron density maps around residues Glu48/Asp51 (or their substituted residues) in the wildtype and all mutant structures. The following sentences have been added in the revised manuscript:

      “To show that the introduced substitutions were unambiguously supported by the crystallographic data, electron density maps around residues 48 and 51 are shown in Supplementary Fig. S2.” (p. 4 line 198-200 in the revised manuscript)

      Reviewer #2 (Public review):

      (1) The authors assessed the consequences of variants in the human MutL homologs PMS2 and MLH1, but various other human GHKL ATPases contain clinically relevant variants, some of which have stronger disease associations than the mutations examined in this study. A broader analysis of the effect (or likely effect) of disease-linked mutations in GHKL ATPases would have strengthened this study.

      We agree that extending the analysis to additional disease-associated variants in other human GHKL ATPases would further strengthen our understanding of the conserved catalytic mechanism and its clinical relevance. However, we believe that such a comprehensive analysis is beyond the scope of the present study, which focuses on establishing the fundamental catalytic mechanism shared between MutL and GyrB. We consider systematic functional and structural analyses of disease-associated variants across the GHKL ATPase family to be an important direction for future research. We have now added a statement to the Results and Discussion section to acknowledge this limitation and highlight this future perspective:

      “Although we focused here on pathogenic variants in the MutL homologs MLH1 and PMS2, extending similar structural and biochemical analyses to disease-associated variants in other human GHKL ATPases will be important for evaluating the generality and clinical relevance of the conserved catalytic mechanism proposed in this study.” (p. 6 line 303-306 in the revised manuscript)

      (2) In MLH1, the E37K mutation completely abolishes ATPase activity, but the corresponding mutations in aqMutL, aqGyrB, and PMS2 do not. It remains unclear why E37K in MLH1 leads to complete loss of activity, as the authors propose that water molecule positioning via the first acidic residue, as well as ATP lid stabilisation and associated conformational changes, should still be possible.

      We agree that the complete loss of ATPase activity caused by the MLH1 E37K variant cannot be explained solely by loss of the catalytic carboxylate. However, we note that the corresponding aqMutL E32K variant analyzed in this study also exhibited essentially no detectable ATPase activity, indicating that this phenotype is not unique to MLH1. It can be thought that the severe defect of the lysine variants arises not merely from loss of the acidic side chain but from charge reversal. We have clarified this point in the Results and Discussion sections:

      “In contrast, the E37K mutation in the MLH1 NTD completely abolished the ATPase activity under our assay conditions (Fig. 5B and Table 1) unlike the corresponding glutamine substitutions, which retained substantial residual ATPase activity in aqMutL, aqGyrB, and PMS2 NTDs. A similar complete loss of ATPase activity was also observed for the E32K mutant form of the aqMutL NTD. These observations suggest that the severe defect caused by the lysine substitution cannot be attributed simply to loss of the catalytic carboxylate. Instead, introduction of a positively charged side chain (charge reversal) is likely to perturb the local electrostatic environment. Structural characterization of the MLH1 E37K and aqMutL E32K mutant forms will be required to clarify the molecular basis of this severe functional defect.” (p. 6 line 281-289 in the revised manuscript)

      (3) The authors do not examine ATP binding in the E32 mutants of aqMutL NTD and the D51 mutants of aqGyrB, or AMPPNP binding of the NLH1 and PMS2 mutants. Hence, the relative contributions of the acidic residues to ATP binding and hydrolysis remain partially unclear.

      We performed additional ATP-binding experiments using the aqMutL NTD E32A and aqGyrB NTD D51A mutant forms. Both mutant forms exhibited ATP-binding activities comparable to those of the corresponding wildtype forms. These results support our conclusion that the second acidic residue primarily contributes to ATP hydrolysis rather than ATP binding, whereas the first acidic residue plays dual roles in ATP binding and catalysis.

      Although we agree that nucleotide-binding analyses of the MLH1 and PMS2 variants would be informative, these experiments were not feasible because the equilibrium dialysis assay requires high protein concentrations, which we were unable to obtain for the recombinant human MLH1 and PMS2 N-terminal domains.

      We have incorporated these new data into the Results and Discussion section:

      “In contrast to the E29A mutant form of the aqMutL NTD, the E32A mutant form exhibited ATP binding ability comparable to that of the wildtype form (Supplementary Fig. S1A), indicating that Glu32 does not contribute to ATP binding.” (p. 3 line 143-145 in the revised manuscript)

      “The D51A mutant form of the aqGyrB NTD retained ATP binding ability comparable to that of the wildtype form, indicating that Asp51 is not required for nucleotide binding (Supplementary Fig. S1B).” (p. 4 line 183-185 in the revised manuscript)

      (4) The ATPase assays for PMS2 and MLH1 (Figure 7 and Table 1) were performed with purification/solubility tags still present. Hence, it cannot be ruled out that these tags influence the measured activities.

      We thank the reviewer for raising this important point. We agree that the possible influence of the purification/solubility tags on the absolute ATPase activities of the PMS2 and MLH1 NTDs cannot be completely excluded. However, the wild-type and mutant forms for each homolog were analyzed using identical constructs under the same experimental conditions. Therefore, the affinity/solubility tags are unlikely to affect the relative comparisons of the mutational effects. Furthermore, because the affinity tags are located at the N terminus and are distant from the ATPase active site, they are unlikely to directly perturb the catalytic center.

      (5) The authors suggest that the two-acidic-residue mechanism proposed in this study could be shared among several GHKL ATPase families, yet they also state that the hydrogen-bonding network was not observed in MutL and MORC family proteins. This raises doubt about how conserved the mechanism is, e.g., in MutL and MORC proteins.

      We thank the reviewer for this insightful comment. Our proposed mechanism is based on the cooperative catalytic roles of the two conserved acidic residues, namely the involvement of the first acidic residue in ATP binding and nucleophilic water positioning and the role of the second acidic residue in proton abstraction. In contrast, the Glu48–Gln340 hydrogen-bonding interaction described in aqGyrB was proposed only as a structural feature that may modulate the contribution of the first acidic residue to ATP binding. It is not an essential component of the catalytic mechanism proposed in this study. Therefore, the absence of this particular hydrogen-bonding network in the currently available structures of MutL and MORC proteins does not argue against conservation of the catalytic mechanism itself.

      Recommendations for the authors:

      Reviewing Editor Comments:

      One of the structures (Crystal Structure of the E48A variant) has relatively poor statistics in the PDB validation report. Please improve this structure.

      We performed additional refinement of the E48A crystal structure. This resulted in a clear improvement in the overall model quality, with the Ramachandran favored residues increasing from 93.4% to 95.4%, the percentage of side-chain outliers decreasing from 6.1% to 1.4%. The refined structural model has been used throughout the revised manuscript, and the updated refinement statistics are provided in Table 2.

      Reviewer #1 (Recommendations for the authors):

      Please show conventional density maps (e.g., sigmaA weighted 2fo-fc maps).

      This comment is closely related to Comment (2) in the Public Review by the Reviewer #1. In response, we have added Supplementary Fig. S2, which presents conventional σA-weighted 2F<sub>o</sub>–F<sub>c</sub> electron density maps around the catalytic acidic residues in the wild-type and mutant aqGyrB structures.

      Reviewer #2 (Recommendations for the authors):

      (1) Regarding the analysis of clinical variants, it would be informative to note that the second allele is lost before tumor growth in Lynch syndrome.

      “Therefore, these variants might contribute to the development of Lynch syndrome by weakening the ATPase-driven regulatory functions of MutL.” (p. 6 line 280-281 in the original manuscript) has been changed to:

      “In individuals carrying these germline variants, subsequent loss or inactivation of the remaining wildtype allele would leave only the ATPase-defective MutL protein, thereby compromising mismatch repair and promoting tumorigenesis.” (p. 6 line 296-298 in the revised manuscript)

      (2) P. 4, in the paragraph "Conserved roles of two acidic residues of aqGyrB in ATP hydrolysis", the E48Q mutant retains approximately one third of the WT activity, not one quarter as stated in the text (Table 1). Additionally, later in the article, the D51 mutant is reported to retain approximately one-sixth (~17%) of the WT activity, rather than ~25% as written.

      We thank the reviewer for carefully identifying these inconsistencies. The text has been corrected to accurately reflect the data presented in Table 1: “…one third of the wildtype activity” (p. 4 line 180) and “…retaining ~16%...” (p. 4 line 186 in the revised manuscript)

      (3) P. 6, lines 275-276, this sentence should be rephrased for clarity, as the authors note at the end of page 5 that not all members of the GHKL ATPase family possess this second acidic residue.

      “…this second acidic residue plays a conserved and functionally significant role in ATP hydrolysis across the GHKL ATPase family.” in the original manuscript has been changed to:

      “…this second acidic residue plays a conserved and functionally significant role in ATP hydrolysis among some members of the GHKL ATPase family.” (p. 6 line 292 in the revised manuscript)

      (4) P. 9, in the "Data Accessibility Statement", the PDB code 23UY is missing. This entry corresponds to the crystal structure of the D51A mutant of aqGyrB NTD and should be included.

      The Data Accessibility Statement has been revised to include the code 23UY. (p. 9 line 454 in the revised manuscript)

      (5) P. 14, the table should be labelled "Table 2. Data collection and refinement statistics for the aqGyrB NTDs", rather than "Supplementary Table 2", to ensure consistency with how it is cited in the main text.

      The table title has been corrected from "Supplementary Table 2" to "Table 2”. (p. 4 line 198 in the revised manuscript)

      (6) It is difficult to determine from the figures whether the magnesium ion is positioned equivalently in aqMutL and aqGyrB. Did the authors observe any differences in ion positioning?

      To facilitate direct comparison of the catalytic Mg<sup>2+</sup> ion between the aqMutL and aqGyrB NTDs, we have added Supplementary Fig. S3, which shows a structural superimposition of the ATPase active sites of the two proteins:

      “Structural superposition of the aqGyrB NTD and aqMutL NTD revealed that the catalytic Mg<sup>2+</sup> ion occupies essentially the same position in the two ATPase active sites (Supplementary Fig. S3), indicating that the metal-binding geometry is highly conserved, where the Mg<sup>2+</sup> ion is coordinated by the side chain of the conserved Asn, AMPPNP, and surrounding water molecules. Neither Glu48 of aqMutL nor Asp51 of aqGyrB directly coordinated the Mg<sup>2+</sup> ion.” (p. 5 line 201-205 in the revised manuscript)

      (7) The authors should discuss the interaction between the aqGyrB NTD, Mg<sup>2+</sup>, and ATP during the binding step. In the case of the E48A mutant, where ATP binding is lost, does E48 directly establish contacts with Mg<sup>2+</sup>, or is another residue involved (with conformational changes preventing this interaction)?

      Our structural analyses indicate that Glu48 does not directly coordinate the catalytic Mg<sup>2+</sup> ion. Instead, as shown in Supplementary Fig. S3, the Mg<sup>2+</sup> ion is coordinated by the side chain of Asn52, AMPPNP, and surrounding water molecules. We have clarified this point in the Results and Discussion sections:

      “Structural superposition of the aqGyrB NTD and aqMutL NTD revealed that the catalytic Mg<sup>2+</sup> ion occupies essentially the same position in the two ATPase active sites (Supplementary Fig. S3), indicating that the metal-binding geometry is highly conserved, where the Mg<sup>2+</sup> ion is coordinated by the side chain of Asn52, AMPPNP, and surrounding water molecules. Neither Glu48 nor Asp51 directly coordinated the Mg<sup>2+</sup> ion.” (p. 5 line 201-205 in the revised manuscript)

      (8) Figures 1 and 3: Use ribbon representation and no shadows, at least for the inset panels, to enhance clarity and interpretability.

      We have revised Figures 1 and 3 by displaying the protein structures in ribbon representation and removing shadows from the inset panels.

      (9) Combine Figures 1 and 2, and combine Figures 3 and 4.

      Following the reviewer's recommendation, we have combined the original Figures 1 and 2 into a single figure and the original Figures 3 and 4 into another single figure.

      (10) Figure 5: Zoom in further and remove shadows. The current panels are not very effective in highlighting how ATP is bound by the different protein variants.

      Figure 5 has been revised by increasing the magnification of the ATP-binding sites and removing shadows from the structural renderings.

      (11) Figure 8. Add a scale bar to show evolutionary distance.

      We thank the reviewer for this helpful suggestion. To provide information on evolutionary distances while preserving the clarity of the main figure, we have added a new Supplementary Fig. S5 showing the same phylogenetic tree with branch lengths proportional to the inferred evolutionary distances and an evolutionary distance scale bar. Figure 6 has been retained in its simplified form with equal branch lengths to facilitate visualization of the ancestral-state reconstruction, and we have clarified this distinction in the Materials and Methods section:

      “For visualization purposes, branch lengths were not scaled and were displayed with equal lengths in Fig. 6. The corresponding phylogeny with branch lengths proportional to the inferred evolutionary distances is provided in Supplementary Fig. S5.” (p. 9 line 429-432 in the revised manuscript)

    1. eLife Assessment

      This study offers valuable insights into brain responses to somatosensory and visual temporal and spatial tasks in the auditory cortex of Deaf and hearing individuals. The evidence for a sensorily-bound representation in the deprived cortex -- rather than abstract -- is solid; however, the study will benefit from some additional analyses to better clarify and contextualise the results. This work will be of broad interest to neuroscientists investigating brain plasticity and development.

    2. Reviewer #1 (Public review):

      Summary:

      The authors conducted a carefully constructed experiment to test reorganization in the auditory cortex in deafness in response to task (spatial vs. temporal working memory) and modality (visual vs. somatosensory). They found a complex pattern of results, which included changes to univariate response strength in deafness that differed between the primary and association auditory cortex. HG showed a preference for the somatosensory working memory task, whereas STG/S responded more in both modalities for the temporal task. They further showed multivariate similarity of their results to models representing both task and modality in both groups, which were increased in deafness. Curiously, the task effect was for a sensorily-bound model, which shows mid-level representation, and not high-level task ("metamodal") code.

      Strengths:

      I appreciated the matched design, rigorous analysis, and careful interpretation of the nuanced results.

      Weaknesses:

      Only minor weaknesses: behavior and residual hearing can be better controlled.

    3. Reviewer #2 (Public review):

      Manini and colleagues present an interesting study on the consequences of early deafness on the organization of temporal regions chiefly engaged in audition in hearing people. Mainly relying on representational similarity analyses, they show that the auditory cortex in deaf individuals represents information about task, sensory modality, and somatosensory frequency. Critically, task and modality representations were also found in the auditory cortex of hearing individuals. There were significant differences between groups, implying that these representations are enhanced as a consequence of deafness.

      Overall, I feel that the paper could gain in clarity and impact if the hypothesis space tested in the introduction and discussion was made clearer, if some new analyses were provided to support some claims, and if the authors better matched their conclusions to the observed results.

    4. Reviewer #3 (Public review):

      Summary:

      This manuscript examines functional plasticity in auditory cortices of people born deaf (or early deaf). Deaf participants (N=13, all native BSL signers) and hearing controls (N=18) performed a delay-match-to-sample working memory task in visual and somatosensory modalities, attending either to frequency (temporal task) or spatial pattern (spatial task). Using fMRI and univariate and representational similarity analysis (RSA), the authors test what type of information is represented in auditory areas of hearing controls and deaf participants. Different types of tasks and stimulus features are examined, including low-level sensory features (vibration/movement frequency and spatial position of stimuli on screen/hand) as well as task features (attending to space vs. frequency) and stimulus modality (somatosensory vs. visual).

      There are a number of interesting findings. In early auditory cortices (right Heschl's gyrus), only Deaf participants show above baseline responses and only in the somatosensory task. In secondary auditory/multisensory STS, only deaf participants show above-rest responses to both visual and somatosensory stimuli. In RSA analysis, d/Deaf, but not hearing participants, show sensitivity to the frequency of somatosensory vibration when finger position is held constant (SFRm model). RSA in auditory cortices finds sensitivity to sensory modality (visual vs. somatosensory) information in both hearing and deaf participants, with an enhanced effect in the d/Deaf group. A similar d/Deafness enhancement effect was observed for task (frequency vs. spatial) but only within modality, not across, suggesting a less abstract representation.

      Strengths:

      The paper has a number of strengths. Running tasks in multiple modalities in the same study is a technical challenge and adds valuable information, since, as it turns out, early auditory areas of deaf people are sensitive to somatosensory but not visual frequency. Moreover, it makes it possible to look for modality coding.

      The RSA analysis finds sensitivity to task and modality features not detectable with univariate analyses.

      Overall, the findings of differences across modalities (visual and somatosensory) in auditory cortices are very interesting. Having parallel tasks in the two modalities makes it all the more important that only somatosensory stimuli activate HG. The discussion of this finding and ideas about alternative possible paths of somatosensory and visual information to the auditory cortices in d/Deaf individuals is also interesting.

      The authors conclude that there is both evidence for shared and different function across hearing and deaf groups, and this makes sense. It's refreshing that the authors acknowledge that the simplicity dichotomy of preservation vs. change present in the literature and presented in the introduction turns out not to explain the findings.

      Weaknesses:

      The participant sample is not very large, but this is a difficult-to-recruit population, and it is within an acceptable range since the authors have taken care to run a robust study design and collect a large amount of data from each person.

      Some weakness of the paper includes incomplete presentation of the results and conclusions that do not follow from the data.

      The results are presented in a way that makes it hard to track what is significant and how the auditory ROIs are different or similar to the control regions. It is also difficult to connect the written text results with the figures. In some cases, the figures look like there is no significant effect, but the results report that there is one.

      It is not clear that three-way interactions were tested for (e.g., group, by task, by modality). This complicates the interpretation of significant two-way interactions, e.g., task by group. For example, in STG a task by group interaction is reported, but the plot suggests that this is driven primarily by the somatosensory modality.

      The paper suggests that frequency/temporal information is coded in auditory areas of d/Deaf participants but not location/spatial information. This is stated in the Results and in the Discussion as a major point. But this is not quite true in the somatosensory case and not true in the visual case at all, as far as I can tell. In the somatosensory case, since frequency is only coded in a finger-dependent manner, location is in fact coded in this regard. This pattern differs from what is observed in the somatosensory cortex, where frequency is coded in a finger-dependent and independent manner as well as location as such. In the visual case, it seems like temporal, i.e., frequency information, is not coded at all in the auditory ROIs. Although it is also puzzling that visual frequency is not coded in the visual 'control' ROI. An incomplete presentation of results motivates some conclusions that are not warranted, e.g., coding in the auditory cortex reflects a preservation of its function - i.e., frequency but not location coding.

      A second related issue is that some dimensions (e.g., visual frequency, modality-independent task) show no neural response anywhere in the brain, including in the canonical visual, somatosensory, and amodal networks. For these dimensions, the current experiment does not offer a good test. This is fine but should be clearly stated in the results and the Discussion so there is no confusion about which hypotheses are really tested. Right now, the results say things like "the control ROI shows the expected pattern", but in some cases it's more complicated and failures to observe effects constrain what can ultimately be expected in auditory areas. This is okay, but needs to be made clear in the results and discussion, and auditory results need to be interpreted in this context.

      Relatedly, the paper has no whole cortex searchlight analyses, and it is not clear why. If nothing comes out in the small sample of n=13, that is okay, but at least the collapsed hearing and d/Deaf sample data should be shown for each dimension. This will give the reader a sense of what is to be expected and contextualize the results.

      Some claims are made which are not supported by the data: "However, it is likely that the representation of somatosensory frequency in deaf individuals relies on the same mechanisms used to represent auditory temporal frequency in hearing individuals." Likewise, the paper goes on to say, 'the underlying computations might be the same'. True, they might be, but they also might be different. No evidence is presented to support one or the other hypothesis, so both should be stated, and it should be stated that these cannot be distinguished based on the presented data.

      Conclusions-wise, the paper sometimes makes sweeping claims that go well beyond what the evidence supports and fails to provide caveats. The first paragraph of the Discussion states: "Overall, these findings suggest that crossmodal plasticity relies on representational and functional configurations that are present across individuals and modulated by sensory experience." Such a sweeping conclusion about cross-modal plasticity in general is not supported or refuted by the present data. There is no evidence that 'representations' or 'functional configurations' are the same across groups. What are functional configurations? What representations are shared? Second, it is far too general to make claims about 'cross-modal plasticity' based on one study with one population.

    5. Author response:

      We would like to thank the editor and reviewers for their thoughtful and constructive feedback. We appreciate the time and care devoted to reviewing our manuscript, as well as the recognition of rigorous experimental design, the technical challenges involved in conducting an fMRI study with two sensory modalities and two tasks in both deaf and hearing participants, and the value of the findings for understanding crossmodal plasticity and cortical organisation in deafness. We are encouraged by the overall assessment of the study, and appreciate the suggestions for strengthening the manuscript. Below, we provide a summary of how we plan to address the reviewers’ comments in our formal revision of the manuscript:

      (1) Additional analyses

      (a) We will incorporate behavioural performance measures into the relevant analyses to disentangle potential behavioural contributions to the observed effects.

      (b) We will calculate the noise ceiling value for each of the RSA analyses. 

      (c) We will conduct a correlation analysis between RDMs of auditory and control regions, to investigate the similarity between these computations and whether this is influenced by sensory experience.

      (d) Regarding the suggestion to conduct whole-brain searchlight analyses, we respectfully do not believe that this approach would address the primary research question of the study, namely whether and how representations within auditory cortex differ between deaf and hearing individuals. Our central hypotheses specifically concern representational content within predefined auditory cortical regions, making the ROI-based approach the most appropriate and sensitive method for testing these questions.

      Furthermore, the searchlight approach would require adequately powered group comparisons at the whole-brain level. Given the challenges associated with recruiting deaf native signers participants and the resulting sample size, we do not believe the study is sufficiently powered to draw reliable conclusions from this analysis. We will further clarify this rationale in the revised manuscript.

      (2) Presentation of the results

      We will revise the presentation of the findings to better guide the reader through the analyses and facilitate interpretation of the figures. In particular, we will ensure that significant effects, interactions, and their relationship to the corresponding figures are described more explicitly throughout the manuscript.

      (3) Revision of the discussion

      Following the reviewers’ feedback, we will revise the Discussion to more clearly distinguish between results that directly support a conclusion and hypotheses that remain speculative.

    1. eLife Assessment

      This study provides a useful investigation of machine learning approaches that can lessen potential gaps in the prediction of behaviour from brain imaging data across majority and minority samples. The authors provide incomplete evidence to suggest that domain adaptation methods can mitigate these gaps. The analysis would benefit from further testing of model generalisability and the inclusion of recommended workflows for how the tested approaches can be used in future research. This work will be of interest to scientists using machine learning in brain imaging.

    2. Reviewer #1 (Public review):

      Summary:

      The present report describes an investigation into the use of machine learning techniques to improve cross-racial/ethnic performance of brain models of cognitive function. The authors tested several approaches to boost prediction of NIH cognitive toolbox scores using brain imaging data (function, structure) for minoritized (Black) participants in the ABCD Study sample compared to white (majority) participants. Structural (e.g., volume) measures showed the greatest performance gap, and a balanced weighting method showed the greatest performance gain across features. The authors conclude that supervised domain adaptive methods can improve models for cognitive prediction and mitigate cross-racial/ethnic performance disparities.

      Strengths:

      This investigation makes some headway into issues by identifying computational methods that may help to improve some models for limited outcome variables (i.e., general cognitive performance). Addressing racial/ethnic disparities in brain imaging research has significant implications for generalizability of findings and for the practical utility of imaging findings in the wider population. A comparative approach to evaluate the improvements in a "prediction gap" across various methods could have benefits for neuroimaging beyond racial/ethnic disparities. The use of the ABCD Study, given its deep phenotyping of individuals, is also a benefit.

      Weaknesses:

      Despite its strengths, there are several large conceptual and related methodological issues that impact its conclusions and the overall utility of the approach. The sample selection approach limits insight into likely drivers of the performance gap (e.g., socioenvironmental disparities known to exist between groups and associated with neurodevelopment), and in so doing ignores a critical component of understanding racial brain differences, particularly in relation to cognitive functions. Further, while the relative gaps in performance of a single cognitive score across features are well described, the actual performance (and therefore relative benefit to these techniques) is unclear. Specific examples include the following.

      (1) The overarching conceptual issue with the manuscript is a lack of engagement with a substantial and growing evidence base on the drivers of racial disparities in brain imaging which impact model performance. Racial/ethnic groups in the US (and other regions of the world) are not equivalent in terms of developmental environments that shape brain function and structure (see Harnett et al., 2023, Neuropsychopharmacology; Ricard et al., 2023, Nature Neuroscience; Cardenas-Iniguez & Gonzalez, 2024, Nature Neuroscience for some overview here). The socioenvironmental disparities inherent to race in the US further shape cognitive development and brain associations with cognitive performance (e.g., Marek et al., 2025, Science). The framing of the manuscript focuses almost exclusively on broad sampling issues, and in doing so treats racial/ethnic variability as if it reflects statistical abnormality rather than a critical component of understanding human brain development. This lack of contextualizing racial disparities significantly impacts the overall utility of the proposed approach and the conclusions of the manuscript.

      (2) In relation to the above, another conceptual issue in this approach of using a majority to inform minority brain associations with cognitive variables is an assumption that minority brain patterns should match the majority, rather than developmental stressors inducing alternative brain-weighting to predict outcomes. This framework does not assess this possibility and may in fact obscure such an outcome, limiting our inferences into neurodevelopment.

      (3) Another conceptual/methodological issue here is the use of "matched groups" for analysis. The specifics of matching are fairly vague, but given the description one would assume the w/B groups are matched on a number of behavioral/socioenvironmental variables, which is a significant issue for interpretability and applicability. As noted, w/B groups in the US (and the ABCD Study) differ substantially across variables; matching has the likely consequence of creating a highly non-generalizable sample, particularly when the minority group is restricted to N = 10.

    3. Reviewer #2 (Public review):

      In the manuscript "Supervised domain adaptation mitigates cross-ethnicity prediction errors in neuroimaging-based cognitive prediction", the authors investigated the efficacy of data adaptation techniques to reduce ethnicity-related prediction bias in neuroimaging-based cognitive prediction. They found that data adaptation algorithms, particularly balanced weighting, contributed to mitigating ethnicity-related performance disparities. Furthermore, these bias mitigations could be achieved without requiring a large set of data from the underrepresented ethnic group. This study addressed an important concern in the field of neuroimaging-based behaviour prediction, providing many intriguing results. Nevertheless, the manuscript also suffers from a lack of coherent methods design, the unorganised presentation of information, and the lack of in-depth discussion of results.

      The conclusions claimed by the authors are sometimes over-generalised and not fully supported by the study outcomes. Overall, this study demonstrated strong technical designs and convincing statistical analysis for the main outcomes, although clearer presentation would be needed to convey the messages in the manuscript.

      The central investigation of this study is whether domain adaptation techniques improve ethnicity-related performance disparities. However, these improvements were only measured against a very weak baseline model, where a small set of African American (AA) subjects were added to the training sample consisting purely of White American (WA) subjects. While the authors recognised that balancing the training sample could already mitigate the ethnicity-related disparities, they considered that such approaches are unfeasible in their experimental scenario, where only a small amount of AA data were available. However, as Li et al. (2022) showed, a balanced sample of around 90-150 AA subjects could already reduce the ethnicity-related bias. Even from a practical standpoint, this balanced sample approach would be a more valid baseline for domain adaptation models to compare against.

      The authors made two main conclusions: that domain adaptation methods reduced ethnicity-related bias, and that balanced weighting performed the best and the most stably. Both claims were over-generalised to some extent. First, the adaptation benefit claimed in the first conclusion is not seen in the functional connectivity (FC) modality, which is the most popular modality for neuroimaging-based prediction of behaviour. This difference in adaptation benefit across modalities is an important finding that is meaningful for future studies, the omission of which also removes interesting insights that the audience could take away from this article.

      Second, the judgement of prediction performance is based on the area under the improvement curve (AUIC) metric, which summarises a model's performance across different availability of labelled AA data. As a result, the analysis of prediction performance naturally favours algorithms that could perform well with a small amount of added AA data. On the one hand, this provides an easy decision point for users to pick an algorithm to use without being concerned about data availability. On the other hand, important insights could be overlooked with the oversimplified recommendation of balanced weighting. As the authors have also observed, in some cases, domain adaptation strategies do not improve ethnicity-related bias more than the non-adaptation baseline. If the message is to recommend simple, low-cost strategies to reduce ethnicity-related prediction bias, it would be misleading not to note that the simplest and lowest-cost strategy could also be non-adaptation methods sometimes.

      Regardless, for the general audience, the underlying assumptions when interpreting the AUIC metric are not immediately clear, which could cause the conclusions to be misleading. Apart from aggregating over different amounts of available AA data, the statistical comparison of AUIC gain across data adaptation algorithms also did not account for the impact of brain phenotype modalities. Even though the upstream analyses have confirmed that adaptation benefits vary greatly across brain modalities, this major observation was not followed in the final analysis where conclusions were made about which algorithm performed the best. Based on visual inspection of Figure 3b, it may be suspected that PRED performed better than or comparably to balanced weighting when task contrasts based on the Destrieux atlas were used.

      Finally, the findings from this study align with the common hypothesis that ethnicity-related prediction bias originates from disparities already manifested during data collection and preprocessing. As the authors have noted, the modalities with the most tendency for ethnicity-related bias are the anatomical ones, including all three volume-based modalities (cortical volume, T1 and T2 subcortical volume) in the top ten phenotypes with the largest performance gap. Most prominently, brain features in the occipital pole, frontal pole, and a range of subcortical areas were found to contribute highly to adaptation gain. Subcortical areas are often reported to show noisier measurements compared to cortical areas, whereas the poles of the brain are likely more strongly warped/distorted during alignment to a standard template. From a data quality perspective, these results support the interpretation that ethnicity-related prediction bias may stem from loss of data quality during data collection or preprocessing. In the prediction models based on anatomical brain features, data adaptation methods may have helped to address these disparities in the data, without the more resource-intensive need to improve the bias in preprocessing pipelines.

      Li, J., Bzdok, D., Chen, J., ... Genon, S. (2022). Cross-ethnicity/race generalization failure of behavioral prediction from resting-state functional connectivity. Science Advances, 8(11), eabj1812.

    4. Reviewer #3 (Public review):

      The manuscript frames its work in fairness and disparities but does not show or directly test that its approach decreases differences between White and African Americans. While it is stated that the objective is not to equalize performance across groups, large parts of the paper repeatedly claim that the methods mitigate cross-ethnicity disparities and improve fairness. Improving prediction in African American participants relative to a non-adapted model is not necessarily the same as reducing the disparity between African American and White American participants. The adapted model should be evaluated in both groups, and the post-adaptation performance gap should be reported directly.

      Additional prediction performance measures are needed. For example, in Li et al, different results and conclusions are made with MSE and the correlation between observed and predicted variables. In that paper particularly, aggression measures showed better correlation in African Americans but better MSE in White Americans. Such differences are important to note as they likely suggest different mechanisms.

      Similarly, characteristics of the cognitive outcome need to be understood. For example, differences in MSE or MAE may reflect a difference in variance between the groups. The group with a larger variance will have a larger MSE. Correlation or other performance measures that are invariant to different mean or variance shifts can be helpful here.

      While the authors note that for the paper they treat racial and ethnic backgrounds interchangeably, I do not think that is the best given the differences between them and the impact and history they have in American culture. Overall, the authors likely need to do a better job conceptualizing their results in the history of minoritized populations in the United States. It is immensely important not to treat them as biological domains without considerable qualification and to avoid language implying that observed domain differences are intrinsic properties of racial groups.

      Changes in feature weight are not a proper way to identify the mechanisms of improved performance. At most, these analyses characterize how model coefficients change when target-group data are incorporated or upweighted.

      The cross-validation strategy is suboptimal. First, the use of the matched splits of the ABCD data introduces data leakage. To match a validation set to the training set in such a manner requires that each split knows about the other split's characteristics. That is data leakage. Though the impact could be small. Second, African American breakdowns are not balanced across sites and scanners. Domain adaption methods may be learning a shortcut or proxy for African American like site, scanner, or something else. A likely better approach would be some sort of leave X sites out approach, where a model is trained on White Americans from a set of sites, adapted with African Americans from those sites, and applied (with and without adaptation) to the White and African Americans from the left-out sites.

      Baseline models for comparisons to the domain adaptation are missing. Some simpler ones include a target-only model trained on the same 10-100 African American participants and a pooled model with a group indicator and group-by-feature interaction. Without these comparisons, it is difficult to know whether balanced weighting is learning target-specific neurobiological information or merely recalibrating the prediction distribution.

      There are a few statistical issues:

      (1) The repeated MAE estimates are therefore not independent observations. Paired t-tests cannot be applied across repetitions. Subject-level bootstrap or permutation procedures that repeat the complete training and testing process are needed

      (2) The caption describes approximate 95% confidence intervals as {plus minus}1.96 × SD/n. Conventionally, the standard error would involve SD/sqrt(n). However, even if corrected, there would still be issues about the dependence among the overlapping resamples.

      (3) Ten repetitions are likely insufficient, especially in the case of ten target participants. Results at n = 10 may be extremely sensitive to which children are selected. The authors should use substantially more repetitions and report the full distribution of results.

      (4) The Friedman and Wilcoxon comparisons treat the 80 imaging phenotypes as the observational units. These phenotypes are highly dependent because they are derived from the same participants, many use overlapping images, and numerous task contrasts and structural measures are strongly correlated. This non-independence can make the comparison among adaptation methods look much more precise than it is. A hierarchical analysis by modality or a resampling strategy that preserves dependence among phenotypes would be more appropriate.

      (5) The gap metric and AUIC are difficult to interpret. Gap is the absolute relative difference between target-group and source-group MAE, normalized by source-group MAE. It is sensitive to the denominator and may produce large values whenever source-group. MAE is relatively small. Reporting signed raw MAE differences and MAE ratios alongside this derived score would help. Similarly, the AUIC combines errors with an arbitrary sequence of target-sample sizes. IStatistically significant differences in AUIC do not necessarily indicate practically meaningful differences among methods.

      (6) The correlation between baseline gap and adaptation gain is partly tautological. Those with the widest gaps have the most room for improvement and likely thus show the greatest improvement. While still of value, the authors may want to tone down their interpretation of the correlation and describe its limitation.

      (7) Given that the sample sizes vary from approximately 4,000 to more than 11,000 depending on modality, the authors may want to consider a reduced sample matched in size across modalities. It is hard to fully know if the conclusion that connectivity is more robust given the wide-scale differences in sample size and feature dimensionality.

      (8) PLS are sensitive to many factors like scaling and collinearity. Many recent papers have been written about their limitations when used for subtyping. Some of these hold for prediction too. I think showing the results are consistent with different prediction algorithms is needed. SVR and ridge regression are two common methods for regression prediction with neuroimaging data.

      (9) The feature interpretation is partly circular. The method with the largest performance gain is selected, and its coefficient changes are then used to explain that gain. A method designed to give target observations greater influence will unsurprisingly change its coefficients more than naïve inclusion.

      (10) The practical and ethical deployment scenario is underspecified. Supervised adaptation requires labelled cognitive outcomes from the target population and, as currently framed, may require choosing a model based on an individual's racial category. What are the implications of deploying race-specific models that need to be considered? It is not self-evident that this approach is preferable to developing a broadly representative model or directly modeling the social and technical sources of distribution shift.

      (11) The paper is worded and interpreted much too strongly. The current study supports the conclusion that, within ABCD, giving a small labelled target-group sample greater influence can sometimes improve held-out target-group MAE relative to naïvely adding the same participants. It does not yet establish that the method improves fairness or identifies mechanisms of racial bias. Likewise, the abstract and conclusion overstate the results. The abstract states that all adaptation methods reduced target-group prediction error, while the Results show near-zero or negative benefits for several functional-connectivity phenotypes and instability of PRED and interpolation below 30 target participants. Similarly, "substantially reduce disparities," "improve equity," "consistently," and "practical path forward" are stronger than the analyses support. Finally, the limitations section is incomplete and omits the more consequential limitations.

      (11) That only ten labelled participants are needed to change the results is troubling. This is a shockingly low number. Giving ten target observations disproportionate influence can move the fitted model, particularly when the balanced-weighting ratio is high. A measurable MAE change is therefore possible, but it may reflect a shift in intercept or slope rather than learning a stable target-group brain-cognition relationship. Further, the manuscript does not report the numerical improvement for the n = 10 condition in the text or a table. Visual inspection of Figure 4 suggests reductions of roughly 0.10-0.25 standardized MAE units for some high-gap structural phenotypes, approximately 10-20%, while low-gap connectivity phenotypes show little or no gain.

      (12) The study lacks genuine external validation, which may be needed to fully convince readers that such a low number of subjects is needed to reduce biases.

    1. eLife Assessment

      This study presents a useful finding on using diverse experimental systems to understand how neuromodulatory signals shape glial inflammatory signaling; however, the strength of evidence is inadequate. The astrocyte-enrichment method used may permit contamination by microglia, oligodendrocyte-lineage cells, or neurons, complicating attribution of TNF expression specifically to astrocytes. This concern is compounded by the strong microglial response to Gi manipulation and the lack of quantitative validation of chemogenetic cell-type specificity in vivo. These weaknesses have hindered further evaluation of the claims.

    2. Reviewer #1 (Public review):

      In this manuscript, the authors explore whether GPCR signaling in astrocytes affects the production of TNF by astrocytes and, to a lesser extent, microglia. Unfortunately, the method used by the authors to acquire astrocyte-enriched cultures is known to result in meaningful rates of contamination by myeloid cells (microglia and others), oligodendrocyte-lineage cells, and neurons. Alternative methods of generating highly enriched astrocyte cultures, as well as purifying astrocytes with little to no neuronal or myeloid contamination across age and brain regions, have shown no evidence of TNF expression by astrocytes (Zhang et al., J Neurosci, 2014; Zhang et al., Neuron, 2016; Clarke et al., PNAS, 2018). In fact, the paper cited by the authors as demonstrating differences between human and rodent astrocytes found no evidence of TNF expression in immature or mature human astrocytes (Zhang et al., Neuron, 2016). The idea that the majority of the observed TNF transcriptomic signal, at least in culture, comes from myeloid or neuronal contamination also aligns with the authors' observation that myeloid-enriched cultures act identically to astrocyte-enriched cultures.

      The authors also use a GFAP virus to drive GPCR signaling in astrocytes and neuronal progenitor cells in their cultures, but, given that these cultures are known to have meaningful contamination by other cell types, such signaling could be due to astrocyte → microglia/neuron signaling or other multicellular pathways that cannot be excluded. Similar concerns mean that we cannot assume the effect of DREADD activation of astrocytes in vivo (Figure 6) reflects a bulk change in TNF expression driven by astrocyte-specific changes rather than by multicellular signaling.

      The most compelling evidence for their claim of astrocyte TNF expression comes from the human-induced astrocytes. However, their antibody staining is not sufficient to claim these cells are truly astrocyte-like. Antibody staining is highly prone to non-specificity, as highlighted by the fact that their ALDH1L1 antibody staining appears perfectly nuclear despite ALDH1L1 being a cytoplasmic protein.

      To address both the purity concerns of the astrocyte-enriched cultures and the concerns about the astrocyte identity of the induced astrocytes, the authors should perform RNA sequencing. By profiling gene expression in these cultures at the genome-wide level, readers can truly assess the degree of contamination and thus the likelihood of the proposed mechanism (i.e., astrocyte-specific TNF production). Importantly, previous studies have suggested that very little neuronal and myeloid contamination is required to dramatically change cellular responses (Foo et al., Neuron, 2011; Liddelow et al., Nature, 2017).

    3. Reviewer #2 (Public review):

      Summary:

      Abbasi et al. examine how signaling through the major G-protein pathways (Gs, Gq, and Gi) influences tumor necrosis factor expression in astrocytes and microglia. Using a combination of pharmacological receptor activation, chemogenetic manipulation, primary rodent glial cultures, human induced pluripotent stem cell-derived astrocytes, and an in vivo astrocyte-targeted Gi manipulation, the authors report a broadly consistent pattern in which Gs- and Gq-associated signaling reduces tumor necrosis factor expression, whereas Gi signaling increases it. The study's cross-species and cross-preparation design, spanning astrocytes and microglia as well as in vitro and in vivo systems, provides a potentially valuable framework for understanding how neuromodulatory pathways may regulate glial inflammatory signaling.

      Strengths:

      A major strength of the study is the breadth of experimental systems used, which includes primary rat glia, human induced pluripotent stem cell-derived astrocytes, and an in vivo manipulation, allowing for comparison across species and levels of biological complexity. The use of chemogenetic receptors in astrocytes provides relatively direct control over Gq and Gi signaling, and these experiments yield consistent effects on both tumor necrosis factor messenger RNA and protein, strengthening the internal validity of the astrocyte findings. The observation that similar directional effects are seen in human-derived astrocytes and in microglial cultures further supports the idea that aspects of this regulatory relationship may be conserved across glial cell types. More broadly, the study addresses an important and timely question about how neuromodulatory signaling pathways interface with glial inflammatory outputs, and it generates a coherent set of observations that could serve as a foundation for more mechanistic work.

      Weaknesses:

      The central claim that Gs, Gq, and Gi signaling broadly and directly constitute a general regulatory code for tumor necrosis factor expression is more expansive than the current evidence fully supports. In particular, the evidence for Gs-dependent effects is indirect, relying on beta-adrenergic receptor activation and forskolin-mediated adenylyl cyclase stimulation rather than direct manipulation of Gs itself, leaving uncertainty about pathway specificity. More generally, the use of different endogenous receptors to represent each G-protein class in microglia complicates interpretation, since individual receptors may engage additional signaling pathways beyond their canonical G-protein coupling, limiting the extent to which the results can be attributed to G-protein class alone.

      The in vivo experiment also does not definitively establish the cellular source of the observed increase in tumor necrosis factor, as measurements are taken from bulk cortical tissue following astrocyte-targeted Gi activation. This leaves open the possibility that the observed changes arise indirectly from other cell types, particularly microglia, which are shown elsewhere in the study to be strongly responsive to Gi-related manipulations. In addition, the specificity of chemogenetic expression in vivo is not quantitatively demonstrated, further limiting cell-type attribution.

      There are also important issues related to experimental design and statistical interpretation. Across several experiments, it is unclear whether reported sample sizes reflect independent biological replicates, technical replicates, or imaging fields, which is especially consequential for the human induced pluripotent stem cell-derived astrocyte experiments where donor-level independence is not clearly established. The in vivo design also appears to treat hemispheres as independent observations despite their paired nature, which may inflate statistical independence given the small sample size.

      Finally, several conclusions would benefit from more cautious framing. The data support differential regulation of tumor necrosis factor relative to interleukin-1 rather than strict cytokine specificity, and measurements based solely on messenger RNA should not be interpreted as direct evidence of cytokine production. The comparison between glial signaling effects and neuronal excitation or inhibition also juxtaposes fundamentally different biological readouts and should not be interpreted as a direct functional opposition. Overall, while the study provides interesting and potentially important observations, the broader pathway-level and cell-type-specific conclusions are not yet fully established by the current experimental evidence.

    1. eLife Assessment

      This study addresses a fundamental question about how large-scale brain networks interact, and specifically how the default mode network exchanges information with sensory cortex. The analyses provide solid evidence for the claims made in the paper. The findings should be of broad interest to researchers studying brain network organization and dynamics.

    2. Reviewer #1 (Public review):

      Summary:

      This paper leverages 7T fMRI data from the Natural Scenes Dataset to investigate whether retinotopic coding the position-selective organization of visual responses structures spontaneous resting-state interactions between the Default Network (DN) and the Dorsal Attention Network (dATN). Using individualized network parcellations and population receptive field (pRF) modeling, the authors show that DN voxels can be split into two subpopulations based on their response to visual stimulation: those with position-specific positive BOLD responses (+pRFs) and those with position-specific negative BOLD responses (-pRFs). Critically, these subpopulations relate differently to the dATN during rest: -pRFs are anticorrelated with the dATN, +pRFs are positively correlated, and non-retinotopic DN voxels show no coupling. The anticorrelation (and positive correlation) is enhanced when DN and dATN voxels share visual field preferences. An event-triggered analysis suggests that retinotopic coding shapes both "top-down" (DN-initiated) and "bottom-up" (dATN-initiated) spontaneous activity transients, supporting the claim that the retinotopic scaffold is intrinsic to the DN. These findings challenge the prevailing view of global DN-dATN antagonism and suggest retinotopic coding as an organizing principle for cross-network communication.

      Strengths:

      The central finding that what looks like network-level independence between DN and dATN decomposes into structured, bivalent interactions organized by voxel-level visual field preferences is a compelling demonstration that macro-scale network descriptions can hide meaningful substructure. The logic of the analysis is clean: pRF properties are estimated from retinotopic mapping data and then used to predict resting-state coupling in completely independent scanning sessions. This cross-session, cross-modality design rules out many circularity concerns.

      The use of individualized multi-session hierarchical Bayesian parcellation (Kong et al.) to define DN and dATN boundaries within each subject is the right methodological choice for this question. Network boundaries in posterior cortex, where DN and dATN interdigitate most closely, vary considerably across individuals, and group-average approaches would introduce exactly the kind of misassignment that would most confound the result.

      The matched-vs-random pRF analysis is well-controlled. The authors demonstrate that cortical distance between matched and randomly matched dATN pRFs does not differ, effectively ruling out spatial proximity on the cortical surface as a confound. tSNR controls further show that signal quality differences do not drive the effect.

      The event-triggered analysis (Figure 3) is creative and adds genuine value. Showing that retinotopically-specific coupling persists during DN-initiated activity transients not only dATN-initiated ones is the key piece of evidence for the claim that the code is intrinsic to the DN rather than passively inherited through bottom-up visual drive.

      The result is observed consistently across all individual participants, which provides strong evidence for the robustness of the qualitative pattern despite the small sample size inherent to densely sampled designs.

      Comments on revised version:

      I'm content with the additional analyses and alterations to the writing that the authors have performed. I'm convinced that this work will spawn a very productive thread in the literature.

    3. Reviewer #2 (Public review):

      Summary:

      Using a public dataset of retinotopic mapping and resting-state data, the authors find that the default mode network has voxels that respond (positively or negatively) to visual stimulation at specific retinotopic positions, and that resting-state activity in these voxels is correlated with activity in more traditional sensory voxels with the same visual-location preference. The retinotopic specificity is bidirectional, such that high activity in default mode voxels drives activity only in voxels with matching receptive fields in sensory cortex, and vice versa. These findings are at odds with traditional views of the default mode network as having abstract (non-retinotopic) representations and competing (rather than cooperating) with external sensory representations.

      Strengths:

      This study continues an intriguing line of research about how default mode regions interact with sensory cortex. Demonstrating that there are structured interactions between these regions at rest, and that these interactions are in fact organized according to retinotopic location (as opposed to traditional views of representational format in the default mode network), provides a new framework for thinking about large-scale internal and external brain networks. The authors make use of a well-powered public dataset that allows for precise estimates of pRFs and individual-specific resting-state networks and develop a number of interesting analyses that characterize the relationships between DN and dATN voxels. The findings are exciting and could have a major impact on future studies in cognitive neuroimaging.

      The authors mention that these findings could shed light on internal/external interactions such as "anticipatory saccades or memory-guided attention," which is true, though I would argue that constructing DN representations of external stimuli is in fact even more fundamental than these specific cases (e.g. see Barnett and Bellana, 2025, "Situation models and the default mode network"). The "highways" identified in this study could play a vital role in real-world perceptual processes that are constantly translating external input into internal mental models.

      Weaknesses:

      (1) The criterion used for defining voxels as retinotopic seems very liberal. The authors show that only 5% of voxels have R^2>0.14 in a null analysis and therefore define voxels with R^2>0.14 as retinotopic. Although all the networks in Fig 1C show voxel distributions that differ from the null, the number of false positives above R^2>0.14 seems problematic, especially for the DN positive pRFs (red distribution) and to a lesser extent the DN negative pRFs (blue distribution). From visual inspection of the plot, the false discovery rate (fraction of voxels labeled as retinotopic that are false positives) looks like it would be greater than 50% for the DN positive pRFs. The authors do show that the positive pRF voxels have above-chance consistency across runs and also show in a supplementary analysis (Fig S5) that applying a stricter R^2 criterion yields similar results. These help to mitigate this concern, providing evidence that there are true positive voxels in this set which are driving the effects.

      (2) The claim that "voxel-level visual response profiles shape DN-dATN coupling during spontaneous resting-state activity" is well-supported for specific sub-groups of DN voxels, though it is unclear whether the overall DN-dATN correlation at rest is primarily driven by the pRF-tuned voxels investigated in this study.

      (3) The event-triggered analysis is effective at testing the bidirectional relationship between DN and dATN, with high activity in either network triggering a response in the other network. However, it would be helpful to show more validation that these "events" are meaningful windows of time to study, and that 13 TRs a typical length of time that activity is elevated during one of these events.

      (4) The framing of this paper relative to the authors past work, such as Steel et al. 2024 ("A retinotopic code structures the interaction between perception and memory systems") could be improved. The primary novelty here is that this paper examines resting-state data and individually defined whole-brain networks, showing that there are widespread spontaneous interactions between broad internal and external networks, but this distinction is not made explicit in the Introduction.

    4. Reviewer #3 (Public review):

      Summary:

      This paper addresses an important question (relationship between DN and dATN, and the role of retinotopic coding) and uses a set of novel analyses.

      Strengths:

      Important question, novel analytical approaches (pRF-informed functional connectivity analysis).

      Weaknesses:

      Some of the analyses are not described with sufficient clarity, especially the final analysis related to Fig. 3.

      Comments on revised version.

      Related to my previous comment 3), the removal of the labels "bottom-up" and "top-down" in the final analysis is a big improvement. However, I still don't fully understand how the 10 most aligned pRFs and the 10 most anti-matched pRFs are selected. The methods section on this has some ambiguity: "the 10 with the smallest Euclidean distance in RF center (x,y)". Does this mean that these are the pRFs closest to fovea? If not, what is the Euclidean distance referring to? Likewise, I don't understand how the anti-matched voxels are selected. This makes the interpretation of Fig 3 difficult.

      My previous comment about baseline activation was to compare the matched voxels with randomly selected voxels, instead of with anti-matched voxels. The authors responded that there was a technical difficulty with this.

    5. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This paper leverages 7T fMRI data from the Natural Scenes Dataset to investigate whether retinotopic coding, the position-selective organization of visual response structures, spontaneous resting-state interactions between the Default Network (DN) and the Dorsal Attention Network (dATN). Using individualized network parcellations and population receptive field (pRF) modeling, the authors show that DN voxels can be split into two subpopulations based on their response to visual stimulation: those with position-specific positive BOLD responses (+pRFs) and those with position-specific negative BOLD responses (-pRFs). Critically, these subpopulations relate differently to the dATN during rest: -pRFs are anticorrelated with the dATN, +pRFs are positively correlated, and non-retinotopic DN voxels show no coupling. The anticorrelation (and positive correlation) is enhanced when DN and dATN voxels share visual field preferences. An eventtriggered analysis suggests that retinotopic coding shapes both "top-down" (DNinitiated) and "bottom-up" (dATN-initiated) spontaneous activity transients, supporting the claim that the retinotopic scaffold is intrinsic to the DN. These findings challenge the prevailing view of global DN-dATN antagonism and suggest retinotopic coding as an organizing principle for cross-network communication.

      Strengths:

      The central finding that what looks like network-level independence between DN and dATN decomposes into structured, bivalent interactions organized by voxellevel visual field preferences is a compelling demonstration that macro-scale network descriptions can hide meaningful substructure. The logic of the analysis is clean: pRF properties are estimated from retinotopic mapping data and then used to predict resting-state coupling in completely independent scanning sessions. This cross-session, cross-modality design rules out many circularity concerns.

      The use of individualized multi-session hierarchical Bayesian parcellation (Kong et al.) to define DN and dATN boundaries within each subject is the right methodological choice for this question. Network boundaries in posterior cortex, where DN and dATN interdigitate most closely, vary considerably across individuals, and group-average approaches would introduce exactly the kind of misassignment that would most confound the result.

      The matched-vs-random pRF analysis is well-controlled. The authors demonstrate that cortical distance between matched and randomly-matched dATN pRFs does not differ, effectively ruling out spatial proximity on the cortical surface as a confound. tSNR controls further show that signal quality differences do not drive the effect.

      The event-triggered analysis (Figure 3) is creative and adds genuine value. Showing that retinotopically-specific coupling persists during DN-initiated activity transients, not only dATN-initiated ones, is the key piece of evidence for the claim that the code is intrinsic to the DN rather than passively inherited through bottom-up visual drive.

      The result is observed consistently across all individual participants, which provides strong evidence for the robustness of the qualitative pattern despite the small sample size inherent to densely-sampled designs.

      Weaknesses

      (1) The nature of negative pRFs requires more scrutiny

      The entire interpretive framework depends on treating negative pRFs in the DN as genuine position-selective neural responses (suppression). However, negative BOLD signals are well known to arise from non-neural sources, specifically, vascular stealing (where activation in nearby tissue diverts blood from adjacent voxels) and macrovascular draining vein effects that produce spatially displaced signal inversions. These concerns are amplified at 7T, where T2*-weighted GEEPI carries substantial macrovascular weighting. The DN and dATN interdigitate extensively in the posterior cortex, often within millimeters. A negative pRF in a DN voxel adjacent to a positive dATN voxel could, in principle, reflect the hemodynamic shadow of its neighbor rather than an independent neural response.

      The spatial dispersion control (matched vs. random pRFs have similar cortical distribution) is valuable but addresses long-range confounds, not local hemodynamic crosstalk. The reliability of sign and center position across runs is reassuring but does not exclude a vascular origin, as vascular architecture is itself stable across sessions. I would encourage the authors to test whether the matched-vs-random effect survives exclusion of voxels near large pial vessels (identifiable from T2* contrast or the venograms available in the NSD). These analyses would not be dispositive, but they would meaningfully strengthen the neural interpretation.

      The reviewer raises an important concern about the interpretation of negative pRFs in the DN, namely that spatially specific negative BOLD responses could, in principle, reflect local vascular effects rather than genuine position-selective suppression. The reviewer suggests excluding voxels near large vessels to address this issue.

      Based on the reviewer’s suggestion, we repeated the pRF matching analysis excluding any voxels within 3mm of a major vein, as identified using the time-of-flight (TOF) MR venography included in the NSD. This analysis therefore tests whether the retinotopically specific DN–dATN coupling persists after removing voxels most likely to be affected by vascular signal.

      Excluding these voxels did not impact our results: we found preferential coupling according to response valence and center position, with stronger correlation between matched +DN and +dATN voxels (t(6) = 6.054, p < 0.001), and a more pronounced negative correlation between matched -DN and -dATN voxels (t(6) = -5.0448, p < 0.01). We have added these results to the supplemental figures (Fig. S7), and also added to the text (Pg. 8). Together with the run-wise reliability of pRF sign and position, and the persistence of the matched-versus-random effect after vessel exclusion, this analysis supports the interpretation that negative DN pRFs reflect structured, spatially specific responses rather than a vascular artifact.

      “Finally, to rule out any possible influences from vascular stealing (i.e. the shunting of blood into active tissue from nearby regions), we repeated the matching analysis after excluding any voxels within a 3mm radius of a major vessel (Fig. S7; see Methods). Both matching effects remained after excluding vascularly susceptible voxels (+DN x dATN: t(6) = 6.054; p < 0.001; -DN x dATN: t(6) = -5.0448; p < 0.01).”

      (2) Amount of retinotopic mapping data and choice of pRF pipeline

      The NSD includes 6 runs of retinotopic mapping (~5 minutes each; 3 baraperture, 3 wedge/ring). The authors use only the 3 bar-aperture runs (~15 minutes total per subject) and fit their own pRFs using AFNI's 3dNLfim procedure, rather than using the pRF estimates provided as part of the NSD release (which were fitted using the analyzePRF toolbox with all 6 runs).

      Fifteen minutes of bar data is quite limited for reliable voxel-wise pRF estimation, especially in regions far from the early visual cortex, where signal-to-noise is inherently lower. Standard recommendations for robust pRF mapping in higherorder regions generally suggest substantially more data. The variance-explained threshold is close to the noise floor by design, meaning that a non-trivial number of the "retinotopic" DN voxels may be poorly estimated. Given that the core analyses depend on both the sign and the center position of these pRFs, the limited data is a significant concern.

      The authors do not explain why they chose to re-fit pRFs rather than use the NSD-provided estimates. If the motivation was methodological (e.g., the NSD pRF pipeline does not readily yield signed amplitude, or the bar-only fits were judged more appropriate for detecting negative responses), this should be made explicit. If the NSD-provided pRFs can reproduce the key findings, this would substantially increase confidence in the results. If they cannot, that divergence itself would be important to understand. I would ask the authors to address this choice and, if feasible, to report whether the core results replicate using the NSDprovided pRF estimates and/or whether using all 6 runs of retinotopy data changes the findings.

      The reviewer raises two related concerns: first, that the amount of retinotopic mapping data available in the NSD may be limited for estimating voxel-wise pRFs in higher-order cortical regions; and second, that we re-fit the pRF model using AFNI rather than relying on the pRF estimates provided with the NSD release. We appreciate the opportunity to clarify both points. We agree with the reviewer that more travelling bar data would be preferable and would likely yield more robust model fits, particularly in higher-order regions with lower SNR. This is a limitation of our paper that we now acknowledge in the discussion section. However, we do not think that more data would fundamentally change the pattern of our results for the following reasons.

      First, we implemented a novel data-driven approach to derive a threshold for thresholding significant pRF fits (a noise floor). Importantly, our noise floor estimation yields a conservative threshold (R<sup>2</sup> > 0.14), which is greater than both our previous work characterizing cortical pRFs (Steel et al. 2024: R<sup>2</sup> > 0.08) and other work exploring visual responses in the default network (Klink et al. 2021: R<sup>2</sup> > 0.05; no threshold: Szinte and Knapen 2020; Knapen 2021).

      Second, the key pRF features used in our analyses – response sign and centre position – were reliable across retinotopic mapping runs. This reliability is important because our central matching analysis depends on voxel-wise estimates of both response valence and visual-field position.

      Third, the matching analysis asks whether pRF parameters estimated from the retinotopic mapping task predict functional coupling measured during independent resting-state scans. Noisy or unstable pRF estimates should weaken this relationship, because they would degrade the accuracy of voxel-wise matching. Thus, parameter instability would be expected to obscure retinotopically specific coupling rather than systematically produce the observed matched-versus-random effects.

      To the reviewer’s question about our decision to re-fit the pRF model using AFNI, the reviewer is correct that this was motivated by the requirements of our analysis: we re-fit the pRF estimates using AFNI because it allows for both positive and negative signed amplitudes. The pRF model fits provided with the NSD do not allow bivalent amplitude estimates. We have made this decision clearer in the text, reproduced below (Pg. 4-5; Pg. 14-15).

      “We chose to re-fit the data using a simple Gaussian approach as implemented in AFNI to allow for both positive and negative signed amplitudes.”

      “The limited amount of pRF mapping task data included in the NSD posed a challenge for establishing reliable visual response estimates. Here, we addressed this issue by developing a novel thresholding method to establish robust voxel-wise model fits. Among voxels that passed this empirical threshold, we observed a significant correlation in voxel-wise estimates of centre position and visual response amplitude. In addition, our pRF matching results were based on the relationship between the voxel-wise estimates of centre position and response amplitude with resting-state fMRI – a completely independent measure. Crucially, noisy estimates of pRF parameters would obscure this relationship and make our results less likely. Therefore, despite the relatively limited pRF mapping data available, unstable pRF estimates are unlikely to drive our results.”

      (3) pRF model adequacy for the Default Network

      The isotropic Gaussian pRF model was developed for and validated in early and mid-level visual cortex, where it captures the dominant spatial selectivity of neuronal populations. In DN voxels where the model explains comparatively little variance, it is less clear that the model is capturing the right quantity.

      Specifically, the negative pRFs could conceivably be described by a model with a dominant suppressive surround (e.g., a difference-of-Gaussians model), in which what appears as a "negative pRF" in the standard model is actually the surround component of a center-surround mechanism whose center is poorly resolved. This distinction matters: a genuine inverted code (negative center response) implies a qualitatively different computation than inherited surround suppression from nearby visual cortex.

      The authors should consider discussing why the standard model is sufficient for the questions asked, or ideally, testing whether the sign distinction survives under alternative pRF model specifications.

      We appreciate the reviewer’s comment about the limitations of a single gaussian pRF model. We chose the single gaussian model as a direct extension of prior work from our lab and others (Steel et al., 2024, Klink et al., 2022, Szinte and Knapen, 2021). We agree that a negative response in this model could, in principle, reflect a more complex spatial profile, such as a dominant suppressive surround. However, adjudicating among alternative pRF models would require more retinotopic mapping data than are available in the NSD, particularly for higher-order cortex. Thus, we feel that it is outside the scope of the current work. We now address this limitation in our discussion (Pg. 15).

      “Relatedly, here we used a single gaussian model, consistent with prior work on negative visual responses in memory systems (31, 33, 34). However, other models of visual receptive fields might offer further insight into the DN’s visual responsiveness, such as double gaussian models of surround suppression (65) or compressive summation (66). Future studies might directly compare different visual models to further refine the computations underpinning visual responses in the DN.”

      (4) Interpreting resting-state transients as top-down vs. bottom-up The event-triggered analysis labels high-amplitude DN pRF activations as "topdown events" and dATN activations as "bottom-up events." This is a reasonable inference given experience-sampling studies showing that rest involves alternation between internal and external attention, but it remains an inference. Without concurrent experience sampling, eye-tracking, or physiological monitoring, we cannot establish that a spontaneous DN transient reflects memory retrieval or internally-directed thought rather than a global arousal fluctuation. Similarly, dATN transients during rest could reflect covert shifts of spatial attention to remembered or imagined locations rather than bottom-up processing per se. I would ask the authors to soften this framing or to discuss what additional data would be needed to validate the top-down/bottom-up attribution.

      The reviewer raises an important concern about the strong interpretation of elevated BOLD activity detected in the DN and dATN as top-down and bottom-up events. We agree that the limitations of fMRI in our current data prevent these strong claims about the origin of these signals. We have therefore softened this framing throughout the manuscript, and we now refer to these events as DN-driven and dATN-driven. We think that this more directly describes the analysis: events were defined by transient high-amplitude activity in DN or dATN pRFs, respectively.  

      (5) The "retinotopic code" vs. "visual field bias" distinction The paper uses the language of a "retinotopic code" throughout and correctly distinguishes this from a "retinotopic map," noting that DN voxels do not form a continuous topographic representation on the cortical surface. This distinction deserves greater emphasis. In vision science, retinotopic maps carry computational significance through their topographic continuity and relationship to cortical wiring. A distributed collection of voxels with coarse visual field preferences but no cortical topography is a fundamentally different organizational feature. Recent reviews have drawn an explicit distinction between retinotopic maps and visual field biases (Groen, Dekker, Knapen & Silson, TiCS 2022), and the present findings may be more accurately characterized as the latter. Perhaps the authors think that the distinction is merely a signal-to-noise distinction, in which case I would invite them to clearly speak to this interpretation. In any case, this is not a criticism of the findings themselves, but clarity on this point would prevent conflation of two different organizational principles and would help position the work for both the vision and network neuroscience communities.

      The reviewer raises a valuable point about the distinction between a retinotopic code, a retinotopic map, and a visual field bias, and we are happy to add discussion of this topic to our manuscript.

      Our results show that the DN does not exhibit a continuous retinotopic map in the sense used in early visual cortex. Rather, our results suggest a distributed voxel-level code for visual-field position: individual DN voxels show reliable spatial preferences, and these preferences predict retinotopically specific functional coupling with dATN voxels. This voxel-level organization is analogous to other distributed spatial codes, such as head-direction coding in retrosplenial cortex, where spatial variables are represented by population activity without requiring a topographic map on the cortical surface. This differs from a coarse visual-field bias, including preferential responses to the contralateral visual field, although we do also observe such biases. We have added text unpacking this important distinction to the Discussion (Pg. 15-16):

      “Prior work has emphasized the visual response bias in regions where voxel-wise retinotopic responses lack a map-like organization(35); overall, the DN does exhibit this kind of bias. However, our results show that the voxel-scale activity underpinning this bias reflects the latent connectivity of those voxels. Thus, we adopt the term “retinotopic coding”, because this voxel-scale coding scheme exists without a map-like organization on the cortical surface. For example, rodent and bat head direction cells are not laid out in a literal ring, but the population code of these neurons forms a ring manifold(68, 69).”

      Reviewer #2 (Public review):

      Summary:

      Using a public dataset of retinotopic mapping and resting-state data, the authors find that the default mode network has voxels that respond (positively or negatively) to visual stimulation at specific retinotopic positions, and that restingstate activity in these voxels is correlated with activity in more traditional sensory voxels with the same visual-location preference. The retinotopic specificity is bidirectional, such that high activity in default mode voxels drives activity only in voxels with matching receptive fields in sensory cortex, and vice versa. These findings are at odds with traditional views of the default mode network as having abstract (non-retinotopic) representations and competing (rather than cooperating) with external sensory representations.

      Strengths:

      This study continues an intriguing line of research about how default mode regions interact with the sensory cortex. Demonstrating that there are structured interactions between these regions at rest, and that these interactions are in fact organized according to retinotopic location (as opposed to traditional views of representational format in the default mode network), provides a new framework for thinking about large-scale internal and external brain networks. The authors make use of a well-powered public dataset that allows for precise estimates of pRFs and individual-specific resting-state networks, and develop a number of interesting analyses that characterize the relationships between DN and dATN voxels. The findings are exciting and could have a major impact on future studies in cognitive neuroimaging.

      The authors mention that these findings could shed light on internal/external interactions such as "anticipatory saccades or memory-guided attention," which is true, though I would argue that constructing DN representations of external stimuli is in fact even more fundamental than these specific cases (e.g., see Barnett and Bellana, 2025, "Situation models and the default mode network"). The "highways" identified in this study could play a vital role in real-world perceptual processes that are constantly translating external input into internal mental models.

      Weaknesses:

      (1) The criterion used for defining voxels as retinotopic seems very liberal. The authors show that only 5% of voxels have R^2>0.14 in a null analysis, and therefore define voxels with R^2>0.14 as retinotopic. Although all the networks in 1C show voxel distributions that differ from the null, the number of false positives above R^2>0.14 seems problematic, especially for the DN positive pRFs (red distribution) and to a lesser extent the DN negative pRFs (blue distribution). From visual inspection of the plot, the false discovery rate (fraction of voxels labeled as retinotopic that are false positives) looks like it would be greater than 50% for the DN-positive pRFs. The authors do show that the positive pRF voxels have abovechance consistency across runs, again providing evidence that there are true positive voxels in this set, but perhaps a stricter criterion (such as having consistent negative fits across runs) would provide more targeted identification of the DN voxels with true retinotopic sensitivity.

      We thank the reviewer for giving us the opportunity to discuss this important decision. We agree with the reviewer that a stricter R<sup>2</sup> criterion could result in more targeted pRF identification. Motivated by the reviewer’s suggestion, we repeated the cross-region pRF matching analysis across multiple R2 thresholds.

      The retinotopic matching effects were not dependent on the original threshold. In fact, we found that the pRF matching effects are enhanced as the R<sup>2</sup> value increases (Fig. S5). This pattern suggests that any false-positive voxels admitted near the original threshold would dilute, rather than drive, the observed matched-versus-random effects. We have added text to the results highlighting this finding (Pg. 7):

      “In contrast, DN voxels that responded positively to visual stimulation (DN positive pRFs, +pRFs) had a positive correlation with the dATN (mean correlation = 0.284±0.152, t(6) = 4.96, p = 0.0025), while DN voxels with systematic negative responses to visual stimulation (DN negative pRFs, -pRFs) were anti-correlated with the dATN (mean correlation = -0.21±0.149, t(6) = -3.75, p = 0.0094). This relationship was further strengthened by adopting more conservative R<sup>2</sup> thresholds up to 0.30 despite the overall number of included voxels decreasing, suggesting that this effect is not driven by false-positive voxels at the edge of our threshold criteria (Fig. S5).”

      (2) The claim that "opponency at rest between the DN and dATN appears to be driven by the subset of DN voxels with negative retinotopic tuning" is not well supported. The fraction of DN voxels with negative pRFs is small: 9.42% of DN voxels have pRFs, and 58.77% are negative, so about 6% of DN voxels have negative pRFs. The fact that any DN voxels have negative pRFs is notable, but the authors do not provide evidence that these 6% are driving the overall behavior of the DN. They do show (e.g., in Figure 2B) that negative and positive pRFs have opposing influences, but the overall correlation with dATN does not look similar to the negative pRF connectivity. I'm also unsure whether "opponency" is a reasonable description for two networks that are "independent (i.e., not correlated)" in this analysis.

      The reviewer raises an important point about whether negative DN pRFs should be described as driving the overall DN–dATN relationship. We agree that this language was too strong. Negative pRFs constitute a small subset of DN voxels, and our analyses show that this subset has a distinct pattern of functional coupling with the dATN, not that it explains the global relationship between the DN and dATN as a whole.

      We have therefore revised the manuscript to avoid implying that negative DN pRFs drive overall DN–dATN opponency. Instead, we now frame these voxels as an important retinotopically tuned subpopulation nested within broader network dynamics. Specifically, our results show that visually responsive DN voxels are not homogeneous: positive and negative DN pRFs show opposing patterns of coupling with dATN pRFs, and these interactions are strengthened when voxels share visual-field preferences. This suggests that a small but structured subset of DN voxels may provide a route for retinotopically specific communication between internally and externally oriented networks, without implying that this subset determines the mean activity pattern of the entire DN:

      “Spontaneous DN and dATN activity during rest is uncorrelated at the network level. However, voxel-scale functional coupling across networks is shaped by the latent visual field preferences of individual voxels in each network, as measured during independent retinotopic mapping.” Abstract (Pg. 2)

      “This result shows that voxel-level visual response profiles shape DN-dATN coupling during spontaneous resting-state activity. Specifically, the DN and dATN activation is independent during rest. However, at the voxel-level, specific sub-groups of DN voxels have distinct coupling patterns with the dATN that depends on the valence of voxels’ visual responses. DN and dATN voxels with positive visual responses show a positive relationship during rest, and a notable subset of DN voxels with negative visual responses display the canonical opponency with dATN voxels. This suggests that retinotopic coding may be a mechanism that enables visual information to be exchanged between these large-scale brain systems. Specifically, opponency at rest between the DN and dATN appears to be driven by the subset of DN voxels with negative retinotopic tuning.” Results (Pg. 7)

      These findings offer a multi-scale account of neural communication, in which interactions among sub-populations of voxels with shared tuning preferences are nested within macro-scale network dynamics. Nesting multiple neural codes might enable ongoing computations within a larger brain system (e.g., attending to internal mental states within the DN during memory recall), while simultaneously allowing for the sharing of fine-grained representations across brain systems (34). Discussion (Pg. 14)

      (3) The event-triggered analysis is effective at testing the bidirectional relationship between DN and dATN, with high activity in either network triggering a response in the other network. However, it would be helpful to show more validation that these "events" are meaningful windows of time to study. First, is 13 TRs a typical length of time that activity is elevated during one of these events? Second, the top-down and bottom-up terminology is perhaps too loaded and not well-justified; if the negative pRFs in the DN reflect a meaningful coding system, then couldn't low (rather than high) activity indicate a top-down event?

      We thank the reviewer for these helpful suggestions. To the best of our knowledge, there is not currently a widely agreed-upon time window for performing event-based fMRI analyses. We chose a 13 TR time window to balance between sufficiently capturing BOLD signal related to the chosen event while also minimizing influence from other signal fluctuations, based on the procedure adopted in Gordon et al. (Nature, 2023) and Mitra et al. (J. Neuro Phys, 2014), which considered temporal relationships among brain regions over comparable timescales. In our analysis, this window considered 6 TRs (9.6s) on either side of the detected event, which we felt comfortably captures the peak BOLD signal that would result from an impulse at the event time, and responses that may reflect upstream activity leading into it.

      The reviewer has raised an additional comment about the terms “top-down” and “bottom-up.” These concerns were shared by Reviewer 1. Based on these comments, we have adopted the terms “DN-driven” and “dATN-driven”, which we think aligns more closely with our analysis approach.

      (4) The framing of this paper relative to the authors' past work, such as Steel et al. 2024 ("A retinotopic code structures the interaction between perception and memory systems"), could be improved. The existence of negative pRFs in the DN and a functional relationship between these pRFs and the sensory pRFs have already been described in prior work. My understanding of the primary novelty here is that this paper examines resting-state data, showing that there are widespread spontaneous interactions between broad internal and external networks, but this distinction is not made explicit in the Introduction.

      We appreciate the opportunity to clarify the novel aspects of our paper. The reviewer correctly identifies the extent of prior work, which identified -pRFs in regions of the canonical default network (Szinte and Knapen, 2021; Klink et al., 2022) and characterized the local interactions between adjacent perceptual and mnemonic regions (Steel et al., 2024). Our current work builds upon these findings in two key ways.

      First, we explore the effect across individually-defined whole brain networks. While the DN and dATN are often adjacent, these networks are spatially discontinuous and are comprised of distinct sub-regions (e.g. in prefrontal cortex). Whether retinotopic patterning of activity would persist in distributed networks could not have been extrapolated from our prior work. We think that finding will be of broad interest to the community studying perception and memory systems, because it offers a mechanistic account of how information is read in/out of memory.

      Second, here we considered whether spontaneous activity across networks would be structured by a retinotopic code. Our previous work characterized activity during tasks that depended on visual information: either scene perception or mental imagery. While the prior work was an important first step, it left open the possibility that retinotopic coding may only be relevant in visual tasks. By demonstrating that the retinotopic coding structures voxel-specific coactivation during rest, which entails no overt visual demands, we provide evidence that retinotopic features are a general, mode-agnostic code between regions.

      (5) The definition of the default mode (DN) in this study aligns with past research, but the definition of the dorsal attention network (dATN) seems at odds with standard terminology. For example, the authors cite Fox et al. 2006, which depicts the dATN as including regions such as IPS, FEF, SMA, and MT+. Here, however, the "dATN" seems to be primarily lateral and ventral visual cortex (e.g., Figure S5). The exact location of these sensory pRFs is not critical to the authors' claims, but this labeling seems incorrect, and the motivation for defining/selecting the sensory network in this way is not described.

      We thank the reviewer for this insightful comment and their careful consideration of our network definition.

      Our method for network identification, and the topography of the resulting networks, are broadly consistent with more recent conceptualizations of the DN and dATN (e.g. Du et al. 2024, Gordon et al. 2017, Braga and Buckner 2017). Relatedly, because we defined brain networks based on the unique connectivity patterns of each individual participant, we expect them to differ from previous group-level network descriptions. The increased resolution of the 7T data in the NSD may also result in greater departure from prior definitions compared to previous work done at 3T.

      Further study into dATN differences between group-level 3T, individualized 3T, and individualized 7T networks could be a valuable future direction, but this is outside the scope of this work.

      Reviewer #3 (Public review):

      Summary:

      This paper addresses an important question (the relationship between DN and dATN, and the role of retinotopic coding) and uses a set of novel analyses.

      Strengths:

      Important question, novel analytical approaches (pRF-informed functional connectivity analysis).

      Weaknesses:

      Some of the key claims are not fully supported by the data presented. There is also a concern about over-interpretation of the results. Key issues:

      (1)  The authors claim that retinotopic coding scaffolds the interaction between DMN and dATN. However, retinotopically tuned voxels account for a mere 9% of DMN voxels. So this appears to be a major overstatement. For instance, the statement that "these findings would position retinotopy as a unifying framework for brain-wide information processing" is not justified given the presented data.

      We appreciate the reviewer’s concern about the framing of our conclusion, which was shared by reviewers 1 and 2. In response to these comments, we have revised our paper to more accurately reflect the observed data. Specifically, we focus on the specific sub-populations of voxels within the DN and dATN that show retinotopic responses, and we have removed references to explaining the overall pattern of activity across networks.

      (2) Given that positive pRF voxels in DMN positively correlate with dATN voxels and negative pRF voxels in DMN negatively correlate with dATN voxels, there is a concern that these results could be contributed to by imprecise brain network parcellations. E.g., could some of the positive pRF voxels in DMN be erroneously assigned to DMN and actually belong to one of the other task-positive networks? There is insufficient validation of network parcellation to put this worry to rest, especially since it depends on ICA, which has a degree of arbitrariness built in.

      We thank the reviewer for the opportunity to clarify our method for network definition.

      Precision functional mapping is a growing field with many methods for defining personalized functional networks for each individual. Because the NSD resting-state data is relatively high resolution, we chose an approach designed to improve the stability of voxel-wise network assignment: Multi-Session Hierarchical Bayesian Modeling approach (Kong et al. 2019; Du et al. 2024). This approach enhances stability of network assignment by including a group-based prior and accounting for both within- and across-subject variability. This approach is more stable than ICA, and, because this approach leverages a prior, there is less concern about arbitrary or idiosyncratic network definitions.

      However, it is still common for network assignments to have lower confidence around the borders between networks. Yet, we also do not think border misassignment is likely to explain the present results for two reasons: first, while DN and dATN nodes are sometimes adjacent, there are many regions where they are spatially distant, such as the IPS for dATN and the lateral temporal lobe for DN. Second, the DN pRFs do not appear to cluster selectively along DN–dATN borders, suggesting that they are not simply misassigned dATN voxels.(Fig. S3) Therefore, we think voxels on the edge of these networks are unlikely to drive the effects observed here (see Fig. S3).

      (3) The claim that retinotopic coding is intrinsic to the DN network is not supported by rigorous analysis and results. The analysis here has many arbitrary factors, including: the threshold of the 99th percentile of resting-state distribution; the designation of DN as "top-down" and dATN as "bottom-up"; the definition of "anti-matched" voxels instead of using randomly selected voxels; and the statistics being paired between matched and anti-matched voxels instead of using comparisons to baseline. Overall, I do not think that the result supports the conclusion that retinotopic coding in DN is intrinsic instead of being bottomup-driven, given the very high threshold (99%) used and the fact that many other networks could also send bottom-up input to DN. Furthermore, the idea that bottom-up inputs only occur when the dATN (or any other RSN)'s spontaneous BOLD activity is above a certain threshold is a huge and unvalidated assumption.

      The reviewer raises several interesting concerns about decisions in our event-detection analysis. Here, we clarify the rationale for several analytic choices:

      (1) The 99th-percentile threshold was chosen to identify sparse, high-amplitude events while minimizing contamination from smaller ongoing fluctuations.

      (2) The other reviewers also noted a concern with the top-down/bottom-up terminology. We have revised these terms to DN-driven and dATN-driven, which we think reflect our approach more accurately.

      (3) We used anti-matched rather than randomly selected voxels because the full event-by-voxel randomization procedure was computationally intractable at the network level.

      (4) We did not understand the reviewer’s contention about activation baseline, but we would welcome clarification.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Minor points

      (1) The reliability analysis (Figure 1D) notes that dATN negative pRF amplitude was not reliable above chance in 2 of 7 participants. This could be discussed more prominently as it suggests that negative pRFs may not be stable features in all networks, which tempers the generality of the sign distinction as a fundamental organizational property.

      We thank the reviewer for raising this point of clarification. It is true that 2/7 participants did not show reliable negative pRFs in the dATN. However, the majority of participants show stable negative pRFs, and even in these 2 participants, the negative result does not indicate that negative pRFs would not be stable in those individuals with additional data. 

      Based on the reviewer’s comment, we have added emphasis to this point, but we do not feel that this warrants greater discussion in the paper. 

      Both positive and negative pRF amplitude was reliable in the DN for all subjects. In the dATN, positive amplitude pRFs were reliable in all participants, and negative amplitude pRFs (which constituted a small proportion of the overall pRFs in this network) were reliable in 5/7 participants. For the remainder of the paper, we only consider positive pRFs in the dATN. Importantly, pRF center position was highly reproducible across runs of pRF data in the dATN and DN in all subjects (Fig. 1D).  Pg. 5

      (2) The paper would benefit from situating the findings more explicitly within the cortical gradient framework (Margulies et al., 2016), which predicts that DN regions have maximally abstract, transmodal codes. The present findings complicate this view productively and deserve to be "situated" within that ongoing debate.

      We agree that the gradient framework is interesting, and we have added discussion of Margulies to our paper. (Pg. 16)

      Relatedly, the DN is considered a transmodal hub for cortical processing, where disparate sensory and motor processes converge (59, 75) The DN’s position at the cortical apex implies connections with and influence over unimodal cortical areas. However, the mechanism for liking unimodal and transmodal networks had been unknown. Prior work posited that sensory coding in transmodal areas might serve this function (31, 35), and our data provide direct empirical support for this account: specific visually-responsive voxels provide an input/output interface linking perceptual and memory systems. This complements work delineating specific affective and effective subregions within the DN that link the DN to other brain areas (76). Thus, while the DN may be “distant from input” (28), these data suggest that it is not disengaged from sensory processing.

      (3) It would be informative to know whether the *proportion* of negative vs. positive pRFs differs between DN-A and DN-B, given their distinct functional roles.

      Despite the functional specialization of DN-A and DN-B, and the slightly higher mean proportion of negative pRFs in DN-A (61% vs 56%), we found no statistically significant difference in the proportion of negative pRFs across the two networks (t(6) = 0.888, p = 0.409).

      (4) Low N is inherent to the densely-sampled NSD design, and the within-subject consistency is a strength. Nevertheless, with 6 degrees of freedom, the precision of specific quantitative estimates (e.g., that 58.77% of DN pRFs are negative) is uncertain, and the authors should be cautious about the generalizability of these point estimates.

      The reviewer raises a concern about the inclusion of specific levels of decimal place in our statistical reporting. We do not think that this is a major issue with the paper, but we are willing to change if the reviewer feels strongly.

      Reviewer #2 (Recommendations for the authors):

      (1) Figure 1C could use an explicit legend (I believe it is following the color convention from the bar plots in 1F?). Also, for consistency, it would be helpful to make all the colormaps in 1F correspond to the bars (i.e., change the dATN colormap to go white->green).

      We thank the reviewer for this suggestion, and we have added explicit labels to Fig. 1C

      (2) Providing a scatter plot, in which each dot is a voxel and the x and y axes are the pRF amplitude estimates in different runs, could help provide evidence that there are voxels with pRFs that have consistently negative amplitudes across runs. This would also go beyond the binary consistency analysis in Figure 1D to show that the magnitudes of the amplitude estimates are also consistent.

      We thank the reviewer for this suggestion. We feel that the binary consistency conveys sufficient information. Because the analysis is done using pairwise correlation, how the scatter plot would reflect the three-way consistency is not clear. 

      (3) For understanding how the overall correlation between DN and dATN could be driven by voxel populations with opposing effects (e.g., Figure 2B), it would be useful to show how the +pRF and -pRF voxels compare to other voxels within the DN. For example, are these the voxels with the strongest negative and positive correlations with dATN, or are there many other DN voxels (among the 90% that do not have pRFs) that also have similarly-strong dATN correlations?

      The reviewer offers a very interesting suggestion. Based on the reviewer’s suggestion, we have refocused our paper on the particular subpopulations of +/- pRFs in the DN, rather than on an explanation for the overall pattern of correlation between the DN and dATN. Because our revised framing focuses on the properties of these retinotopically defined voxel populations, rather than on explaining whole-network DN–dATN coupling, we have not added this additional analysis. We have revised the relevant text to avoid implying that these pRF subpopulations drive the overall network-level relationship.

      (4) Initially, the baseline comparison pRFs for the matched pRFs are labeled "random" pRFs, which seems misleading; these are closer to "mismatched"/"anti-matched" pRFs since they are selected from the 1/3 that are farthest away. Then the comparison switched to using the anti-matched pRFs that are the 10 very farthest away, though I didn't understand the rationale that "the large number of pRFs made the random matching procedure impractical" - in what way is the number of pRFs larger in this analysis? Having a more consistent baseline (e.g., just using the 10 anti-matched pRFs the whole time) would be easier to interpret.

      We thank the reviewer for this suggestion. We have compared the results between the randomly-sampled bottom ⅓ matched versus the 10 worst matched, and the pattern of results is identical (the effect is strongest in the 10 worst matched). Therefore, we include the bottom ⅓ matched in the main text as a more conservative test of this effect. We are happy to include this as a supplemental figure if the reviewer feels it is essential. 

      (5) In the past, I have only seen the terminology "bootstrapped" to refer to sampling with replacement from the data sample, producing samples/statistics that are centered on the observed data. Here (lines 704-708), the sampling is coming from the null distribution of randomly-chosen voxels, and therefore the term "bootstrapped" would not apply (and could just be replaced with "null").

      We have revised this terminology in the manuscript.

      Reviewer #3 (Recommendations for the authors):

      (1) Abstract and Discussion should be significantly toned down. E.g., the claim that "These findings challenge the prevailing view of global DN-dATN antagonism" is not really supported by the data provided. The claim that "retinotopic coding underpins the dynamic coordination of perception and thought" is also unsupported by the presented data.

      We have revised the manuscript in light of this comment.

      (2) Line 233-235: The null statistical result cannot support the claim reached here. Correlation analysis or Bayesian statistics should be used.

      We have revised the manuscript in light of this comment.

      (3) Line 250-254: Comparison to baseline should be used, in addition to comparing matched and random voxels.

      We agree that baseline comparisons can be useful in event-triggered analyses. However, for the pRF-matching analysis discussed here, the critical question is whether shared visual-field preference influences resting-state functional coupling between DN and dATN voxels. For this question, we believe that the appropriate baseline is the coupling observed for pRFs that do not share visual-field preferences. We therefore compare retinotopically matched pRFs to randomly matched pRFs drawn from the same networks. 

      (4) Line 271: "not" is missing.

      We have revised the manuscript in light of this comment.

    1. eLife Assessment

      This important study combines chromatin accessibility and genomic DNA sequence conservation data from low-coverage genome sequencing of related species (without assembly), for the in-silico identification of cis-regulatory elements in large genomes. The approach and results are compelling and well supported by the experimental validations. The work will be of interest to researchers working in the field of gene regulation and evolution, particularly because the methodology proposed can be applied to a large variety of experimental organisms.

    2. Reviewer #1 (Public review):

      Summary:

      Forbes et al. developed an integrated approach to identify cis-regulatory elements (CREs) in the large (3.6 Gbp) genome of the crustacean Parhyale hawaiensis, addressing the challenge of pinpointing these regions among large regions of non-coding sequences. They combined ATAC-seq chromatin accessibility profiling (both bulk and single-nucleus) across embryonic and adult tissues with low-coverage genome sequencing of three congeneric species (P. aquilina, P. darvishi, P. plumicornis). Without assembling congener genomes, they mapped reads with low stringency to the P. hawaiensis reference, identifying about 55k conserved islands that overlap ATAC peaks more than expected by chance. This dual filter was used to select CRE candidates for transgenic reporter validation, yielding 6 functional elements (out of 11 tested) driving ubiquitous, neuronal, or muscle-specific expression, a major advance for non-model systems with large genomes.

      Strengths:

      Forbes et al. generated high-quality ATAC data across multiple scales. Using bulk ATAC-seq (from whole embryos, developing and adult legs) they identified tens of thousands of open chromatin peaks across the assembled P. hawaiensis large genome. Moreover, using single-nucleus ATAC-seq from adult legs, they could resolve differentially accessible chromatin profiles across more than 15 cell types previously identified by scRNA-seq, enabling cell-type-specific candidate selection.

      Furthermore, their innovative low-coverage comparative genomics method mapped 0.46-6.4% of congener reads to P. hawaiensis without genome assembly, revealing hundreds of thousands of conserved non-coding islands, including about 55k showing conservation in all four species, far exceeding random expectation.

      Using the developed approach, the authors could validate 6 (out of 11 candidates) reporter constructs, driving robust ubiquitous and tissue-specific expression, succeeding where prior promoter-only screening failed and providing immediately useful genetic tools for the Parhyale community.

      Weaknesses:

      The primary limitation is that functional CRE testing was performed only in P. hawaiensis. While the conservation maps provide a valuable resource for comparative analyses, functional validation in congener species was not performed, so the extent to which the identified CREs or the prioritization strategy can be functionally generalized across related species remains to be established.

      The approach did not successfully identify developmental CREs among the candidates tested. None of the candidates selected using the combined ATAC-seq and conservation filtering drove reporter expression matching the expected endogenous patterns. The authors appropriately discuss possible technical and biological explanations.

      Overall Assessment:

      Forbes et al. fully succeed with their integrated approach to (1) generate an ATAC-seq atlas plus functional CRE discovery and (2) innovative low-coverage sequencing for conservation mapping in the large 3.6 Gbp genome of Parhyale hawaiensis. Their combination of ATAC-seq chromatin accessibility profiling (bulk and single-nucleus) across embryonic and adult tissues with low-coverage genome sequencing of three congeneric species (P. aquilina, P. darvishi, P. plumicornis), without congener genome assembly, drastically shrank the CRE search space. Using this approach, the authors could validate six out of 11 candidate transgenic reporters (ubiquitous, neuronal, and muscle-specific) where prior promoter-only screening failed.

      The low-coverage mapping innovation cuts cost and labour while snATAC-seq provides cell-type resolution, making these resources valuable for building new genetic and imaging tools in Parhyale.

      This compelling method also has the potential to enable labs with limited resources to identify and characterize regulatory elements in more non-model organisms, advancing our understanding of their evolution while establishing a scalable pipeline for large-genome systems.

      Comments on revised version.

      The authors have adequately addressed all my previous comments. I have no further specific suggestions or requests.

    3. Reviewer #2 (Public review):

      The manuscript by Forbes, Skafida, Karapidaki et al. concerns the in-silico identification of cis-regulatory elements (CREs) in large genomes using chromatin accessibility (ATAC-seq) and sequence conservation (genomic DNA sequencing) data. They exemplify this method by applying it to identify novel CREs in Parhyale hawaiensis, which they validated using reporter constructs.

      The results are convincing and are well supported by the data and validations. Identified CREs are valuable for researchers interested in the regulation of the expression of genes they control.

      The methodology on the whole is also valid, as suggested by the results and previous publications on various taxa. Sequence conservation, as stated by the authors, was long used as a method to identify regions of non-coding DNA with functional and evolutionary constraints. The same applies to ATAC-seq data, which has also been used as a proxy for functional regions in different animals such as sea urchins and amphioxus. The methodology proposed is likely to be successfully used by researchers working on a variety of experimental organisms.

      The authors do not use existing genome assemblies and use short-read sequencing to identify conserved regions, and while it is not conceptually novel, such an approach is becoming more and more viable and useful considering the recent advances in next generation sequencing technology and the decrease in price of short-read sequencing.

      The authors have addressed and discussed the limitations and weaknesses of the approach as well as explicitly indicated the advantages.

      All in all, the authors provide a valid method to strengthen CRE identification via sequence conservation without the need of multiple complete close species genome assemblies, making it a compelling option for non-model organism research.

    4. Reviewer #3 (Public review):

      Summary:

      Forbes et al. present a new approach for identifying cis-regulatory elements in large genomes. Using Parhyale hawaiensis, a crustacean with a large genome (~3.6 Gb, comparable in size to the human genome), the authors show that current methods for identifying cis-regulatory elements, effective in smaller genomes, are markedly inefficient in organisms with large genomes. To address this limitation, they combine bulk ATAC-seq and single-cell (sc) ATAC-seq to identify chromatin regions that are either ubiquitously accessible or specifically accessible in particular cell types. They further integrate comparative genomics across multiple Parhyale species (P. hawaiensis, P. aquilina, and P. darvishi), selected at appropriate phylogenetic distances (20-95 million years divergence), to pinpoint conserved open chromatin regions likely under functional constraint.

      Using this strategy, the authors predict a set of ubiquitous and cell-type-specific cis-regulatory elements. Importantly, they validate these predictions using rigorous transgenic reporter assays, convincingly demonstrating that their approach can successfully identify functional regulatory elements where previous methods had failed.

      Strengths:

      The approach introduced by Forbes et al. is conceptually straightforward, efficient, and readily transferable to other organisms. The validation experiments show not only that a substantial proportion of the predicted elements are functional, but also that the method is capable of identifying both ubiquitous and cell-type-specific regulatory elements. Given that the identification of regulatory regions remains a major bottleneck in understanding the molecular mechanisms underlying processes of development and regeneration, this work has the potential to make a significant impact in developmental and regeneration biology, particularly for studies involving non-model organisms with large genomes.

      An additional strength is the demonstration that only the genome of the focal species requires high-quality sequencing and assembly. In contrast, species used solely for comparative analysis can be sequenced at low coverage without assembly, substantially reducing costs and increasing the accessibility of the approach.

      Weaknesses:

      While the method is effective in identifying regulatory elements that are active ubiquitously or in differentiated cell types, it failed in detecting elements associated with developmentally regulated genes. This may be due to trivial reasons, such as very low level of expression of the selected genes. However, as acknowledged by the authors, it may also indicate inherent challenges in identifying regulatory elements associated with developmentally dynamic gene regulation, compared to those associated with genes expressed in differentiated cell types.

      A second limitation, also acknowledged by the authors, is the absence of chromatin conformation capture data, which would help link distal regulatory elements to their target genes. This limitation may be particularly relevant for developmentally regulated genes, where long-range regulatory interactions may be critical.

      Addressing these limitations will be an important direction for future work. Nonetheless, the approach as presented in this manuscript represents a key contribution that sets the stage for further methodological advances in the identification of cis-regulatory elements in large genomes.

      Comments on revised version.

      I am fully satisfied with the current version of the manuscript.

    5. Author response:

      The following is the authors’ response to the original reviews.

      Reviewer #1 (Public review):

      Summary:

      Forbes et al. developed an integrated approach to identify cis-regulatory elements (CREs) in the large (3.6 Gbp) genome of the crustacean Parhyale hawaiensis, addressing the challenge of pinpointing these regions among large regions of non-coding sequences. They combined ATAC-seq chromatin accessibility profiling (both bulk and single-nucleus) across embryonic and adult tissues with low-coverage genome sequencing of three congeneric species (P. aquilina, P. darvishi, P. plumicornis). Without assembling congener genomes, they mapped reads with low stringency to the P. hawaiensis reference, identifying about 55k conserved islands that overlap ATAC peaks more than expected by chance. This dual filter was used to select CRE candidates for transgenic reporter validation, yielding 6 functional elements (out of 11 tested) driving ubiquitous, neuronal, or muscle-specific expression, a major advance for non-model systems with large genomes.

      Strengths:

      Forbes et al. generated high-quality ATAC data across multiple scales. Using bulk ATAC-seq (from whole embryos, developing and adult legs), they identified tens of thousands of open chromatin peaks across the assembled P. hawaiensis large genome. Moreover, using single-nucleus ATAC-seq from adult legs, they could resolve differentially accessible chromatin profiles across over 15 cell types previously identified by scRNA-seq, enabling cell-type-specific candidate selection.

      Furthermore, their innovative low-coverage comparative genomics method mapped 0.46-6.4% of congener reads to P. hawaiensis without genome assembly, revealing hundreds of thousands of conserved non-coding islands, including about 55k showing conservation in all four species, far exceeding random expectation.

      Using the developed approach, the authors could validate 6 (out of 11 candidates) reporter constructs, driving robust ubiquitous and tissue-specific expression, succeeding where prior promoter-only screening failed and providing immediately useful genetic tools for the Parhyale community.

      Weaknesses:

      The primary limitation is that functional CRE testing was performed only in P. hawaiensis. While conservation maps are valuable resources, the manuscript lacks functional validation in congener species, limiting claims about broad applicability across related genomes/species.

      The approach also failed to validate developmental CREs. None of the candidates from combined ATAC and conservation filtering drove reporter expression matching endogenous patterns. The authors appropriately hypothesize technical limits (low expression) or biological factors (long-range enhancers, shadow enhancers).

      Overall Assessment:

      Forbes et al. fully succeed with their integrated approach to (1) generate an ATAC-seq atlas plus functional CRE discovery and (2) innovative low-coverage sequencing for conservation mapping in the large 3.6 Gbp genome of Parhyale hawaiensis. Their combination of ATAC-seq chromatin accessibility profiling (bulk and single-nucleus) across embryonic and adult tissues with low-coverage genome sequencing of three congeneric species (P. aquilina, P. darvishi, P. plumicornis), without congener genome assembly, drastically shrank the CRE search space. Using this approach, the authors could validate six out of 11 candidate transgenic reporters (ubiquitous, neuronal, and muscle-specific), where prior promoter-only screening failed.

      The low-coverage mapping innovation cuts cost and labour while snATAC-seq provides cell-type resolution, making these resources valuable for building new genetic and imaging tools in Parhyale.

      This compelling method also has the potential to enable labs with limited resources to identify and characterize regulatory elements in more non-model organisms, advancing our understanding of their evolution while establishing a scalable pipeline for large-genome systems.

      We thank the reviewer for their comments and valuable feedback.

      Reviewer #1 (Recommendations for the authors):

      (1) Standardize terminology and introduce acronyms properly:

      I suggest standardizing technique names throughout the manuscript (e.g., ATAC-seq rather than ATACseq, RNA-seq rather than RNAseq, ChIP-seq rather than ChipSeq). Please also introduce technical terms with their full name at first use, such as 'Assay for Transposase-Accessible Chromatin using sequencing (ATAC-seq)' rather than just the acronym.

      Corrected.

      (2) Correct author name:

      On page 15, "Lo brutto" appears to be misspelt. Please review and fix throughout the text and the corresponding reference in the bibliography if this is the case.

      Corrected.

      (3) Revise the title to reflect dual contributions:

      The paper delivers two key advances: (a) ATAC-seq atlas plus functional CRE discovery in P. hawaiensis, and (b) innovative low-coverage sequencing for conservation mapping across congenerics. Current title highlights only the first. Consider modifying the title to better reflect both.

      The title reflects our overall objective, without highlighting one of the two approaches (advances) in particular. We would like to keep this concise title and invite readers to read about the two approaches in the abstract.

      (4) Emphasize the combined power of the approach in the Discussion and Conclusions:

      A significant innovation of this manuscript is integrating low-coverage comparative genomics with ATAC-seq to prioritize functional CREs in a large non-model genome. The abstract highlights this well, but the Discussion and Conclusions could better emphasize the power of this pipeline over ATAC-seq alone. The authors could also add 2-3 sentences quantifying cost savings versus traditional assemblies and reiterating how this complements ATAC-seq for efficient CRE prioritization in non-model species.

      We have added the following text in the Discussion to describe complementary contributions of ATAC-seq and sequence conservation to CRE discovery: "Previous efforts to identify cis-regulatory elements in Parhyale relied on reporter constructs carrying a few kb of sequences upstream of selected target genes, an approach that has worked well in animals and plants with relatively small genomes. As presented earlier, however, this approach was often unsuccessful in Parhyale: from tens of reporters tested, only four robust native drivers had so far been identified (refs). The present work adds 5 new drivers to that collection, including ones with ubiquitous, neuron- and muscle-specific activities. This result comes from combining information on genome-wide chromatin accessibility and evolutionary conservation profiles.

      We cannot at this point distinguish the relative contributions of chromatin profiling and sequence conservation to CRE discovery, because these sources of information were not tested separately. At minimum, we can state that (1) ATAC-seq profiles serve to identify robustly the promoters of candidate genes that are active in a particular cellular context (cell type and stage) and (2) coupling this information with sequence conservation narrows down candidate promoters and distant CREs by a factor of 4 to 10, since only a fraction of ATAC-seq peaks show sequence conservation (Figure 3D). This represents a great improvement in our ability to select candidates to test by transgenesis, the most labour-intensive step in the process."

      Further, we have added this text to explain the advantages and cost savings of our low coverage sequencing strategy: "Our strategy of mapping regions of sequence conservation by direct mapping of short sequence reads across species is much more accessible than conventional strategies that rely on genome assembly. The latter require much higher sequence coverage (> 50x) from multiple libraries, long-read sequencing or other scaffolding methods, and complex bioinformatic pipelines to assemble large genomes. Moreover, these approaches are often compromised by high levels of polymorphism found in natural populations. We estimate that our approach is 5- to 10-fold cheaper than assembly-based methods, even excluding labor costs."

      (5) Improve figure readability:

      The authors could improve figure readability by introducing a schematic representation of the specimens and more references in Figures 4, 5 and Supplementary Figure 5. A schematic representation would help non-experts in Parhyale understand what they are looking at. Some figures might also benefit from improved color contrast (e.g., Figure 3 has very similar orange/red colors; black dots on a dark grey background are hard to distinguish).

      We have added additional labels and explanations in the legends of Figures 4, 5 and Supplementary Figure 5, which we think will make the images more intelligible to the readers. In Figure 3 we modified the colouring in panels B and C to improve contrast.

      (6) Quantify reporter validation efficiencies:

      The authors should add a summary table/plot (e.g. n surviving, n fluorescent) or label the figures (n of specimens showing that pattern/n of specimens that do not show the pattern) with exact numbers to explicitly illustrate the observation. For example, Supplementary Table 3 contains excellent data for the putative developmental CREs tested, but the main text lacks equivalent quantification for successful reporters.

      This information is already provided in Table 3.

      (7) Discuss the rapid evolution of developmental CREs:

      The failure to validate developmental CREs using conserved candidates may also reflect the rapid evolution and turnover of developmental enhancers, which can erode detectable sequence conservation over these phylogenetic distances. As a result, functionally relevant elements may have been excluded during candidate selection. It may be worth discussing this possibility alongside the proposed long-range and shadow enhancers hypothesis.

      We added the phrase “or the rapid evolution of these enhancers leading to low sequence conservation" in the relevant part of the Discussion.

      Reviewer #2 (Public review):

      The manuscript by Forbes, Skafida, Karapidaki et al. concerns the in silico identification of cis-regulatory elements (CREs) in large genomes using chromatin accessibility (ATAC-seq) and sequence conservation (genomic DNA sequencing) data. They exemplify this method by applying it to identify novel CREs in Parhyale hawaiensis, which they validated using reporter constructs.

      The results are convincing and are well supported by the data and validations. Identified CREs are valuable for researchers interested in the regulation of the expression of genes they control.

      The methodology on the whole is also valid, as suggested by the results and previous publications on various taxa. Sequence conservation, as stated by the authors, was long used as a method to identify regions of non-coding DNA with functional and evolutionary constraints. The same applies to ATAC-seq data, which has also been used as a proxy for functional regions in different animals such as sea urchins and amphioxus. The methodology proposed is likely to be successfully used by researchers working on a variety of experimental organisms.

      The authors do not use existing genome assemblies and use short-read sequencing to identify conserved regions, and while it is not conceptually novel, such an approach is becoming more and more viable and useful considering the recent advances in next-generation sequencing technology and the decrease in price of short-read sequencing.

      We thank the reviewer for their comments and valuable feedback.

      Two major weaknesses are:

      (1) The novelty of the approach and its advantages should be more explicitly stated.

      (2) The authors do not discuss in depth the strength of using a combination of two methods rather than either of the two, especially considering that previously known CREs do not overlap with conserved sequences.

      We have added two paragraphs at the start of the Discussion to address the reviewer's comments 1 and 2 more explicitly (see response to reviewer 1, comment 4).

      Previously known CREs do include some conserved sequences, see Suppl. Figure 7.

      Reviewer #2 (Recommendations for the authors):

      In addition to addressing the two above-mentioned weaknesses, the authors should address the following minor issues:

      (1) It is difficult to refer to particular regions of text without line numbers.

      Spelling of ChIPseq is inconsistent in the Introduction.

      Spelling corrected. (Sorry for not including line numbering, we'll try to remember next time.)

      (2) "6.4% of reads from P. aquilina, 4.1% of reads from P. darvishi, and 0.46 % of reads from P. plumicornis could be mapped unambiguously to the P. hawaiensis genome" seems quite low for closely related species. Do the authors expect such low rates?

      Neutral nucleotide substitution rates in multicellular animals are in the order of 1 per site per 100 million years (e.g. https://pubmed.ncbi.nlm.nih.gov/12949132/) or a little lower (e.g. https://pubmed.ncbi.nlm.nih.gov/11792858/, https://pubmed.ncbi.nlm.nih.gov/34049492/). With the evolutionary times separating P. hawaiensis from P. aquilina/darvishi and P. plumicornis estimated at roughly 50 and 180 million years (2x25 and 2x90 million years, respectively), we expect a large fraction of neutrally evolving nucleotides in these genomes to have changed. We performed the read mapping using bowtie2, which requires a ~20 nt long perfect match with the reference sequence. We were therefore not surprised to obtain such low rates of read mapping. In fact, these low mapping rates (long divergence times) are important for islands of sequence conservation to stand out.

      (3) "Of these, 37% are found in introns, 54% in intergenic regions, and 1% overlap with promoters (TSS), marking regions that evolve at a lower rate than surrounding non-coding sequences". The authors explain in the Methods why they omit exons, but in the Results and Discussion, it is not stated. In addition, discussing the conservation with exons would be helpful, and the % in exons should be compared to non-coding regions.

      We have added "Of these, 8% are found in exons, likely reflecting conservation in protein-coding sequences".

      (4) "Two of the 7 reporters we tested, named neuro5 and neuro6," if I understood correctly, neuro5 and neuro6 are CREs, however, they are named quite ambiguously, and their names can be mistaken for gene names.

      Indeed, neuro5 and neur6 are the names of the CRE reporters. We have now added the names of the corresponding genes ("carrying CREs associated with the genes αTub and Cdk5α, respectively"). The gene names are also given in Table 3.

      (5) Why was single-end sequencing done for E24?

      We now explain this in the Methods: "Sequencing was carried out on an Illumina NextSeq 500 sequencer; we carried out single-end 76 bp sequencing for the first sample we generated (E24), and then switched to paired-end 76 bp sequencing for the other samples, because this leads to more specific read mapping."

      (6) Syntax related to in-line references should be double checked as the following sentences are broken by parentheses, e.g., "updated in (Almazán et al. 2022))".

      Corrected.

      (7) Could the authors discuss the P. aquilina genome size, which was estimated to be 3-times less than P. hawaiensis? Considering that in their phylogeny these two species are closest, it is quite surprising that they have such differing genome sizes. Do you expect it to be true? If yes, what could be the reason?

      As we explain in the manuscript, our estimates of genome size were obtained by dividing the total number of nucleotides sequenced by the estimated genome coverage, for each species. This method could overestimate genome sizes if there was a significant fraction of contaminating DNA in the preps, or a high degree of sequence variation that would prevent efficient mapping to BUSCO genes (both would underestimate genome coverage), but we find no evidence of this when we estimate the genome size of P. hawaiensis (see manuscript). We used the same method to estimate genome size in all four Parhyale species and have no reason to think that the method would be biased in one species and not in others. We therefore think that we have comparable estimates of genome size for the 4 species and the size difference is real.

      Variations in genome size can be driven by changes in the fraction of repetitive sequences found in a genome. We therefore checked the proportion of repetitive elements in each Parhyale genome using dnaPipeTE (https://github.com/clemgoub/dnaPipeTE). Based on this method (which likely underestimates the repetitive genome content) we find that the genomes of P. hawaiensis, P. aquiline, P. darvishi and P. plumicornis contain 31%, 22%, 18% and 39% of repetitive sequences, respectively. These figures do not fully account for the differences in genome size (particularly since P. darvishi appears to have even fewer repetitive sequences than P. aquiline). We therefore hesitate to add this very preliminary analysis to the manuscript.

      Of note, such rapid change in genome size is not unprecedented: in fruit flies genome size can vary more than 3-fold in species that have diverged over about 30 million years (https://elifesciences.org/articles/66405).

      (8) Wording "and found a genome coverage of 5.8x, corresponding to a genome size of 3.0 Gbp instead of 3.6 Gbp" is confusing and unclear as to what the authors exactly did here.

      We modified the sentence: "As a control, we followed the same procedure for P. hawaiensis, for which genome size is known (ref), and found a genome size of 3.0 Gbp instead of 3.6 Gbp (with a genome coverage of 5.8x)."

      (9) In the figures and supplementary figures, the genome browser screenshots should also include tracks of macs2 called peaks (those in narrowPeak format).

      Each ATACseq and sequence conservation track has its own set of peaks; we think that adding more tracks would overcrowd the figures. All the tracks (including called peaks) are provided as genome-browser-readable files in Supplementary Data files 1-3, so readers should be able to explore the data and reconstruct the figure panels without much effort.

      Reviewer #3 (Public review):

      Summary:

      Forbes et al. present a new approach for identifying cis-regulatory elements in large genomes. Using Parhyale hawaiensis, a crustacean with a large genome (~3.6 Gb, comparable in size to the human genome), the authors show that current methods for identifying cis-regulatory elements, effective in smaller genomes, are markedly inefficient in organisms with large genomes. To address this limitation, they combine bulk ATAC-seq and single-cell (sc) ATAC-seq to identify chromatin regions that are either ubiquitously accessible or specifically accessible in particular cell types. They further integrate comparative genomics across multiple Parhyale species (P. hawaiensis, P. aquilina, and P. darvishi), selected at appropriate phylogenetic distances (20-95 million years divergence), to pinpoint conserved open chromatin regions likely under functional constraint.

      Using this strategy, the authors predict a set of ubiquitous and cell-type-specific cis-regulatory elements. Importantly, they validate these predictions using rigorous transgenic reporter assays, convincingly demonstrating that their approach can successfully identify functional regulatory elements where previous methods had failed.

      Strengths:

      The approach introduced by Forbes et al. is conceptually straightforward, efficient, and readily transferable to other organisms. The validation experiments show not only that a substantial proportion of the predicted elements are functional, but also that the method is capable of identifying both ubiquitous and cell-type-specific regulatory elements. Given that the identification of regulatory regions remains a major bottleneck in understanding the molecular mechanisms underlying processes of development and regeneration, this work has the potential to make a significant impact in developmental and regeneration biology, particularly for studies involving non-model organisms with large genomes.

      An additional strength is the demonstration that only the genome of the focal species requires high-quality sequencing and assembly. In contrast, species used solely for comparative analysis can be sequenced at low coverage without assembly, substantially reducing costs and increasing the accessibility of the approach.

      Weaknesses:

      While the method is effective in identifying regulatory elements that are active ubiquitously or in differentiated cell types, it failed in detecting elements associated with developmentally regulated genes. This may be due to trivial reasons, such as a very low level of expression of the selected genes. However, as acknowledged by the authors, it may also indicate inherent challenges in identifying regulatory elements associated with developmentally dynamic gene regulation, compared to those associated with genes expressed in differentiated cell types.

      A second limitation, also acknowledged by the authors, is the absence of chromatin conformation capture data, which would help link distal regulatory elements to their target genes. This limitation may be particularly relevant for developmentally regulated genes, where long-range regulatory interactions may be critical.

      Addressing these limitations will be an important direction for future work. Nonetheless, the approach as presented in this manuscript represents a key contribution that sets the stage for further methodological advances in the identification of cis-regulatory elements in large genomes.

      Reviewer #3 (Recommendations for the authors):

      I have no specific comment for the authors. While in my opinion the study has two limitations (as described in the public review), these are clearly acknowledged and properly discussed in the manuscript.

      The manuscript is extremely well written. It has been a great pleasure to read it. Excellent job!

      Thank you!

    1. eLife Assessment

      The report by Liu and colleagues reports on an analysis of environmental adaptation across diverse lineages of the grass Phragmites australis differing by their level of ploidy. These results represent important findings for understanding the environmental adaptation of species complexes with mixed ploidy and the analysis reports solid evidence that lineages with distinct levels of ploidy occupy different climate niches. The use of regional survey in tandem with common garden experiment represents a convincing approach to suggest a correlation between ploidy and climate adaptation. This manuscript will be of interest to a broad community of ecological genomicists interested in how structural variation in gene dosage potentially affects the pattern of adaptation.

    2. Reviewer #1 (Public review):

      Summary:

      The article is testing the relative advantages of plant lineages with differing ploidy and admixture across environmental gradients. The results show that intraspecific variation in ploidy and admixture between lineages impacts plant traits that may enable persistence and range expansion.

      Strengths:

      Suitable marker panel size and strong results that include attempts to analyse mixed ploidy level data which is a challenge.

      Weaknesses:

      The sample sizes of the common garden experiments are very low making it difficult to draw robust conclusions.

    3. Author response:

      The following is the authors’ response to the original reviews.

      Overview of Revisions

      We thank the editors and reviewers for their constructive and insightful comments, which have substantially improved the manuscript. We have carefully addressed every point raised. The major revisions include:

      (1) Methods 2.1: completely reorganized to clarify the allotetraploid genome structure of Phragmites australis, the rationale for single-chromosome-anchored microsatellite markers, the maximum distinguishable alleles per ploidy level, and the conservative Ploidies(mydata) <- 4 setting in polysat.

      (2) Methods 2.3 / Results 3.3 / Discussion 4.4: clarified common garden sample sizes, added Cohen's d effect sizes, and acknowledged the correlational nature of the lineage-level comparisons.

      (3) Introduction: added a new paragraph on the eco-evolutionary significance of gene flow in mixed-ploidy systems, and another paragraph emphasizing the novelty of integrating SDMs with physiological and common garden experiments.

      (4) Discussion 4.1 and 4.4: reframed all "polyploidy-driven" language to "polyploidy-associated", explicitly acknowledging that ploidy is confounded with genetic background and was not experimentally manipulated.

      Public Reviews:

      Reviewer #1 (Public review):

      (R1-P1) Inadequate explanation of allele dosage for ploidy levels

      Inadequate explanation of allele dosage for ploidy levels, some of which do not match the allele counts expected for genome copy number.

      We appreciate this comment and have substantially revised Methods 2.1. The key clarifications are:

      (A) Marker specificity: All 42 microsatellite markers were aligned to the P. australis reference genome, and each marker mapped to a single unique chromosome (Table S2). This confirms that each marker amplifies a locus specific to one subgenome only. Therefore, in tetraploids each marker detects at most two alleles (the two homologous copies of that chromosome from one subgenome), while in octoploids (autopolyploid derivatives with four copies of the same chromosome) each marker detects up to four alleles.

      (B) Conservative ploidy setting: We set Ploidies(mydata) <- 4 for all samples in polysat because the exact ploidy of many samples could not be confidently assigned a priori. This uniform treatment is conservative: for actual tetraploids, the two unobserved "copies" are scored as null; for actual octoploids, all four detected alleles are accommodated.

      (C) Dosage estimation: Allele dosage was estimated from high-coverage sequencing read counts (mean >5,000× per locus per sample) using the SSRSeq V1.1 pipeline (Cui et al., 2022), which applies stutter correction, amplification bias correction, and ploidy-optimized dosage calling, not inferred from allele presence/absence alone.

      Methods 2.1, second paragraph onward

      Phragmites australis has a base allotetraploid genome. All 42 microsatellite markers used in this study were aligned to the P. australis reference genome, and each marker mapped to a single unique chromosome (Table S2), confirming that each marker amplifies from one subgenome only. Therefore, in tetraploids each marker detects at most two alleles (the two homologous copies of that chromosome from the target subgenome) while the homologous region from the other subgenome is not amplified (Saltonstall, 2003). In Asia, the prevalent octoploids are most likely autopolyploid derivatives of tetraploids, carrying four homologous copies of the same chromosome and thus capable of up to four distinguishable alleles per locus (Liu et al., 2022; Wang et al., 2024). Hexaploid individuals are rare and occur primarily in contact zones, likely originating from inter‑lineage hybridization (Wang et al., 2024).

      In practice, the 42 selected markers very rarely produced more than four alleles in any single individual (Table S2), consistent with a ploidy ceiling of octoploid. Because the exact ploidy of many samples could not be confidently assigned a priori (ploidy was inferred from a combination of chloroplast haplotype, geographic origin, and flow cytometry from prior studies; Lambertini et al., 2020; Liu et al., 2022), we consistently set Ploidies(mydata) <- 4 in the polysat R package (Clark & Jasieniuk, 2011), treating every individual as having four homologous copies. This uniform treatment is conservative: for an actual tetraploid (two copies per locus), the two unobserved "copies" are simply scored as null (missing data) in the dosage matrix; for an actual octoploid, all four detected alleles are accommodated. The allele dosage itself was estimated from high-coverage sequencing read counts (mean >5,000× per locus per sample) using the SSRSeq V1.1 pipeline (Cui et al., 2022), not inferred from allele counts alone.”

      (R1-P2) Common garden setup and sample sizes unclear

      The setup and sample sizes of the common garden experiments are very unclear. The numbers implied are extremely low to draw robust conclusions.

      We agree that the original description was insufficiently detailed. We have made three modifications:

      (A) Methods 2.3: clarified that each population contributed one rhizome segment planted in one pot (one biological replicate per population per site), and emphasized that the key inference rests on the direction and consistency of differences across four climatically distinct sites rather than on significance at any single site.

      (B) Results 3.3: added Cohen's d effect sizes (verified from the raw data), showing that where lineage differences are present, they are biologically substantial (d = 1.10–1.43 at three of four sites).

      (C) Discussion 4.4: added a paragraph acknowledging the limited number of populations per lineage and the need for future confirmation with a larger panel.

      Methods 2.3

      “(1) A previously published common garden experiment (Song et al., 2021) conducted in 2017 across Jinan (36.43°N, 117.45°E) and Panjin (41.20°N, 122.02°E), using CN (n = 11) and FEAU (n = 9) lineages, for which we determined the haplotype information of all samples; (2) A new common garden experiment established in 2021 across Qingdao (36.36°N, 120.69°E) and Shanghai (30.20°N, 121.29°E), with CN (n = 9) and FEAU (n = 8) lineages (Table S3). Each rhizome segment (2–3 buds per segment, one segment per population) was transplanted into an individual 20 L pot, yielding one biological replicate per population per site. Although the number of populations per lineage is modest, the key inference rests on the direction and consistency of lineage differences across four climatically distinct sites rather than on the statistical significance at any single site.”

      Results 3.3

      “Similarly, plant height was significantly greater in the FEAU lineage than in the CN lineage in Jinan, but not in the other common gardens (Figure 3C). Effect sizes for total biomass were large in Jinan (Cohen's d = 1.10), Panjin (d = 1.12), and Qingdao (d = 1.43), but negligible in Shanghai (d = 0.37), confirming that the lineage differences, where present, are biologically substantial.”

      Discussion 4.4

      “We also acknowledge that the common garden experiments, while replicated across four climatically distinct sites, involved a limited number of populations per lineage (9–11 CN and 8–9 FEAU), which constrains our ability to fully separate lineage-level effects from population-level variation. The consistent direction of biomass differences across three of four sites, supported by large effect sizes, nonetheless provides robust evidence for a lineage-level performance advantage that merits further confirmation with a larger, more geographically representative panel of populations.”

      (R1-P3) How allele dosage is determined

      Unclear how allele dosage is determined. Given it's so central to many analyses, it would be useful to see how this is done rather than use a citation.

      We agree that a self-contained description is warranted. We have rewritten the relevant paragraph in Methods 2.1 to describe the three core steps of the SSRSeq V1.1 pipeline (Cui et al., 2022): (i) stutter correction based on empirically estimated slip ratios; (ii) amplification bias correction across alleles of different repeat lengths; and (iii) ploidy-adjusted allele dosage calling. The pipeline's source code and full documentation are available at https://github.com/ccoo22/SSRseq_count.

      Methods 2.1

      “Microsatellite genotyping was performed using the SSRSeq V1.1 pipeline (Cui et al., 2022; https://github.com/ccoo22/SSRseq_count). Briefly, the pipeline takes the per-locus per-sample read count table generated from high-throughput sequencing and processes it through three core steps fully described in Cui et al. (2022): (i) stutter correction, which reallocates a fraction of reads from each allele to its adjacent repeat class based on empirically estimated slip ratios; (ii) amplification bias correction, which normalizes read counts across alleles of different repeat lengths using locus-specific bias coefficients; and (iii) allele dosage calling, which selects the maximum number of alleles consistent with the specified ploidy (four in this study) and assigns integer dosages (0–4) by comparing corrected read ratios to a ploidy-adjusted threshold optimized to minimize both allelic dropout and false positives. The final output is a genotype matrix with integer allele dosages for all samples and loci, which was used directly as input to the polysat R package for subsequent population genetic analyses.”

      Reviewer #2 (Public review):

      (R2-P1) Polyploidy has no causal evidence; confounded with genetic background

      First, no data support the claims that polyploidy has any causal effect. The ploidy levels are, in fact, completely confounded with other genetic differences, so it is not possible to eliminate genetic variation, independent of ploidy, as the causative factor. As the authors note, ploidy was not manipulated in the reported experiments. Thus, the focus on polyploidy in the introduction and elsewhere distracts from the novel and informative experiments that were conducted.

      We fully acknowledge this critical limitation and thank the reviewer for this important critique. We have revised the manuscript at four locations to reframe all claims from "polyploidy-driven" to "polyploidy-associated" and to explicitly state that ploidy is confounded with lineage identity and was not experimentally manipulated.

      Abstract (last sentence)

      “These results demonstrate that climate change interacts with intraspecific variation among polyploidy-associated lineages, manifested through differences in thermal tolerance, biomass production, and asymmetric gene flow, to drive potential lineage replacement within a native range”

      Introduction (polyploidy paragraph)

      “The octoploid FEAU lineage is distinguished from its tetraploid relatives not only by ploidy level but also by its distinct evolutionary history, genomic background, and geographic origin. Polyploidy has been shown in other systems to generate genetic novelty, alter gene expression, and enhance physiological stress tolerance (Bureš et al., 2024; Cheng et al., 2021; Kolář et al., 2017; Van de Peer et al., 2017), potentially pre-equipping polyploid lineages to occupy new geographical ranges and endure environmental shifts (Cheng et al., 2021; López-Jurado et al., 2019). The FEAU lineage's superior thermal tolerance and biomass are consistent with such polyploidy-associated effects, although ploidy is correlated with, rather than experimentally separable from, the broader genetic identity of each lineage.”

      Discussion 4.1 (title and opening paragraph)

      “Our findings demonstrate that the octoploid FEAU lineage of P. australis possesses greater heat tolerance and biomass production than the tetraploid CN lineage. Under a high emission scenario (SSP5-8.5), the projected suitable habitat for the FEAU lineage expands by 18.6%, while the CN lineage exhibits a much smaller relative increase. Several non-mutually-exclusive mechanisms could explain these lineage-level differences, including increased gene dosage from whole-genome duplication, divergent selection histories, and/or standing genetic variation in thermal tolerance loci unlinked to ploidy (Bures et al., 2024; Cheng et al., 2021; Van de Peer et al., 2017). Our data cannot fully partition these factors, but the strong association between lineage identity and both physiological performance and projected range dynamics highlights the importance of incorporating intraspecific lineage information into ecological forecasts, regardless of the ultimate causal mechanism.”

      Discussion 4.4 (limitations paragraph)

      “Crucially, ploidy was not experimentally manipulated in this study; it is inherently confounded with the distinct evolutionary history and genomic background of each lineage. While the observed thermal tolerance and biomass differences are consistently associated with the octoploid FEAU lineage, we cannot formally exclude the possibility that these traits are driven by genetic factors independent of ploidy per se. Future studies using experimental approaches that can partition ploidy effects from lineage-specific genetic effects, such as common gardens with synthetic polyploids or transcriptomic analyses comparing gene expression dosage responses, are needed to strengthen causal inference (Wei et al., 2020). Similarly, the common garden results should be interpreted as lineage-associated rather than ploidy-causal performance differences. The potential role of admixture in facilitating the adaptive introgression of heat tolerance alleles also warrants deeper investigation (Suarez-Gonzalez et al., 2018).”

      (R2-P2) SDMs treat lineages as homogeneous entities

      Second, the manuscript indicates that intraspecific variation is critical for the evolutionary potential of a species to respond to environmental change, but intraspecific variation is seldom considered in species distribution models... the manuscript performs species distribution modeling on a small number of sub-specific lineages, essentially treating them as homogeneous "species"—thus the analysis commits the same oversimplification that the manuscript highlights, but does so at a finer evolutionary scale than species.

      We acknowledge this important limitation and agree that it deserves explicit discussion. While disaggregating the species into three major genetic lineages is a step forward from species-as-monolith approaches, within-lineage variation in thermal tolerance, growth, and dispersal capacity is plausible given the broad geographic ranges of the CN and FEAU lineages. We have added a new paragraph in Discussion 4.4 to address this point.

      Discussion 4.4 (new paragraph)

      “We also recognize that our SDM approach, while disaggregating the species into three major genetic lineages, still treats each lineage as a homogeneous entity. This simplification parallels—albeit at a finer scale—the species-as-monolith assumption that we critique in the Introduction. Within-lineage variation in thermal tolerance, growth, and dispersal capacity is plausible, particularly given the broad geographic ranges of the CN and FEAU lineages. By modelling each lineage as a uniform group, our projections may overestimate the precision of range forecasts and underestimate the evolutionary potential of standing variation within lineages (Chardon et al., 2020). Future frameworks that incorporate trait variation at multiple hierarchical levels (population, lineage, ploidy) will be necessary to capture both the adaptive potential and the ecological constraints that shape species' responses to climate change.”

      (R2-P3) Asymmetric introgression and thermal tolerance lack context in Introduction and Discussion

      The title suggests that asymmetric introgression and thermal tolerance are the most important findings of the work. However, the introduction contains no explanation of the potential importance of gene flow (other than to say that asymmetric gene flow was suggested by some preliminary analyses), and the discussion offers only a limited explanation of either the potential mechanisms underlying the asymmetric gene flow or its importance for the long-term evolution of the species.

      We agree that the evolutionary significance of asymmetric gene flow was underdeveloped. We have added two substantial new passages:

      (A) Introduction: a new paragraph explaining the dual role of gene flow in climate adaptation: introgression of adaptive alleles vs. asymmetric introgression as a mechanism of gradual lineage replacement. This paragraph explicitly connects genome dosage differences (octoploid vs. tetraploid) to the natural directionality of backcrossing.

      (B) Discussion 4.2: a new paragraph extending the discussion of asymmetric introgression into its long-term evolutionary consequences, including the potential erosion of the CN lineage's genetic distinctiveness and the risk of losing cold-adapted alleles under future climate volatility.

      Introduction (new paragraph)

      “Gene flow between lineages of differing ploidy can play a dual role in climate adaptation. Introgression may introduce adaptive alleles (e.g., heat tolerance loci) into a recipient lineage, facilitating its persistence under warming (Suarez-Gonzalez et al., 2018). Conversely, if introgression is asymmetric, such that one lineage's genome is disproportionately represented in admixed populations, it can drive a gradual but systematic shift in genetic composition within the contact zone—effectively functioning as a mechanism of lineage replacement without requiring complete competitive exclusion (Bartolić et al., 2024; Zohren et al., 2016). In mixed-ploidy systems, genome dosage differences create a natural directionality in backcrossing: hybrids tend to backcross more frequently with the high-ploidy parent (Bartolić et al., 2024). In the present study, we test whether such a bias exists between the octoploid FEAU and tetraploid CN lineages and examine its consequences for future distribution under climate warming.”

      Discussion 4.2 (new paragraph)

      “From an evolutionary standpoint, asymmetric introgression can erode the genetic distinctiveness of the minority lineage (CN) while enriching the majority lineage (FEAU) with alleles that may have been locally adapted in the CN genomic background. This could reduce the species' overall evolutionary potential, even if the FEAU lineage itself thrives—because cold-adapted alleles from the CN lineage, which may be valuable under future climate volatility (including extreme cold events), risk being diluted or lost (Exposito-Alonso et al., 2022). The directionality of introgression is also not fixed; it could shift if environmental conditions alter hybrid fitness or if the demographic balance between lineages changes. Long-term genomic monitoring of the CN–FEAU contact zone will be essential to determine whether the asymmetric gene flow documented here represents a transient phase or a persistent trajectory toward genomic homogenization.”

      Recommendations for the authors:

      Reviewing Editor Comments:

      (RE-1) Explain how ploidy level is inferred

      We invite the authors to clearly explain how the level of ploidy is being inferred (Reviewer #1).

      We agree that the rationale for ploidy assignment and its relationship to allele counts needed greater clarity. This has been addressed by the comprehensive revision of Methods 2.1 described in response to R1-P1 (Part 1, Reviewer #1 Public Reviews). The revised text now presents a complete step-by-step logical chain: (i) P. australis has an allotetraploid base genome; (ii) all 42 markers map to a single unique chromosome in the reference genome, confirming that each marker amplifies from only one subgenome; (iii) tetraploids therefore show at most two distinguishable alleles per locus, while octoploids (autopolyploid derivatives of tetraploids) show up to four; (iv) because many samples lacked independent ploidy confirmation, we uniformly set Ploidies(mydata) <- 4 in polysat as a conservative treatment that accommodates both tetraploids (two observed copies + two null) and octoploids (four observed copies).

      See the full revised text under R1-P1 (Part 1) above.

      (RE-2) Common garden results are correlational

      We note that the findings of the common garden experiment, although interesting, are mostly correlational (not causative) and rely on a relatively small sample size and confound lineage isolation and adaptive differentiation (both Reviewers).

      We fully acknowledge this limitation. Because all octoploids belong to the FEAU lineage and all tetraploids to CN, ploidy and lineage identity are inherently confounded. This has been addressed by the four-part revision described in response to R2-P1 (Part 1, Reviewer #2 Public Reviews). Specifically:

      The Abstract now frames the findings as "polyploidy-associated" rather than "rooted in polyploidy."

      The Introduction now explicitly states that ploidy is correlated with—but not experimentally separable from—the broader genetic identity of each lineage.

      The Discussion 4.1 title was changed to "Polyploidy-associated thermal tolerance" and the opening paragraph now presents multiple non-mutually-exclusive mechanisms rather than asserting a causal role for polyploidy.

      The Discussion 4.4 now includes an expanded limitations paragraph acknowledging that ploidy was not experimentally manipulated and that common garden results should be interpreted as lineage-associated rather than ploidy-causal.

      In addition, the Methods 2.3 and Discussion 4.4 revisions described in response to R1-P2 (Part 1) address the sample size concern by clarifying the experimental design and adding a dedicated acknowledgement of the limited population replication.

      See the full revised text under R2-P1 and R1-P2 (Part 1) above.

      Reviewer #1 (Recommendations for the authors):

      (R1-R1) "Large morphological traits"

      Line 85. Large morphological traits. Does this mean physically large? Or higher values of some trait.

      We agree the original phrasing was ambiguous. We have replaced "large morphological traits" with explicit trait descriptions.

      Introduction

      “The octoploid FEAU lineage exhibits greater shoot height, larger leaf size, and thicker stems (K. Chen et al., 1993; Guo et al., 2025; Liu et al., 2021b, 2026; Yin et al., 2024), along with stronger salt tolerance and higher thermal tolerance”

      (R1-R2) Why 2 allele copies expected for a tetraploid

      Line 119. Not clear why this allele copy number is expected. A tetraploid can have up to 4 unique alleles (e.g., ABCD), not two.

      This comment arises from the same conceptual gap addressed in R1-P1. The key point is that P. australis is an allotetraploid with two subgenomes, and our 42 markers each map to a single unique chromosome (one subgenome). Therefore, the marker only amplifies the two homologous copies from that subgenome, giving at most two distinguishable alleles. A true autotetraploid would indeed show up to four alleles—but that is not the genomic architecture of P. australis. The revised Methods 2.1 (see R1-P1 in Part 1) now explicitly explains this logic.

      Fully addressed by the Methods 2.1 revision in R1-P1.

      (R1-R3) Theoretical expectation of allele number vs. ploidy

      Line 127-129. As above, this is unclear and not what we expect theoretically. If there is a reasonable number of alleles, there should be a maximum of 4 for tets, 6 for hex, and 8 for octs.

      Same point as R1-R2. The reviewer's expectation (4 for tetraploids, 6 for hexaploids, 8 for octoploids) is correct for autopolyploids with markers that amplify all homologous copies. The discrepancy arises because P. australis is an allotetraploid and our markers are single-chromosome-anchored (each amplifying from only one subgenome). The revised Methods 2.1 now clarifies this distinction explicitly.

      Fully addressed by the Methods 2.1 revision in R1-P1.

      (R1-R4) Typo "makers"

      Line 136. Should be 'markers' not 'makers'

      We have performed a full-text search and corrected all instances of "makers" to "markers" in the manuscript.

      Full-text search and replace.

      (R1-R5) Reason for removing markers with >4 alleles

      Line 142. The reason for the removal of more than 4 alleles is not clear. What about hexaploids and octoploids? They can carry 6 or 8 alleles, respectively.

      We agree the original text did not adequately justify this quality-control step. In our study, the maximum expected distinguishable alleles (given the allotetraploid genome and single-chromosome-anchored markers) is two for tetraploids and four for octoploids. The observation of five or more alleles in multiple individuals is therefore diagnostic of multi-locus amplification (the marker amplifying more than one genomic locus), not of high ploidy. This is a quality-control filter, not a ploidy assignment criterion.

      Methods 2.1

      “During genotyping, eleven markers (including four of the five multi-mapping markers) were removed because more than ten samples exhibited more than four alleles per sample at these loci. Because the maximum number of distinguishable alleles expected under our ploidy model is two (tetraploid) to four (octoploid), the observation of five or more alleles in multiple individuals indicates that these markers amplify more than one genomic locus, rendering them unsuitable for dosage-based genotyping. This filtration is a quality-control step, not a ploidy assignment criterion.”

      (R1-R6) Unclear sample sizes in common garden

      Line 240. Unclear sample sizes. If these are the numbers, they are a very low level of replication expected for a common garden experiment.

      Same point as R1-P2 (Public Review). Please see the full response under R1-P2 in Part 1, where we have (A) clarified the experimental design in Methods 2.3, (B) added Cohen's d effect sizes in Results 3.3, and (C) acknowledged the sample size limitation in Discussion 4.4.

      Fully addressed by the three-part revision in R1-P2.

      Reviewer #2 (Recommendations for the authors):

      (R2-A1) De-emphasize polyploidy

      De-emphasize polyploidy, as it's not manipulated in the study and is entirely confounded with the genotypes of the distinct lineages, and the putative links between polyploidy and heat tolerance are circumstantial and lacking in a clear mechanism.

      We agree fully. This has been addressed comprehensively across four locations in the manuscript (Abstract, Introduction, Discussion 4.1, Discussion 4.4). See the full response under R2-P1 (Part 1) for the revised text at each location.

      Fully addressed by the four-part revision in R2-P1.

      (R2-A2) Emphasize the novelty of combining SDMs with experiments

      Emphasize the novelty of combining SDMs with experiments (or, if I'm not up on the literature and they are more common, explain how they have been used to make new insights).

      We appreciate this suggestion and agree that explicitly stating the novelty of our integrative approach strengthens the manuscript. We have added a new paragraph at the end of the Introduction.

      Introduction (end, before "Here, we integrate population genomics…")

      “Studies that combine species distribution models with physiological or common garden experiments remain surprisingly uncommon (but see López-Jurado et al., 2019). Such integration is essential for transforming correlative SDM projections into mechanistically grounded predictions. In the present study, we adopt this integrative approach: common garden experiments directly test growth performance under controlled conditions, heat-tolerance measurements identify the specific physiological thresholds (T<sub>crit</sub>, T<sub>50</sub>) underlying lineage-specific climate responses, and SDMs project how these experimentally documented differences translate into spatial dynamics under future warming. By linking experimental data with spatial forecasting, we move beyond correlative climate matching toward a trait-based understanding of how intraspecific variation shapes species' future distributions.”

      (R2-A3) Elaborate on the importance of gene flow

      Elaborate on the importance of gene flow and the potential connections between gene flow and evolving species (or lineage) geographical limits.

      This has been addressed by the two new paragraphs described under R2-P3 (Part 1)—one in the Introduction on the dual role of gene flow in climate adaptation (introgression of adaptive alleles vs. asymmetric introgression as a mechanism of lineage replacement), and one in Discussion 4.2 on the long-term evolutionary consequences of asymmetric introgression.

      Fully addressed by the two-part revision in R2-P3.

    1. eLife Assessment

      This study provides valuable insights into the cellular dynamics underlying accelerated tooth regeneration in a vertebrate model. Using single-nucleus RNA sequencing across multiple time points, the authors present a well-structured analysis of cell populations, trajectories, and intercellular signaling events associated with this process. The strength of evidence is solid but only partially supported, as the conclusions are primarily supported by computational inference, without experimental validation of key findings.

    2. Reviewer #1 (Public review):

      Summary:

      The authors used single-nucleus RNA sequencing (snRNA-seq) to investigate accelerated tooth replacement following tooth plucking in cichlid fish. They analyzed four stages of regeneration using elegant and well-designed approaches to characterize cellular trajectories and interactions within the dental epithelium and mesenchyme during the accelerated replacement process. Their analyses identified cell type-specific gene expression profiles and intercellular signaling interactions associated with whole-tooth regeneration.

      Strengths:

      This is a highly interesting and thoughtfully executed study that provides compelling and convincing insights into the mechanisms underlying accelerated tooth regeneration.

      Comments on revised version.

      I noted in my initial review that "the manuscript currently lacks experimental validation of the single-nucleus RNA-seq data." In response, the authors have added a statement indicating that their cell-type annotations and pathway interpretations are supported by extensive prior experimental work in the cichlid tooth model, including histology, in situ hybridization, immunohistochemistry, and pharmacological perturbation of major developmental pathways. They have also acknowledged this limitation in the Study Limitations and Future Directions section, stating that direct experimental validation of the single-nucleus RNA-seq findings will be the focus of future studies.

      The authors have carefully addressed my comments, particularly the Major Points (2), (3), and (4), as well as all of the Minor Points. I appreciate their efforts to further characterize the mesenchymal landscape surrounding the putative successional lamina and to provide additional evidence supporting the presence of a specialized stromal microenvironment associated with tooth regeneration. Overall, the revisions have substantially strengthened the manuscript.

    3. Reviewer #2 (Public review):

      Summary:

      Mubeen and colleagues study the cellular basis of tooth regeneration in cichlid fish. Using an elegant tooth plunking strategy followed by single nucleus RNA-sequencing, the authors were hoping to achieve an atlas of cellular and transcriptional changes that occur within and between cells during whole tooth replacement.

      Strengths:

      The major strengths of the methods and results are high novelty in the approach in a vertebrate with continuous tooth replacement, the temporal analysis of analyzing at plucking and three later time points, the thorough and sophisticated analysis of the snRNA-seq data including the inferring of trajectories and signaling events, and the robust signal of transcriptional differences induced by tooth plucking.

      Weaknesses:

      The major weaknesses of the methods and results are no validation of any of the inferred cell types, no functional tests of whether any of the changes in signaling pathways affect the plucking-induced tooth replacement process, and perhaps no clear take-away message for biologists not necessarily interested in tooth replacement.

      Conclusions:

      The authors achieved their aims of identifying the changes in gene expression and cellular composition that occur during whole tooth replacement accelerated by plucking. Overall, the results support their conclusions, although some slight semantic qualifiers should probably be added (e.g. referring to "cell types" as "putative cell types").

      The work should have high impact in the field of tooth and organ regeneration, and the novel methodological paradigm established here of accelerating tooth replacement three-fold by plucking has great promise for future follow up studies to further study this process. The work also could have strong impact by the computational methods used here to infer trajectories and signaling interactions. Specific pathways, genes, and cell types could be tested in other fish such as zebrafish to test function during tooth replacement.

      The work is unique and interdisciplinary and also has significance by establishing that robust phenotypically plastic accelerations in regeneration rates occur upon tooth removal. There are very few studies like this one that combine genetic x environmental studies of regeneration. The result that three different species of cichlid fish that normally have very different tooth patterns all accelerate tooth replacement threefold upon tooth plucking also has significance in revealing a highly conserved plucking response.

    4. Author response:

      The following is the authors’ response to the original reviews.

      Many thanks to the three reviewers and the editors for their thoughtful comments and careful evaluation of our manuscript. These are fair, consistent and largely expected comments. On behalf of my co-authors, we provide this response to the public reviews to summarize the main issues raised and the corresponding revisions we have made in the revised manuscript.

      (1) The main consistent comment from all three referees was that our single-nucleus RNA-seq data should be further validated. The reviewers differ in the detail of exactly what they think should be validated, but collectively these comments referred to validation of: (1) the identified cell types, (2) pathways inferred from trajectory analysis, (3) differentially expressed genes between plucked and control conditions across the four sampled time points, and/or (4) inferred ligand–receptor pairs from the cell–cell communication analysis.

      We believe that we are on strong footing for some of these points because of extensive work we’ve done in the past in the cichlid fish model.

      In the references cited in the manuscript and highlighted below (References 1, 10, 11, 29, 30, 31), we tally 29 figures with 273 individual figure panels presenting histology, in situ hybridization, and immunohistochemistry featuring genes expressed in cichlid (replacement) teeth. Most of these genes are markers of dental competency and/or indicative of regenerative potential.

      In addition, in multiple of these papers, we use pharmacology to manipulate the role of key pathways (Hh, BMP, Wnt, Notch) in cichlid tooth development and replacement. Validation of cell types in the present study therefore draws on these published data in cichlids (and other vertebrates), as well as on an unbiased comparative approach, SAMap, which identifies homology between cichlid and mouse dental cell types based on shared gene expression.

      In short, experiments to validate cell types and pathways active in cichlid teeth have been published and are referenced herein. We recognized, however, that these references (some of which include Gareth Fraser as an author, when he was a postdoc in my group; for Reviewer 2) were cited primarily in the Introduction, rather than in the Rationale/Methods or Results sections. We have therefore clarified these connections in the revised manuscript (line 173-74).

      We have not validated nor analyzed functionally the ligand-receptor pairs we inferred from cell-cell communication analysis. This work is beyond the scope of the current paper, and we now state more clearly that these inferences represent hypotheses to be tested in future studies, although many of these ligand–receptor pairs have been noted in other tooth-related publications cited in the manuscript.

      (2) The biggest weakness of our manuscript, noted by referees, is that we do not provide serial histology to accompany our snRNA-seq time course after plucking. We previously described this as a limitation in the “Study limitations and future direction” section of the Discussion, but we have now strengthened this discussion. In particular, we more explicitly acknowledge that we do not directly document the histological progression of tissue responses across the plucking time course or the degree of tissue damage caused by the plucking paradigm at each sampled time point.

      In the “study limitations” section, we note both issues 1 and 2 and suggest that a spatial transcriptomics experiment across the timespan of plucking<>recovery would address simultaneously the desire to understand cellular context of plucking and cellular/spatial differences in plucked vs control cell-type gene expression.

      (3) Reviewers also asked about the presence and interpretation of stromal cells in our snRNA-seq data. In response, we re-examined the mesenchymal compartment and added additional analyses to better characterize stromal/mesenchymal populations and their inferred trajectories in the revised manuscript. This includes a revised Figure 4, revised text around Figure 4 and revised Supplementary Figures.

      (4) Multiple (minor) suggestions for clarification in text and figures have been adopted throughout the revised manuscript, figure legends, and supplemental materials.

      Overall, we do not anticipate that further reviewer engagement will be necessary, and we believe that editorial review of the revised manuscript should be sufficient.

      References cited in the manuscript, highlighted here:

      (1) Fraser, G. J. et al. An Ancient Gene Network Is Co-opted for Teeth on Old and New Jaws. PLoS Biol. 7, e1000031 (2009).

      (10) Fraser, G. J., Bloomquist, R. F. & Streelman, J. T. Common developmental pathways link tooth shape to regeneration. Dev. Biol. 377, 399–414 (2013).

      (11) Bloomquist, R. F. et al. Developmental plasticity of epithelial stem cells in tooth and taste bud renewal. Proc. Natl. Acad. Sci. 116, 17858–17866 (2019).

      (29) Streelman, J. T., Webb, J. F., Albertson, R. C. & Kocher, T. D. The cusp of evolution and development: a model of cichlid tooth shape diversity. Evol. Dev. 5, 600–608 (2003).

      (30) Fraser, G. J., Bloomquist, R. F. & Streelman, J. T. A periodic pattern generator for dental diversity. BMC Biol. 6, 32 (2008).

      (31) Bloomquist, R. F. et al. Coevolutionary patterning of teeth and taste buds. Proc. Natl. Acad. Sci. 112, (2015).

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The authors used single-nucleus RNA sequencing (snRNA-seq) to investigate accelerated tooth replacement following tooth plucking in cichlid fish. They analyzed four stages of regeneration using elegant and well-designed approaches to characterize cellular trajectories and interactions within the dental epithelium and mesenchyme during the accelerated replacement process. Their analyses identified cell-type-specific gene expression profiles and intercellular signaling interactions associated with whole-tooth regeneration.

      Strengths:

      This is a highly interesting and thoughtfully executed study that provides compelling and convincing insights into the mechanisms underlying accelerated tooth regeneration.

      Weaknesses:

      The manuscript currently lacks experimental validation of the single-nucleus RNA-seq data.

      We thank Reviewer #1 for the thoughtful and positive assessment of our study, including the recognition that our snRNA-seq time course provides insight into cellular trajectories, cell-type-specific gene expression, and inferred intercellular signaling during accelerated tooth replacement in cichlid fish. We also appreciate the reviewer’s central concern that the manuscript would be strengthened by additional experimental validation of the single-nucleus RNA-seq data.

      As summarized above and discussed in more detail in our point-by-point responses below, we have clarified how the present cell-type annotations and pathway interpretations are supported by extensive prior experimental work in the cichlid tooth model, including histology, in situ hybridization, immunohistochemistry, and pharmacological perturbation of major developmental pathways. We have also added analyses demonstrating reproducibility across biological test subjects and consistency of representative differentially expressed genes between paired plucked and control samples. Finally, we have strengthened the Study Limitations section to more clearly state that future spatial transcriptomic, histological, and functional validation experiments will be important next steps.

      Reviewer #2 (Public review):

      Summary:

      Mubeen and colleagues studied the cellular basis of tooth regeneration in cichlid fish. Using an elegant tooth plunking strategy followed by single-nucleus RNA-sequencing, the authors were hoping to achieve an atlas of cellular and transcriptional changes that occur within and between cells during whole tooth replacement.

      Strengths:

      The major strengths of the methods and results are high novelty in the approach in a vertebrate with continuous tooth replacement, the temporal analysis of analyzing at plucking and three later time points, the thorough and sophisticated analysis of the snRNA-seq data, including the inference of trajectories and signaling events, and the robust signal of transcriptional differences induced by tooth plucking.

      Weaknesses:

      The major weaknesses of the methods and results are no validation of any of the inferred cell types, no functional tests of whether any of the changes in signaling pathways affect the plucking-induced tooth replacement process, and perhaps no clear takeaway message for biologists not necessarily interested in tooth replacement.

      Conclusion:

      The authors achieved their aims of identifying the changes in gene expression and cellular composition that occur during whole tooth replacement accelerated by plucking. Overall, the results support their conclusions, although some slight semantic qualifiers should probably be added (e.g., referring to "cell types" as "putative cell types").

      The work should have a high impact in the field of tooth and organ regeneration, and the novel methodological paradigm established here of accelerating tooth replacement three-fold by plucking has great promise for future follow-up studies to further study this process. The work could also have a strong impact through the computational methods used here to infer trajectories and signaling interactions. Specific pathways, genes, and cell types could be tested in other fish, such as zebrafish, to test function during tooth replacement.

      The work is unique and interdisciplinary, and also has significance by establishing that robust phenotypically plastic accelerations in regeneration rates occur upon tooth removal. There are very few studies like this one that combine genetic and environmental studies of regeneration. The result that three different species of cichlid fish that normally have very different tooth patterns all accelerate tooth replacement threefold upon tooth plucking also has significance in revealing a highly conserved plucking response.

      We thank Reviewer #2 for the careful and constructive evaluation of our manuscript and for highlighting the novelty of the cichlid tooth-plucking paradigm, the temporal design of the snRNA-seq experiment, and the computational analyses used to infer cellular trajectories and signaling interactions during accelerated tooth replacement. We also appreciate the reviewer’s comments regarding validation of inferred cell types and interpretation of signaling pathways.

      In response, we have revised the manuscript to clarify that our cell-type annotations are supported by marker-gene expression, previously published cichlid tooth studies, and an unbiased comparative approach, SAMap, which relates cichlid and mouse dental cell types based on shared gene-expression structure. We have also clarified that inferred ligand-receptor interactions represent computational hypotheses rather than functionally validated mechanisms. In addition, we revised Figure 6 and Figure 7A to improve the readability and interpretation of inferred signaling results, and we edited the relevant text and figure legends to make these results easier to follow. These points are addressed in greater detail in the point-by-point responses below.

      Reviewer #3 (Public review):

      Summary:

      This is an interesting paper. The process of tooth exfoliation and replacement in vertebrates remains an intriguing and fascinating subject of inquiry. As the scientists noted, there are no mammalian models that can be used to examine signaling pathways in real time.

      Strengths:

      This work integrates in vivo and high-resolution transcriptomics. The study confirms previous findings and emphasizes the need for additional research into the processes that drive the restoration of missing teeth for future therapeutic uses.

      Weaknesses:

      I disagree with the use of the phrase "plucking". Instead, the authors use tooth extraction or tooth removal, which is clinically more correct for the procedure they are doing.

      The inspiration for our ‘plucking’ experiment is work done in the hair follicle model (lines 73-74). Because cichlid teeth are so numerous, are very small, and lack dental roots, this is an accurate description of the procedure. We opt to retain the phrasing.

      The title is rather broad and appears to be more appropriate for a review than an original research work. I would advise specifying the species under research and/or the sort of damage model used in the transcriptome analysis.

      We opt to retain the title.

      It's uncertain whether the findings are exclusively based on regeneration. The presence of tooth remnants, as well as unintended harm to surrounding tissues, may have triggered repair mechanisms, thereby biasing the current data. How did the authors handle this issue? The oral cavity was under severe manipulation, increasing the inflammatory stimuli, a situation that does not take place in physiological exfoliation.

      In the revised manuscript, we have more clearly acknowledged that our plucking paradigm may induce tissue damage and repair-associated responses in addition to accelerated tooth replacement. We have strengthened the Study Limitations section to state that we do not directly document the histological progression of tissue responses across the time course or the degree of damage caused by plucking at each sampled time point. One caveat, however, is that bone remodeling and immune response is likely triggered on the ‘control’ side of the jaw also, just not to the same degree as after plucking.

      The authors indicated the use of microCT analysis; however, no such information appears in the main text. In fact, this manuscript lacks anatomical information. It is required to conduct histological examinations of the regenerated teeth at various time points.

      microCT data were included as a Supplemental Figure to demonstrate the dental formulae of our chosen species; but we did not characterize post-plucking recovery using this technique (see above summary and below point-by-point comments).

      Although the current findings confirm previously found and verified signaling pathways, the absence of functional data lends uniqueness to this work.

      In the revised manuscript, we also clarify that, while our transcriptomic analyses identify candidate cell states, pathways, and signaling interactions associated with accelerated replacement, the functional roles of these inferred pathways remain to be verified in future studies.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Major Points:

      (1) Figure 1 should include representative H&E staining images comparing the left control side and the regenerated side at 7 days post-plucking. This would provide important histological context for the regeneration process and help readers better interpret the molecular findings.

      This would indeed be valuable information, but we did not carry out histology of paired control vs plucked jaws to accompany our pulse-chase and dissections for single-nucleus isolation. This comment is similar to that below about validation of what is happening on plucked vs. control jaw halves and is the first limitation we discuss in the “study limitations and future directions” section (from line 538).

      (2) Each tooth position consists of a functional tooth, a replacement tooth, and the dental (successional) lamina. On the control side, the successional lamina contains teeth at different developmental stages, analogous to the mammalian bud, cap, and bell stages. Can the snRNA-seq analysis distinguish among tooth families at different developmental stages, as well as the individual components within a single tooth unit? Clarification of this point would enhance the developmental interpretation of the dataset.

      No, our approach does not distinguish among teeth at different stages, nor among teeth in even vs odd positions that tend to be synchronized in replacement cycles. Theoretically, one could do this by dissecting individual teeth and pooling by tooth stage, but we did not.

      (3) Identifying successional lamina cells is critical, and the authors report putative SL cells within the VEE cluster. However, the stromal cells surrounding the successional lamina are also known to play important roles in tooth regeneration. Can the authors further annotate and characterize stromal cell populations in the snRNA-seq dataset? Additional analysis of these supporting cells would strengthen the conclusions regarding epithelial-mesenchymal interactions.

      We thank the reviewer for this insightful suggestion. In response, we further characterized mesenchymal subpopulations and included these new analyses in the revised manuscript (updated Figure 4 and Supplemental Figure 8). Specifically, pseudotime and CellRank analyses identified a mesenchymal subpopulation enriched for Twist1, Dnmt1, and Runx2, which we interpret as putative dental ectomesenchyme (DEM) based on the established roles of these genes in odontogenic mesenchymal development and differentiation, as well as their reported expression in mouse and human tooth single-cell transcriptomic studies. Notably, this putative dental ectomesenchymal population resides within the broader dental follicle compartment identified in our dataset (Figure 4A-C).

      To further assess supporting stromal populations, we examined the expression of established stromal marker genes, including Lum, Col6a3, Aspn, and Vegfc. These markers were broadly restricted to mesenchymal populations and showed strong enrichment overlapping the newly identified putative dental ectomesenchymal region (Supplemental Figure 8B). Consistent with these observations, differential expression analysis identified additional DEM-enriched genes that substantially overlap canonical stromal markers, including extracellular matrix-associated genes, supporting a close transcriptional relationship between the putative dental ectomesenchyme and the surrounding stromal mesenchymal compartment (Supplemental Figure 8C). Together, these findings refine the mesenchymal landscape surrounding the putative successional lamina and support the presence of a specialized stromal microenvironment associated with tooth regeneration.

      Consistent with this interpretation, our CellChat analysis identified significantly increased interactions between the dental ectomesenchyme and cycling ameloblast populations on the plucked side at Day 0 (Supplemental Figure 8D). These interactions were enriched for signaling pathways including SEMA4, EPHB, SLIT and SPP1, all of which have established roles in tissue remodeling, extracellular matrix organization, and regenerative processes. (Supplemental Figure 8E). Because Day 0 contained sufficient biological replicates and cell numbers for robust statistical comparison, we focused our interaction analyses on this time point. Collectively, these additional analyses provide a more comprehensive characterization of the stromal compartment and further support the conclusion that a specialized dental ectomesenchymal population actively participates in epithelial-mesenchymal communication during the earliest stages of tooth regeneration. So, in total, Figure 4 was revised, the text on lines 281-309 was revised, and Supplemental Figure 8 was added.

      (4) The manuscript currently lacks experimental validation of the single-nucleus RNA-seq data. The authors should validate the expression of major signature genes using RNAscope or immunostaining, ideally comparing regenerated samples with the left-side control. Such validation would significantly enhance the robustness of the conclusions.

      We did not validate up- or down-regulation of differentially expressed genes in intact tissue, owing in part to (1) the complexity of this experiment, (2) the fact that the majority of DEGs, or ‘major signature genes’ have been observed to be expressed in dentitions generally, and often by us in previous work on cichlid teeth, and the fact that (3) independent biological replicates were strongly consistent in the direction of effects (see below). In the “study limitations” section, we note this issue and suggest that a spatial transcriptomics experiment across the timespan of plucking<>recovery would address simultaneously the desire to understand cellular context of plucking and cellular/spatial differences in plucked vs control cell-type gene expression.

      Minor Points

      (1) In Figure 1, the color scheme used in the schematic drawing (Figure 1A) should match the corresponding structures shown in Figure 1B to improve clarity and consistency.

      We appreciate the reviewer’s thoughtful suggestion regarding the color consistency between the schematic (Figure 1A) and the fluorescence images (Figure 1B). However, the color scheme in the schematic (Figure 1A) was intentionally selected to maximize accessibility, particularly for readers with color vision deficiencies, and therefore differs from the magenta and green fluorescence channels used in Figure 1B. In the fluorescence images, the magenta and green colors reflect the native display colors used for the Alizarin Red and Calcein labeling channels in the pulse-chase experiment. Directly matching the schematic colors to the fluorescence images could reduce the visual contrast between key anatomical structures and compromise accessibility for some readers. We have therefore retained the current color scheme in Figure 1A while ensuring that the corresponding structures are clearly identified through consistent labels and annotations across both panels.

      (2) The abbreviation for successional lamina (SL) should be defined upon first use in the Introduction.

      We thank the reviewer for catching this omission. We have now defined the abbreviation “successional lamina (SL)” upon its first appearance in the Introduction.

      (3) Regarding biological replicates, the authors should provide data demonstrating the consistency and reproducibility across replicated samples.

      We thank the reviewer for this suggestion. To demonstrate the consistency and reproducibility across biological test subjects, we have added analyses summarizing sequencing quality metrics, test subject contributions, integrated clustering, and representative differential gene expression across individual samples (see Figure S4, panels C, D & E and Author response image 1). Panel A shows that nuclei from different biological test subjects are well integrated across clusters rather than segregating by sample origin. Finally, Panel B presents representative differentially expressed genes from multiple cell populations, demonstrating consistent expression differences between paired plucked and control samples across biological test subjects.

      Author response image 1.

      (A) UMAP embedding of dental nuclei. Each point represents a single nucleus, colored by test subject. (B) Representative differentially expressed genes show consistent expression differences between plucked and control samples across biological replicates. Paired boxplots of average gene expression for representative differentially expressed genes from multiple cell populations at Days 0, 1, 3, and 7. Each point represents one biological replicate (test subject), with paired plucked and control samples connected by dashed lines. The y-axis shows average gene expression, and the x-axis indicates the experimental condition. These representative examples illustrate the consistent direction of differential expression across biological replicates, supporting the reproducibility of the single-nucleus RNA-seq dataset.

      Reviewer #2 (Recommendations for the authors):

      (1) Figure 1: Can the panels to the right of panel B be labeled? It's not clear what these six images are showing, so giving them letters and explaining briefly in the legend what the point of each panel is would clarify. "Right, example of individually classified teeth" - can the authors elaborate on what each tooth is an example of (i.e., how each tooth shown was classified"?) For clarity, the graphs in panels C and D should have the y-axes labeled

      We thank the reviewer for this helpful suggestion. In response, we revised the Figure 1B legend to clarify the classification criteria used for dye incorporation analyses and to better describe the representative fluorescence images. Specifically, teeth positive for both Alizarin and Calcein were classified as pre-existing old teeth, whereas teeth positive only for Calcein were classified as newly formed teeth. We additionally clarified that the images to the right of panel B show representative individually classified teeth, with the top row representing pre-existing old teeth and the bottom row representing newly formed teeth. We also added y-axis labels to panels C and D to improve figure clarity and readability.

      (2) Figure 2 legend: should "the cell type" instead be "the putative cell type"? Without validation for all cell types, it seems adding some sort of qualifier is in order here. Can the authors comment further on examples of validation from other studies? For example, Gareth Fraser has published numerous studies that show Pitx2 expression marking dental epithelium in different fish, yet none of these older papers are cited.

      Identification and validation of cell types make use of multiple published datasets in cichlids (for markers matched to mouse), as well as an unbiased computational approach (SAMap) that draws homology between cichlid and mouse dental cell types, based on shared global patterns of gene expression. There is perhaps a philosophical debate to be had about the validity of ‘cell types,’ generally, but our data are validated using two methods. We edited the text in lines 167-177 to clarify, including citing references to our own work (these studies include Gareth Fraser as an author, when he was a postdoc with Streelman).

      (3) Figure 6 is extremely complicated. Can any portions of rows or columns in these tables be highlighted in the figure to help the reader follow the proposed signaling interactions highlighted in the text?

      We thank the reviewer for this helpful suggestion. To improve the readability of Figure 6 and better guide readers through the dynamic signaling patterns described in the text, we revised the figure by visually highlighting the key sender-receiver interaction regions discussed in the Results. Specifically, we annotated the interactions involving mesenchymal subpopulations and alveolar bone (OST) signaling toward CYC-AMB at Days 0 and 7, mesenchymal signaling toward NK/T cells at Day 1, and epithelial cross-talk centred around ES-2 at Day 3. These visual annotations allow readers to more readily identify the signaling interactions highlighted in the text and relate them to the corresponding regions of the interaction heatmaps.

      (4) In Figure 7A, what does the black font indicate (if grey is up in control and red is up in plucked)? I'd guess not up in either, which then makes it unclear whether the sets in black are different or why they are being presented.

      We thank the reviewer for pointing out this ambiguity. In Figure 7A, blue and red labels indicate signaling pathways identified by CellChat as condition-specific, with blue representing pathways detected only in the control condition and red representing pathways detected only in the plucked condition. In contrast, pathways shown in black represent signaling pathways detected in both conditions but exhibiting significant differences in inferred communication probability between conditions. Thus, the black labels denote shared signaling pathways whose activity differs significantly between control and plucked samples, rather than pathways unique to either condition. We have revised the figure legend to clarify this distinction and improve interpretability.

      Reviewer #3 (Recommendations for the authors):

      (1) I encourage the authors to offer information on the histological differences between teeth during physiological and accelerated replacement. I'm curious if the eruption's accelerated rate has any effect on the mineralization of those teeth.

      We did not examine the histology of individual teeth, and so can’t comment on differences in mineralization.

      (2) The findings section contains multiple sentences that should be moved under material and techniques.

      We expect the reviewer is referring to paragraph lines 104-114, which was a tricky paragraph to place in the manuscript. In the end, we believe it represents important context necessary to interpret findings (which could be missed if moved to ‘methods’) and so we’ve chosen to keep this paragraph in its place.

      (3) It would be useful to include a table showing sample distribution by experimental design.

      We thank the reviewer for this suggestion. Sample distributions across experimental conditions, time points, biological test subjects, and identified cell populations are already provided in Supplementary Table 1. To improve clarity and accessibility, we have revised the table legend to more explicitly describe the experimental design and sample annotations represented in the table.

      (4) The writers did a nice job with the graphics in Figure 8; however, the schematics in Figure C are difficult to follow and are not adequately discussed anywhere. Please note that this text may be of great interest to the dentistry community, including clinicians, and that a clear and succinct explanation of the schemes at the end would be quite beneficial.

      We thank the reviewer for this helpful suggestion. We have revised the Figure 8 legend to more clearly explain Panel C as a summary schematic of inferred cell–cell communication events associated with accelerated tooth replacement after plucking. The updated legend clarifies that the pathway labels in Figure 8C summarize results directly from Figure 7A: red pathway labels indicate plucked-only signaling events, corresponding to pathways shown as full red bars in Figure 7A, while black pathway labels indicate signaling interactions detected in both plucked and control conditions but showing significant differences in interaction probability between conditions. Panel C also includes a cell-type legend at the bottom to identify the relevant cell populations.

    1. eLife Assessment

      This manuscript provides an important contribution to the field of platelet biogenesis, and the convincing evidence will advance our understanding of signal transduction driving the development of late megakaryopoiesis and platelet reactivity that results in bleeding diathesis. The paper is noteworthy for analyzing two related tyrosine phosphatases, using single or combined conditional gene knockouts at different developmental stages. Because SHP1 is a negative regulator and SHP2 is an activator, the synergistic effects found in the double knockout were surprising.

    2. Reviewer #3 (Public review):

      Summary:

      In this manuscript, Barré et al utilize the Gp1ba-Cre transgenic mouse model to build upon previous findings in a Pf4-Cre system to investigate the effects of individual and combined Shp1 and Shp2 deletion in megakaryocytes and platelets. They report decreased megakaryocyte maturation, macrothrombocytopenia, and increased blood loss primarily in association with the Shp1/Shp2 double-knockout condition. The authors further show that this phenotype appears to be driven primarily by Shp2 and implicate dysregulation of Tpo signaling and downstream Ras/MAPK pathways, including ERK1/2. They propose that Shp1 may be functioning through a distinct pathway that has yet to be identified, opening up areas for future study.

      Strengths:

      Overall, the experiments combine in vitro, in vivo, and ex vivo approaches and appear to have been carefully designed and carried out, with multiple technical and biological replicates where relevant. The authors make a compelling argument for using the Gp1ba-Cre as opposed to the Pf4-Cre system and demonstrate both the dose- and stage-dependent effects of Shp1 and Shp2 on megakaryopoiesis and thrombopoiesis. They find that Shp1 and Shp2 are required in late-stage megakaryocyte maturation and that even low levels of expression compared to baseline are likely sufficient to yield generally normal megakaryocytes. Their findings also lead to specific future directions, such as the mechanism by which Shp1 regulates megakaryopoiesis and thrombopoiesis that is distinct from Tpo-mediated signaling. Figure 8 is particularly effective in summarizing the different models and pathways presented.

      Weaknesses:

      The effects of Shp1 and Shp2 knockouts are described as "synergistic," but it is not always clear that the effects are synergistic vs. additive, especially as the specific mechanism by which Shp1 functions in megakaryocyte development has yet to be identified. On a more minor point, although a significant part of the introduction focuses on the role of Mpl signaling in human disease, there is ultimately limited reference to Mpl (although there is of course a strong focus on Tpo) and the potential clinical implications of the findings presented here.

    3. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      This manuscript provides an important contribution to the field of platelet biogenesis, and the convincing evidence will advance our understanding of signal transduction driving the development of late megakaryopoiesis and platelet reactivity that results in bleeding diathesis. The paper is noteworthy for analyzing two related, either singly or in combination, tyrosine phosphatases in this conditional, stage development gene knockout. Because SHP1 is a negative regulator and SHP2 is an activator, the synergistic effects found in the double knockout were surprising.

      We thank the reviewer for acknowledging the importance and novelty of our findings.

      Public Reviews:

      Reviewer #1 (Public review):

      Barré et al. investigated the role of Shp1 and Shp2 in megakaryocytes (MKs) and platelets by conditional knock-out of Shp1, Shp2, or both under the control of the Gp1ba promoter. Deletion of Shp1 and Shp2 in MKs and platelets was almost complete. The Shp1/Shp2 double knock-out mice displayed macrothrombocytopenia and increased bleeding, whereas the single knock-outs did not show significant defects. Platelet function was aberrant in DKOs, but not in single knock-outs, and so was ligand-induced signaling, particularly Syk phosphorylation.

      Megakaryocyte maturation was impaired in Shp1/Shp2 DKO mice. Ligand-induced signaling was impaired in Shp2 knock-out and DKO. Ex vivo formation of platelets and in vivo maturation of MKs were impaired in DKO mice. Pharmacological inhibitors of Shp1 and Shp2 had largely similar effects as observed in the single knock-outs. The authors conclude that Shp1 and Shp2 have synergistic functions in the MK/platelet lineage, and that Shp2 may be a potential therapeutic target in myeloproliferative neoplasms.

      Strengths:

      The data clearly show effects of the Shp1/Shp2 double knock-out on MKs and platelets.

      Weaknesses:

      There appears to be a discrepancy between the results with the Shp2 single knock-out and the Shp2 inhibitor: the Shp2 knock-out does not affect MKs and platelets, except Erk1/2 signaling, whereas the Shp2 inhibitors appear to affect MK function.

      This work is interesting and may have potential from a therapeutic point of view.

      Pharmacological effects do not always correlate with congenital anomalies arising for genetic defects. The Shp2 allosteric inhibitors used in our study only inhibit catalytically inactive Shp2, whereas targeted deletion of Ptpn11 results in a loss of total Shp2 expression, including catalytic and non-catalytic related functions, with developmental consequences. Further, Gp1ba-Cre+; Shp2fl/fl megakaryocytes express approximately 22% of normal Shp2 level, which likely also contributes to differences observed between pharmacological inhibition and genetic ablation of Shp2.

      We thank the reviewer for recognizing the therapeutic potential of our findings.

      Reviewer #2 (Public review):

      Summary:

      In this manuscript, Barré et al. investigate the roles of the phosphatases Shp1 and Shp2 in the megakaryocyte and platelet lineage using genetic depletion in mice. By employing Gp1ba-Cre-based models, the study builds on the authors' previous work and addresses some limitations associated with earlier Pf4-Cre approaches. The authors report relatively mild alterations in megakaryocyte and platelet parameters in mice lacking either Shp1 or Shp2 alone, whereas combined deletion of both phosphatases results in macrothrombocytopenia, mild bleeding, and impaired GPVI-dependent platelet aggregation accompanied by reduced Syk phosphorylation. The functional platelet defects are linked to reduced expression of GPVI and integrin α2, while thrombocytopenia is associated with impaired megakaryocyte maturation, reduced ploidy, defective proplatelet formation, and altered TPO-dependent Ras/MAPK signaling. Similar effects on megakaryopoiesis are also observed in vitro following treatment with newly developed Shp2 inhibitors.

      Strengths and Weaknesses:

      The study addresses an important biological question and presents a substantial dataset that could contribute to a better understanding of Shp1 and Shp2 function in platelet biology. However, several aspects of data presentation and interpretation would benefit from additional clarification. In particular, while the authors conclude that single genetic deletion or pharmacological inhibition of Shp1 has a limited impact and that the major phenotypes are specific to combined Shp1/2 deletion or Shp2 inhibition, some of the data suggest more nuanced effects that may warrant further discussion.

      We thank the reviewer for raising this point. The manuscript is being revised accordingly, including highlighting the potential role of Shp1 in megakaryopoiesis and thrombopoiesis under steady-state and stressed conditions, requiring more detailed investigation.

      Reviewer #3 (Public review):

      Summary:

      In this manuscript, Barré et al utilize the Gp1ba-Cre transgenic mouse model to build upon previous findings in a Pf4-Cre system to investigate the effects of individual and combined Shp1 and Shp2 deletion in megakaryocytes and platelets. They report decreased megakaryocyte maturation, macrothrombocytopenia, and increased bleeding primarily in association with the Shp1/Shp2 double-knockout condition. The authors further show that this phenotype appears to be driven primarily by Shp2 and implicate dysregulation of Mpl signaling and downstream Ras/MAPK pathways, including ERK1/2. Given the key role of these pathways in human diseases such as myeloproliferative neoplasms and the challenges associated with modulating such a central pathway, identification of a specific regulator of Mpl signaling poses intriguing questions for future studies on clinical applicability.

      We thank the reviewer for acknowledging the importance and novelty of our findings.

      Strengths:

      Overall, the experiments combine in vitro, in vivo, and ex vivo approaches and appear to have been carefully designed and carried out, with multiple technical and biological replicates where relevant. The authors make a compelling argument for using the Gp1baCre as opposed to the Pf4-Cre system and demonstrate both the dose- and stagedependent effects of Shp1 and Shp2 on megakaryopoiesis and thrombopoiesis. They find that Shp1 and Shp2 are required in late-stage megakaryocyte maturation and that even low levels of expression compared to baseline are likely sufficient to yield generally normal megakaryocytes. Their findings also lead to specific future directions, such as the mechanism by which Shp1 regulates megakaryopoiesis and thrombopoiesis that is distinct from TPO-mediated signaling.

      Weaknesses:

      While the experiments have been thoughtfully designed and carried out, there is limited background explanation on relatively complex or niche pathways/mechanisms, such as the relationship between P-selectin, CRP, and PAR4p; the interactions between SFK, Syk, GPVI, and CLEC-2; and TPO, MPL, ERK1/2, AKT, and STAT3, which, while likely intuitive to experts in their respective fields, may be less obvious to a reader approaching this manuscript with a global interest in megakaryopoiesis/thrombopoiesis and thus detract from the impact of the findings.

      We thank the reviewer for raising this point. The manuscript is being revised to better explain the rationale and molecular mechanisms linking these pathways and functions.

      With regard to the science itself, some of the conclusions feel premature based on the available data.

      (1) The section "Aberrant ITAM signaling in Shp1- and Shp2-deficient platelets" is challenging to follow for those not well-versed in ITAM signaling and associated pathways, and may take additional outside reading to follow the conclusion that Syk-dependent signaling is modulated downstream of GPVI and CLEC-2 based on lack of change in Src p-Tyr418, especially considering that Src p-Tyr418 was previously introduced as a measure of SFK rather than Syk. In the introduction, Shp1 is specifically mentioned as a negative regulator of the ITAM/Syk/phospholipase pathway. However, in Figure 4Ai and Bi, Syk phosphorylation/activation in Shp1 knockout cells did not appear to be different from Shp2 knockout cells, and is lower than the control, which is surprising for a negative regulator. It is also not clear why, in the section (Figure 4A-B), there is reduced Syk activation in Shp1 and Shp2 single knockout cells upon CLEC2 stimulation (but apparently not with CRP) when there was no difference in response to CLEC2 (but a difference in response to CRP) in the previous section (Figure 3A, C).

      We thank the reviewer for raising these important points. The manuscript is being revised accordingly, including clarifying the roles of SFKs, Shp1 and Shp2 in the ITAM-Syk-PLCγ2 signaling pathway.

      Briefly, SFKs are essential for phosphorylating ITAMs, allowing SH2-dependent docking of Syk. Reduced reactivity of Shp1/2 DKO platelets to CRP and collagen is likely due to downregulation of the ITAM-containing GPVI-FcR γ-chain complex and integrin α2 subunit, and concomitant reduction in Syk phosphorylation.

      However, the marginal albeit significant reduction in Syk phosphorylation downstream of CLEC-2 in Shp1 and Shp2 KO platelets was not determined and was insufficient to impact CLEC-2-mediated platelet aggregation under the conditions tested.

      Differences in the stoichiometry and docking of Syk to phosphorylated GPVI-FcR γ-chain and CLEC-2 likely contribute to the differences in platelet reactivity and Syk phosphorylation downstream of the two receptors in the absence of Shp1 and Shp2.

      (2) In the section "Reduced Tpo signaling in Shp1/2-deficient MKs," only Western blot data for (p)ERK1/2, AKT, and STAT3 are presented before concluding that decreased ERK1/2 activity is a mechanistic explanation for thrombocytopenia seen in the Shp1/2 doubleknockout condition. Such a statement would benefit from additional experiments, such as protein or transcriptional levels of ERK1/2 targets specifically relevant to megakaryopoiesis, such as ETS, FOS, and JUN, to assess the consequences of decreased phosphorylated ERK1/2.

      We thank the reviewers for these constructive comments. Further experiments are being planned to determine the biological and transcriptional consequences of reduced ERK1/2 phosphorylation during megakaryopoiesis and thrombopoiesis.

      (3) Suggesting that "inhibiting Shp2 will not have any bleeding consequence in patients" and that Shp2 may be a therapeutic target in myeloproliferative neoplasms when none of these studies have been carried out in a human model is a bold conclusion. There are no data presented on, for example, whether Shp2 inhibition can help reverse the MPL/JAK/STAT pathway in the setting of gain-of-function mutations specifically associated with myeloproliferative neoplasms.

      This conclusion is being tempered in the revised manuscript. Genetic- and pharmacological-based approaches will be used to establish the therapeutic potential of inhibiting Shp1 and Shp2 in mouse models of MPN, including Jak2 gain-of-function mice. Bleeding and thrombotic complications of inhibiting Shp1 and Shp2 will be explored as part of these studies.

      Recommendations for the authors:

      Reviewing Editor Comments:

      Altogether, we feel that this is an important study for those in the fields of hematology or signal transduction. Your important study characterizes the roles in late megakaryopoiesis and platelet biogenesis of single or combined conditional deletion of two tyrosine phosphatases, Shp1 and Shp2. Strengths include technical advances in single and combined deletions, the somewhat surprising results of synergy between the two phosphatases, focusing on the critical stage of late megakaryopoiesis, and clinical implications in bleeding diathesis.

      Weaknesses are mostly minor, but the numerous points raised by reviewer 3 need to be addressed and typographical errors corrected. Further discussion should include the relevance or dissimilarity in megakaryopoiesis and platelet biogenesis between murine and human blood health and disease. Since SHP1 is a negative regulator and SHP2 is a positive activator, additional discussion about how they coordinate and fine-tune ("nuanced") signal transduction in TPO- or GPVI-induced signaling in an explicitly stated pathway.

      We invite you to respond to the critiques and submit a revised manuscript.

      Sincerely,

      Seth Corey, MD MPH

      We thank the editor for the positive evaluation of our study and for highlighting its relevance to the fields of haematology and signal transduction. We have carefully addressed all comments raised by Reviewer 3 and corrected typographical errors throughout the manuscript.

      As suggested, we expanded the Discussion to better address the relevance of our murine findings to human megakaryopoiesis and platelet biogenesis. While our study relies on mouse models, key components of TPO/MPL signaling and platelet production are conserved between mice and humans, although differences in megakaryocyte maturation dynamics and platelet biology are acknowledged and now discussed.

      We also clarified the coordinated roles of Shp1 and Shp2 in signaling. Although Shp1 generally acts as a negative regulator and Shp2 as a positive mediator of signal transduction, our results suggest that they function in a complementary manner to optimize signaling downstream of TPO/MPL and GPVI pathways, thereby ensuring appropriate regulation of late megakaryopoiesis, platelet production and activation.

      These additional considerations have been incorporated into the revised manuscript to provide a clearer conceptual framework for how Shp1 and Shp2 cooperate to regulate platelet biogenesis.

      Reviewer #1 (Recommendations for the authors):

      (1) The effects of the Shp1/Shp2 DKO are clear, but the effect of the Shp2 single knock-out is less clear on all parameters that were tested. The exception is ERK1/2 phosphorylation, which was reduced in the Shp2 knock-out as well as the Shp1/Shp2 DKO. Why do the authors conclude that Shp2 may be a potential therapeutic target, while the data show that knock-out of Shp1 and Shp2 is required for the observed effects?

      We agree that the most pronounced phenotypes were observed in the Shp1/Shp2 DKO. However, Shp2 single knock-out consistently reduced ERK1/2 phosphorylation, indicating that Shp2 contributes to MPL downstream signaling in megakaryocytes. The absence of a strong phenotype in Shp2 single knock-out may be due to residual Shp2 protein. However, given the established role of the Shp2–ERK pathway in megakaryopoiesis and the observation that pharmacological Shp2 inhibition significantly affected MK ploidy, proplatelet formation, and ERK1/2 phosphorylation, our data support a contribution of Shp2 to these processes and suggest it as a potential therapeutic target.

      (2) Inhibitors of Shp1 and Shp2 had largely similar effects as Shp1 and Shp2 single knock-outs, respectively. The effect of Shp2 knock-out on MK ploidy is not clear, cf. Figure 5Ai (no effect) and Figure 5Aii (reduction, which is not significant), whereas a clear and significant effect was reported for the Shp1/Shp2 DKO. In contrast, in Figure 7Ciii, the Shp2 inhibitors SHP099 and RMC-4550 clearly affect MK ploidy and the percentage of MKs forming proplatelets. The discrepancy between the effect of Shp2 knock-out and Shp2 inhibitors suggests that the inhibitors may affect other targets. The authors should consider using the Shp2 inhibitors on the Shp2 knock-out to prove or disprove that the effects of the Shp2 inhibitors are mediated exclusively by Shp2.

      Pharmacological inhibition does not necessarily phenocopy genetic deletion. The allosteric Shp2 inhibitors used in our study (SHP099 and RMC-4550) stabilize Shp2 in an inactive conformation and inhibit its catalytic activity, whereas Ptpn11 deletion results in complete loss of the Shp2 protein, including both catalytic and scaffolding functions. These mechanistic differences may lead to distinct biological outcomes and could explain the discrepancy observed between Shp2 knockout and inhibitor treatments.

      (3) Since the most profound effects were found in the Shp1/Shp2 DKO, it would be interesting to use combinations of the Shp1 and Shp2 pharmacological inhibitors to mimic the effect of the Shp1/Shp2 DKO.

      We thank the reviewers for these constructive comments. Further experiments are indeed being planned to use combinations of the Shp1 and Shp2 pharmacological inhibitors to mimic the effect of the Shp1/2 DKO.

      Reviewer #2 (Recommendations for the authors):

      Major points:

      (1) Additional details on the strategy used to isolate megakaryocyte progenitors from mouse bone marrow would improve clarity, including sorting approach, gating strategy, and assessment of population purity.

      We thank the reviewer for this suggestion. We have now expanded the Methods section to provide a more detailed description of the strategy used to isolate megakaryocyte progenitors from mouse bone marrow.

      Briefly, bone marrow cells were first enriched for hematopoietic progenitors and stained with antibodies against lineage markers and megakaryocyte-associated markers. Megakaryocyte progenitors were then isolated by flow cytometric sorting based on established surface marker combinations, including c-Kit and CD41 expression. The gating strategy excluded lineage-positive cells and debris before selecting the progenitor population of interest.

      (2) Platelet GPVI expression appears reduced not only in Shp1/2 double-knockout mice but also, to some extent, in single Shp1- or Shp2-deficient models. A more detailed quantitative comparison and discussion would be helpful.

      We thank the reviewer for this observation. Although the most pronounced reduction in GPVI surface expression was observed in Shp1/Shp2 double knock-out platelets, minor variations may appear in the single knock-out models. To address this, we performed additional statistical analyses comparing WT platelets with each single knock-out genotype. These analyses did not reveal any significant statistical differences in GPVI expression between WT and either Shp1- or Shp2-deficient platelets, indicating that the apparent variations fall within the range of biological variability.

      (3) The aggregation traces shown in Figures 3A and 3B would benefit from clarification regarding their representativeness relative to the corresponding quantitative analyses.

      We thank the reviewer for this comment. The aggregation traces in Figures 3A and 3B represent experiments selected from independent replicates included in the quantitative analysis. The figure legends have been revised to clarify that these traces are representative of the experiments summarized in the quantification panels, which include data from multiple independent mice.

      (4) In several experiments, statistical significance may be influenced by differences in sample size across genotypes (e.g., Figures 2Ci, 3Ai, 3Di, and 6Ai). Using comparable numbers of replicates would strengthen the interpretation.

      We appreciate the reviewer’s attention to statistical rigour. The differences in sample size between genotypes reflect the availability of animals from the different breeding cohorts. Importantly, all statistical analyses were performed using appropriate tests that account for unequal sample sizes. The observed differences remain consistent across independent experiments.

      (5) The rationale for assessing only P-selectin exposure following CRP and PAR4p stimulation is not fully explained. Including integrin αIIbβ3 activation, or clarifying its exclusion, would provide a more complete assessment of platelet activation.

      We thank the reviewer for this suggestion. P-selectin exposure was used as a primary readout because it provides a robust measure of α-granule secretion downstream of GPVI and PAR signaling. Integrin αIIbβ3 activation was not assessed in these experiments because platelet aggregation assays were performed in parallel, which already provide a functional readout of integrin activation, as aggregation requires αIIbβ3 engagement. Nonetheless, we agree with the reviewer that direct measurement of integrin activation (e.g., fibrinogen binding) would provide complementary information and will be considered in future studies.

      (6) Figure 3Dii is described as an aggregation assay, although it appears to report P-selectin exposure; this distinction should be clarified.

      We thank the reviewer for identifying this inconsistency. Figure 3Dii reports indeed P-selectin exposure measured by flow cytometry, rather than platelet aggregation. We have corrected the description in the Results section.

      (7) The suggestion of compensatory extramedullary hematopoiesis based on splenomegaly would be strengthened by immunophenotypic analysis of splenic hematopoietic progenitor populations.

      We appreciate this important suggestion. In the current study, the evidence for possible compensatory extramedullary hematopoiesis is mainly based on the splenomegaly observed in Shp1/2 DKO mice. We agree that detailed immunophenotypic analysis of splenic hematopoietic progenitors would provide additional mechanistic insight; however, this was beyond the scope of the present study, which focuses on the intrinsic role of Shp1 and Shp2 in the megakaryocyte and platelet lineage. We have therefore revised the Discussion to present this interpretation more cautiously and to indicate that further studies will be required to determine whether splenic hematopoiesis contributes to compensatory platelet production in this model.

      (8) In Figure S3, differences in platelet recovery kinetics among genotypes appear evident. Clarification of the statistical tests used to assess these differences would be useful.

      We thank the reviewer for this comment. Platelet recovery kinetics were analyzed using two-way ANOVA with appropriate post hoc tests. No statistically significant differences between genotypes were observed. These details have been added to the Methods and figure legend for clarity.

      Reviewer #3 (Recommendations for the authors):

      Overall, the manuscript suffers from multiple typographical and grammatical errors that distract from the data being presented.

      We have carefully revised the manuscript to correct typographical and grammatical errors throughout, improving clarity and readability.

      (1) Figure S1: I believe this should be referenced in the first paragraph of the results section.

      We have now referenced the Supplemental Figure S1 in the first paragraph of the results section as suggested.

      (2) Figure 2A: Although the individual points for the replicates are informative, they do make it difficult to appreciate the SEM, and to my eye it appears that, for example, there may not be a difference between Shp2 and Shp1/2 or that there may be a difference between Shp1 and Shp1/2 in (ii), as Table S2 suggests. In other words, it seems that the increased MPV (as well as the leukocyte phenotype) may be driven by the knockout of Shp2; are there statistical analyses that could be performed to show that the increased MPV is specific to the double knockout?

      We thank the reviewer for this comment. Despite the slightly higher MPV observed in Shp2 single knockouts, statistical analysis using one-way ANOVA, which is appropriate for comparing means across multiple independent groups, and taking all individual data points into account, revealed no significant differences between Shp2 or Shp1 single KO and the Shp1/2 DKO.

      (3) Figure 2Bi: Is this missing a statistical significance bar, or was there no significant difference in cumulative bleeding time between the conditions? If the latter, this should be clarified in the main text (although the specific sentence regarding bleeding time only claims "mildly prolonged," the preceding sentence indicates "significant increase in bleeding").

      Thank you for this comment. There was no statistically significant difference in cumulative bleeding time between the groups. We have now modified the text accordingly to clarify this point and to indicate that, while bleeding time was not significantly different, blood loss was significantly increased in Shp1/2 DKO mice.

      (4) Figure 2Ci: What was the extent (statistically) of GPVI reduction in the Shp1 and Shp2 single knockout mice compared to the control? It seems that although there was no change in alpha2 expression in the single-knockout conditions, the contributions of Shp1 and Shp2 loss may be additive on GPVI (although I acknowledge that this is not necessarily borne out in Figure 3Ai).

      Thank you for this comment. After reanalyzing the data using an appropriate statistical test (one-way ANOVA followed by Tukey’s post hoc test), we found that GPVI expression is significantly reduced in both Shp1 and Shp2 single knockout platelets compared with controls. However, this reduction did not result in detectable functional consequences on platelet aggregation, as shown in Figure 3Ai.

      (5) Figure 3Ai: It seems that the individual replicates for the Shp1/2 double knockout cluster in two populations, extreme non-responders and arguably normal responders to CRP. Are there any biological or technical explanations for this?

      We thank the reviewer for this observation. We agree that the distribution of individual replicates in the Shp1/2 DKO group suggests the presence of two subpopulations, with some samples showing markedly impaired aggregation and others retaining near-normal responsiveness to CRP. While all experiments were performed under standardized conditions, subtle differences in platelet preparation, agonist sensitivity, or assay timing could also contribute to dispersion within this group. Importantly, despite this variability, the overall trend indicates a significant reduction in aggregation in the Shp1/2 DKO condition compared to controls, supporting a critical and partially redundant role for Shp1 and Shp2 in GPVI-mediated platelet activation.

      (6) "Aberrant functional responses of Shp1/2-deficient platelets": It may be helpful, in the last paragraph of this section, to briefly explain the relationship between P-selectin, CRP, and PAR4p. If short on space/words, the introduction likely does not need an explanation of platelet function and definitions of megakaryopoiesis and thrombopoiesis.

      We thank the reviewer for this suggestion. We have revised the last paragraph to clarify that P-selectin surface expression reflects α-granule secretion following platelet activation. We now specify that CRP activates platelets via GPVI signaling, whereas PAR-4 peptide signals through thrombin receptors, providing context for the differential responses observed in Shp1/2-deficient platelets.

      (7) "Aberrant ITAM signaling in Shp1- and Shp2-deficient platelets": Is there a cartoon figure panel that could be added to clarify how SFK (which, as an aside, is not defined as an acronym), Syk, GPVI, CLEC-2 receptor, Shp1, and Shp2 are interrelated? In addition to the comments left in the public review, I was perplexed by Figure 4Bi, as the band for the Shp1/2 double knockout condition appears to be stronger than the other 3 conditions, but this is not what is depicted in the bar graph on the right.

      We thank the reviewer for this helpful comment. We have now added a schematic cartoon (new Figure 8) to clarify the relationships between SFKs, Syk, GPVI, and the regulatory roles of Shp1 and Shp2. All acronyms, including SFK, are now defined at first mention to improve accessibility.

      Regarding Figure 4Bi, we appreciate this observation. The apparent discrepancy between the representative blot and the quantification reflects variability across experiments. The bar graph represents the average of independent replicates.

      (8) I would also recommend considering reshuffling the panels in Figure 4 so that the 2 assays measuring Syk phosphorylation and the 2 assays measuring Src phosphorylation are next to each other, as opposed to grouped by agonist. They should also be presented in the order of the text, which states that SFK activation was measured via Src before mentioning Syk (but the data are presented in reverse).

      We thank the reviewer for this suggestion. We have reorganized Figure 4 so that the panels measuring Src and Syk phosphorylation are presented together and, in the order, described in the text. The manuscript text has also been updated accordingly to match the revised figure layout.

      (9) GPVI overexpression experiments in these megakaryocytes or, conversely, Syk inhibition in control cells, to reverse or recapitulate the phenotype, respectively, may be additionally informative.

      We thank the reviewer for this suggestion. We agree that modulating GPVI or Syk activity could provide additional mechanistic insight. While these experiments were beyond the scope of the current study, we plan to explore GPVI overexpression and Syk inhibition in follow-up studies to further validate the pathway’s role in the observed phenotype.

      (10) "Reduced Tpo signaling in Shp1/2-deficient MKs": In addition to the comments left in the public review, I would suggest moving this section to after "Defective proplatelet formation and MK maturation in Shp1/2-deficient mice" so that the 2 sets of proplatelet and ploidy data are consecutively presented.

      We thank the reviewer for this helpful suggestion. We have now revised the manuscript accordingly by reorganizing both the text and figures. The ploidy and proplatelet formation data are now presented together in Figure 5, followed by the Tpo signaling data in Figure 6, improving the overall flow and clarity of the results section.

      (11) Figure 6Cii: Why does Shp1 add up to >100%?

      The reason the Shp1 bar exceeds 100% is due to how the data were quantified and normalized. Each segment represents the mean from separate experiments. Stacking these means can exceed 100% because the sum of averages is not equal to the average of the total.

      (12) Figure 7D: How do you reconcile these findings of impaired AKT phosphorylation with the addition of a Shp2 inhibitor but no change with Shp2 knockout (Figure 5C)? Would you attribute it to the residual Shp1 and Shp2 in the Cre-Lox MKs?

      Pharmacological effects do not always correlate with congenital anomalies arising for genetic defects. The Shp2 allosteric inhibitors used in our study only inhibit catalytically inactive Shp2, whereas targeted deletion of Ptpn11 results in a loss of total Shp2 expression, including catalytic and non-catalytic related functions, with developmental consequences. Further, Gp1ba-Cre+; Shp2fl/fl megakaryocytes express approximately 22% of normal Shp2 level, which likely also contributes to differences observed between pharmacological inhibition and genetic ablation of Shp2.

    1. eLife Assessment

      This important study demonstrates how ablation or silencing of hilar mossy cells in the mouse influences the primary location where the mossy cells project, the inner molecular layer of the dentate gyrus. The anatomical findings are convincing and include altered adult-born granule cells and the shrinkage of the inner molecular layer following mossy cell ablation. However, the mechanisms and their functional significance are unclear, so more of these types of experiments/analyses would strengthen the study, especially the support for the broader conclusions.

    2. Reviewer #1 (Public review):

      Summary:

      This study provides valuable evidence that hilar mossy cells play important roles in maintaining the structural organization of the dentate gyrus and regulating the maturation of adult-born granule cells. The evidence for the structural reorganization and for the accelerated dendritic maturation of adult-born granule cells is convincing: it rests on converging anatomical, viral tract-tracing, retroviral birth-dating, and electrophysiological measurements, with appropriate controls for viral spread, off-target CA3 expression, and axonal degeneration. Support for the study's broader interpretive claim - that the dentate circuit functionally compensates for mossy cell loss - is incomplete. That claim rests on two null results obtained under baseline conditions (home-cage cFos and PTZ seizure metrics) in small cohorts, without behavioral assessment and without a stimulus-driven activity readout, and the manuscript does not engage with published work showing that mossy cells regulate neural stem cell activation and are required for stimulus-evoked neurogenic and behavioral responses.

      Strengths:

      (1) The study is technically rigorous and employs multiple complementary approaches, including selective genetic manipulations, viral tracing, immunohistochemistry, retroviral labeling of adult-born neurons, electrophysiology, and anatomical analyses. The comparison between complete mossy cell ablation and chronic synaptic silencing is particularly powerful, allowing the authors to examine the significant role of mossy cells in structural and functional organization in the dentate gyrus.

      (2) One of the most notable findings is the identification of a previously unrecognized collapse of the inner molecular layer following extensive mossy cell ablation. This observation substantially expands current understanding of dentate gyrus structural plasticity. The demonstration that adult-born granule cells undergo accelerated dendritic maturation after both mossy cell loss and silencing also provides important insight into how mossy cells regulate adult neurogenesis.

      Weaknesses:

      (1) The functional significance of the observed structural remodeling remains incompletely addressed. Mossy cells have been strongly implicated in pattern separation, spatial information, and emotional behavior, yet no behavioral analyses were conducted. Consequently, it remains unclear whether the dramatic anatomical changes observed following mossy cell ablation translate into meaningful behavioral alterations.

      (2) The conclusion that the dentate gyrus exhibits remarkable homeostatic compensation is reasonable but remains indirect. Although cFos expression and PTZ-induced seizure susceptibility are unchanged despite altered E:I balance, the mechanisms responsible for maintaining network stability are not investigated. Additional analyses of inhibitory circuit remodeling or compensatory synaptic adaptations would strengthen this conclusion.

    3. Reviewer #2 (Public review):

      Summary:

      The authors examine how hilar mossy cells (MCs) influence adult-born dentate granule cell (abDGC) maturation and dentate gyrus (DG) structural integrity. Using both MC ablation and chronic functional silencing, they find that lacking MC inputs accelerates early abDGC maturation without altering mature cellular or intrinsic properties. MC silencing specifically decreased inner molecular layer (IML) spine density, whereas MC ablation led to IML collapse and an increased E/I ratio. However, neither intervention altered overall network excitability (measured via c-Fos and seizure induction) or seizure thresholds. These results advance our understanding of DG circuit plasticity during neurodegeneration.

      Strengths:

      (1) The side-by-side comparison of ablation vs. silencing provides a clear distinction between structural synapse loss and functional inactivation.

      (2) The multi-level analysis spanning structural anatomy, single-cell physiology, and network-level assays yields a rich, comprehensive dataset.

      Weaknesses:

      (1) Measuring composite E/I ratios without parsing isolated EPSCs and IPSCs limits direct evaluation of MC-driven excitatory inputs. Furthermore, electrical stimulation in the IML likely recruits local interneuron axons directly alongside MC fibers, complicating the attribution of these responses solely to feed-forward MC circuits.

      (2) The dramatic structural reorganization and IML collapse observed following MC ablation make it difficult to attribute changes in the E/I ratio purely to functional synaptic remodeling rather than physical circuit distortion.

      (3) Layer boundary shifts following MC ablation complicate the interpretation of site-specific spine density (Figure 4); without accounting for IML collapse, classifying spine loss purely by traditional layer boundaries rather than proximal vs. distal dendrites may obscure local structural changes.

      (4) The convulsive dosing protocol used for the seizure threshold test lacks the sensitivity required to reveal subtle changes in excitability.

    4. Author response:

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This study provides valuable evidence that hilar mossy cells play important roles in maintaining the structural organization of the dentate gyrus and regulating the maturation of adult-born granule cells. The evidence for the structural reorganization and for the accelerated dendritic maturation of adult-born granule cells is convincing: it rests on converging anatomical, viral tract-tracing, retroviral birth-dating, and electrophysiological measurements, with appropriate controls for viral spread, off-target CA3 expression, and axonal degeneration. Support for the study's broader interpretive claim - that the dentate circuit functionally compensates for mossy cell loss - is incomplete. That claim rests on two null results obtained under baseline conditions (home-cage cFos and PTZ seizure metrics) in small cohorts, without behavioral assessment and without a stimulus-driven activity readout, and the manuscript does not engage with published work showing that mossy cells regulate neural stem cell activation and are required for stimulus-evoked neurogenic and behavioral responses.

      Thank you and we agree with all points. We specifically focused on the structural aspects of dentate rearrangement, the impact of mossy cell loss/silencing on dentate neurogenesis, and dentate function at the circuit level. Although our home-cage cFos and PTZ susceptibility assays are limited in terms of their sensitivity, these assays were chosen to address the role of mossy cells in controlling overall dentate activity levels and seizure susceptibility, and our data demonstrate no dramatic changes in overall activity levels or increased/decreased seizure susceptibility.

      Strengths:

      (1) The study is technically rigorous and employs multiple complementary approaches, including selective genetic manipulations, viral tracing, immunohistochemistry, retroviral labeling of adult-born neurons, electrophysiology, and anatomical analyses. The comparison between complete mossy cell ablation and chronic synaptic silencing is particularly powerful, allowing the authors to examine the significant role of mossy cells in structural and functional organization in the dentate gyrus.

      (2) One of the most notable findings is the identification of a previously unrecognized collapse of the inner molecular layer following extensive mossy cell ablation. This observation substantially expands current understanding of dentate gyrus structural plasticity. The demonstration that adult-born granule cells undergo accelerated dendritic maturation after both mossy cell loss and silencing also provides important insight into how mossy cells regulate adult neurogenesis.

      We were also surprised by the inner molecular layer (IML) collapse, as disease models that produce mossy cell loss often involve granule cell axon (mossy fiber) sprouting (and maintained IML thickness) rather than IML collapse. It is unclear whether axon sprouting, reduced degrees of mossy cell loss, or other signaling pathways drive the differences between our selective ablation and translational disease models. We agree that the differential effects of mossy cell ablation and silencing on adult neurogenesis and proximal spine formation highlight the remarkable plasticity in this circuit and provide insights into both the functional and structural circuit roles of mossy cell inputs.

      Weaknesses:

      (1) The functional significance of the observed structural remodeling remains incompletely addressed. Mossy cells have been strongly implicated in pattern separation, spatial information, and emotional behavior, yet no behavioral analyses were conducted. Consequently, it remains unclear whether the dramatic anatomical changes observed following mossy cell ablation translate into meaningful behavioral alterations.

      Our functional assays were primarily focused on the circuit (synaptic) level, with additional assessment of how mossy cell manipulations affected overall dentate activity levels (as reflected by cFos expression). Our limited behavioral analysis focused on seizures, based on prior foundational work on the roles of mossy cells in seizures/epilepsy. We tested the hypothesis that seizure susceptibility might be markedly changed in the near absence of mossy cells, using a “threshold” dose of PTZ that is just above that required to produce seizures, and which can produce dramatically enhanced seizures in hyperexcitable mice. Alternative seizure assays (dose-response curves, continuous monitoring, different seizure-inducing protocols), measurements of granule cell activity in response to environmental contingencies, and the assessments of the response of the dentate stem cell pool to neurogenesis-enhancing stimuli might absolutely produce further insights into how functional mossy cell inputs control stimulus-related dentate activation and/or neurogenesis. Our resubmitted manuscript will clarify that the preserved basal level of dentate gyrus activity after mossy cell loss does not preclude altered activity-dependent activation in other settings or behavioral/learning changes. This could be uncovered with additional behavioral testing or seizure modeling, and is something that we expect to address in future studies.

      (2) The conclusion that the dentate gyrus exhibits remarkable homeostatic compensation is reasonable but remains indirect. Although cFos expression and PTZ-induced seizure susceptibility are unchanged despite altered E:I balance, the mechanisms responsible for maintaining network stability are not investigated. Additional analyses of inhibitory circuit remodeling or compensatory synaptic adaptations would strengthen this conclusion.

      We believe that there are many potential mechanisms that could explain how the nearly complete loss of a major population of dentate neurons is not accompanied by dramatic changes in overall activity levels. Although a fully comprehensive functional assessment of dentate circuit elements is prohibitive, we will undertake what we believe to be the highest-yield analyses in this regard. We propose to stain tissue for inhibitory circuit markers such as VGAT and PV, to determine whether mossy cell loss alters the density or localization of inhibitory synapses as well as circuit elements involved in feed-forward inhibition. We also plan to perform additional electrophysiological experiments to directly assay whether changes in feed-forward inhibition, overall synaptic inhibition (sIPSCs) and/or tonic inhibition might accompany functional mossy cell loss. We will incorporate the outcomes from these additional assays into a revised manuscript. This will shed light on whether inhibitory circuit remodeling also contributes to compensation after mossy cell loss, and hopefully provide additional insights relevant to translational disease models that involve mossy cell loss.

      Reviewer #2 (Public review):

      Summary:

      The authors examine how hilar mossy cells (MCs) influence adult-born dentate granule cell (abDGC) maturation and dentate gyrus (DG) structural integrity. Using both MC ablation and chronic functional silencing, they find that lacking MC inputs accelerates early abDGC maturation without altering mature cellular or intrinsic properties. MC silencing specifically decreased inner molecular layer (IML) spine density, whereas MC ablation led to IML collapse and an increased E/I ratio. However, neither intervention altered overall network excitability (measured via c-Fos and seizure induction) or seizure thresholds. These results advance our understanding of DG circuit plasticity during neurodegeneration.

      Strengths:

      (1) The side-by-side comparison of ablation vs. silencing provides a clear distinction between structural synapse loss and functional inactivation.

      (2) The multi-level analysis spanning structural anatomy, single-cell physiology, and network-level assays yields a rich, comprehensive dataset.

      Thank you for these positive assessments of our study.

      Weaknesses:

      (1) Measuring composite E/I ratios without parsing isolated EPSCs and IPSCs limits direct evaluation of MC-driven excitatory inputs. Furthermore, electrical stimulation in the IML likely recruits local interneuron axons directly alongside MC fibers, complicating the attribution of these responses solely to feed-forward MC circuits.

      We fully expect electrical stimulation of the proximal molecular layer to directly recruit local interneuron axons in addition to feed-forward inhibition. Thus, our experimental design did not distinguish between directly stimulated and feed-forward inhibitory circuits, and we were only able to conclude that mossy cell loss caused circuit rearrangement without clearly attributing the differences specifically to feed-forward mechanisms. To provide additional insights into the underlying changes, we plan to examine both inhibitory circuit structure (using immunohistochemistry) and function using assays designed to distinguish between directly stimulated vs. feed-forward inhibitory mechanisms, which will be incorporated into the revised manuscript.

      (2) The dramatic structural reorganization and IML collapse observed following MC ablation make it difficult to attribute changes in the E/I ratio purely to functional synaptic remodeling rather than physical circuit distortion.

      We actually consider physical circuit remodeling after MC ablation to be the primary explanation for the E/I ratio changes, in that the proximal translocation of MEC synapses following mossy cell ablation allows them to be electrically stimulated in the proximal molecular layer. Thus, the altered E/I ratio of proximal synapses after MC ablation largely represents the fact that we are stimulating proximal MEC inputs rather than mossy cell inputs (which are now absent). Our MML terminal stain (VGlut2) and MEC viral labeling support this interpretation, which we will clarify in the results and discussion of this data.

      (3) Layer boundary shifts following MC ablation complicate the interpretation of site-specific spine density (Figure 4); without accounting for IML collapse, classifying spine loss purely by traditional layer boundaries rather than proximal vs. distal dendrites may obscure local structural changes.

      We initially kept the classic nomenclature (IML vs OML) to avoid confusion for readers, and defined “IML” vs “OML” spines based on proximity to the inner and outer edges of the molecular layer (the innermost and outermost 40 µm; see Methods). Thus, in the setting of IML collapse after mossy cell ablation, these spines almost certainly occurred in regions innervated by the MEC (formerly “MML”). To avoid obscuring this aspect of the data, we will clarify this in the Results and Figure 4, making the proximal vs. distal designations clear.

      (4) The convulsive dosing protocol used for the seizure threshold test lacks the sensitivity required to reveal subtle changes in excitability.

      Our PTZ dose (40 mg/kg i.p.) is just above a dose (30 mg/kg i.p.) that almost never causes seizures in healthy mice in our hands, making it potentially able to detect seizure resistance. This 40 mg/kg dose causes short, limited seizures with a relatively consistent latency, and in other (unrelated) experiments, mice with genetic hyperexcitability have dramatically increased seizure duration and accelerated seizure onset (and sometimes mortality) at this dose, indicating that it is sensitive to at least some forms of increased seizure susceptibility. That stated, we agree that this single-dose PTZ protocol could miss subtle changes in dentate excitability or seizure susceptibility. These could be unmasked by a more detailed dose-response analysis or by other induced seizure assays; we will clarify this limitation in our manuscript.

    1. eLife Assessment

      This important study proposes a framework in which tutor-song memorization and performance-error computation arise from a shared predictive-cancellation circuit. The approach is solid, but the choice of learning rules, circuit architecture, and sparsity assumptions is not always sufficiently justified, and the model would benefit from a clearer mechanistic derivation and more concrete experimental predictions. This study would be relevant for researchers studying vocal learning and motor learning in general.

    2. Reviewer #1 (Public review):

      Summary:

      This manuscript addresses how internally generated evaluative signals can arise during self-guided learning in the absence of external reward. Using zebra finch song learning as a model system, the authors propose that tutor-song memorization and vocal performance evaluation are not separate processes, but instead emerge from a shared local circuit that learns to predictively cancel tutor-song-related auditory input. The comparison across several candidate circuit architectures, the quantitative comparison to experimental calcium imaging data, and the decomposition of the learned recurrent connectivity into modes shaping the error landscape are all strong aspects of the work. The final demonstration that the learned error signal can guide a downstream reinforcement learning agent also provides a useful proof of principle.

      Strengths:

      The idea that tutor-song memorization and performance evaluation can emerge from a shared predictive-cancellation circuit is interesting, and the combination of circuit modeling, comparison to experimental data, and error-landscape analysis is compelling.

      Weaknesses:

      (1) A central conclusion of the manuscript is that the E→I→E model best matches experimental data. This establishes model fit, but it does not yet explain why E→I and I→E plasticity are important for tutor-song cancellation and error-signal formation. Does E→I plasticity primarily teach the inhibitory population to represent tutor-song-related excitatory activity? Does I→E plasticity then implement the negative image required to cancel expected excitatory responses? Does the closed E/I loop primarily control gain, shift the minimum of the error landscape, or both? A useful analysis would be to compare models in which only E→I synapses are plastic, only I→E synapses are plastic, both are plastic, or neither is plastic.

      (2) The analysis in Fig. 5 does not yet explain how the identified modes arise from the specific E→I/I→E plasticity mechanism. For example, are the landscape modes mainly produced by E/I gain-control dynamics? Are the memory modes related to an inhibitory negative image of the tutor song? Are these modes localized to particular blocks of the recurrent connectivity, such as E→I or I→E weights, or are they distributed across the full network?

      (3) The manuscript emphasizes the emergence of sparse population error codes. However, in Fig. 6, the downstream actor-critic model uses the population mean excitatory activity as a scalar negative reward. This compresses the high-dimensional sparse population response into a single scalar. If the downstream system only uses the mean response, why is a sparse high-dimensional error code functionally important, beyond matching the observed response distribution? Conversely, if the sparse population pattern contains richer information about the direction or structure of vocal errors, how might downstream reinforcement pathways read out this information?

      The manuscript should clarify whether sparsity is proposed to have a functional role in motor learning, or whether it is primarily a biological feature of the evaluative circuit. This point is particularly important because the broader framing of the paper concerns internal evaluative signals, whereas the final reinforcement learning demonstration uses a scalar reward.

      (4) The actor-critic model in Fig. 6 is useful because it demonstrates that the learned error signal contains enough information to guide motor learning. However, the reinforcement learning module is attached downstream of the auditory circuit and is highly simplified. Therefore, it remains somewhat ambiguous whether Fig. 6 should be interpreted as a circuit model of song learning or as a demonstration of sufficiency. The latter interpretation seems appropriate and valuable, but the manuscript should state this more explicitly. The central contribution appears to be the bootstrapping of an internal evaluative signal, rather than a complete model of sensorimotor song learning. Clarifying this distinction would prevent overinterpretation of the actor-critic results.

    3. Reviewer #2 (Public review):

      Summary:

      The paper proposes a network model that explains how birdsong learning can be guided by reinforcement signals.

      Strengths:

      It is well known that self-generated motor actions typically suppress their associated sensory input (for example, in the mammalian auditory cortex; see Eliades & Wang, 2003). This study presents a mechanism that effectively reverses this process. The theory posits that, initially, the motor signal generated in HVC, although not yet sufficient to produce an accurate song, nevertheless sends an efference copy to auditory areas, where it acts to cancel external auditory input from the tutor. This establishes a "scaffold," such that only an accurate replica of the tutor song can successfully suppress the corresponding auditory activity.

      During learning, poorly generated plastic songs produce residual auditory activity that cannot be fully suppressed. This remaining activity then serves as an error signal that guides the refinement of motor output. The idea is elegant and is supported by experimental evidence.

      Weaknesses:

      The authors compare several possible sites of synaptic plasticity within the auditory network and conclude that the E-to-I-to-E model provides the best fit to the existing data. In this model, the auditory network consists of recurrent excitatory (E) and inhibitory (I) neurons, and Hebbian plasticity at E-to-I and I-to-E synapses is required to establish the cancellation pattern necessary to reproduce the tutor song.

      However, the manuscript's presentation of the underlying plasticity mechanisms is somewhat puzzling. The authors repeatedly emphasize anti-Hebbian learning, even though their most successful model fundamentally relies on Hebbian plasticity. Although the resulting functional relationship may be described as anti-Hebbian, the biological learning mechanism implemented in the model is Hebbian. The repeated emphasis on anti-Hebbian learning therefore distracts from the central message and may confuse readers about the actual mechanism responsible for learning.

      This emphasis may reflect an effort to distinguish the present work from previous anti-Hebbian models, but I suggest restructuring the manuscript. The authors should first present the optimal E-to-I-to-E model in detail, clearly explaining its mechanism and biological interpretation. Subsequent sections could then compare this model with the less successful alternative architectures. Such a reorganization would substantially improve the clarity and overall structure of the manuscript.

      Finally, the abstract presents self-guided reinforcement learning as a novel concept, although this general idea has been described in previous work (e.g., Fiete et al., 2007). The abstract should therefore be revised to more precisely identify the specific novelty and contribution of the present study, rather than attributing novelty to the broader concept of self-guided reinforcement learning.

    1. eLife Assessment

      This valuable study analyzes how recurrent neural network can flexibly switch between multiple tasks. The evidence is solid, but reviewers have raised questions about the impacts of transient dynamics, whether the activity is actually chaotic, can inputs be used to switch between tasks, what determines overlap between tasks and others. In addition, there were more minor questions pertaining to whether the analyses are actually analytic or primarily based on numerical simulations, as well as the relationship to prior studies where the same network was optimized to perform multiple tasks simultaneously with different readouts.

    2. Reviewer #1 (Public review):

      Summary:

      Marschall et al. develop a theoretical framework for analyzing multi-task dynamics in nonlinear recurrent neural networks (RNNs). In the RNN model, recurrent connectivity is a linear superposition of multiple non-overlapping low-rank components, each corresponding to a separate "task". Each task is an autonomous dynamical system that does not incorporate external inputs. The "multi-task computation" setting examines whether multiple tasks (dynamical systems) can operate concurrently or how a network can switch between tasks.

      Within this framework, the authors show that when connectivity consists of two low-rank components implementing two tasks (a limit cycle and a bistable attractor), the network exhibits winner-takes-all dynamics, resulting in only one task being active and the other one suppressed. A task consistently dominates this competition when the magnitude of its low-rank connectivity ("task strength") exceeds that of the other task. When one task is dominant and many other tasks with weaker low-rank connectivity are present, increasing the number of tasks destabilizes the dominant-task dynamics, leading to chaotic fluctuations in network activity. Supported by the dynamical mean-field theory analysis, the authors show that as the dominant-task strength increases, the network transitions through three dynamical regimes: chaotic spontaneous activity, chaotic task-selected dynamics, and non-chaotic task-selected dynamics. The theoretical analysis additionally predicts how latent task-related dynamics manifest in single-neuron activity and how the dimensionality of population activity changes across the three dynamical regimes.

      The results are interesting, the analyses and simulations are rigorous, and the text is clear and easy to follow. Overall, this study is a significant and timely contribution to the literature on low-rank RNNs, an influential model class for low-dimensional neural dynamics in computational neuroscience.

      Main comments:

      (1) In this modeling framework, only one dominant task can be selected while all other tasks are suppressed. In contrast, several previous studies constructed RNNs (either through gradient-descent optimization or reservoir computing) that simultaneously generate outputs for multiple tasks across the corresponding task-specific readouts. Of course, what counts as a task is arbitrary, and one could view the dynamics of a reservoir network as implementing a single high-dimensional "task" with multiple readouts. Nevertheless, it would be helpful to explicitly clarify the distinction and similarities between the current modeling framework and networks that simultaneously solve multiple tasks.

      (2) The role of external inputs in task selection appears to be underdeveloped. It is only briefly examined in Fig. S4, with the conclusion that external inputs aligned with the task subspace cannot enable selection of the desired task. However, previous multi-task RNN models (e.g., optimized through gradient descent) are clearly able to switch across many tasks using external inputs. In these models, external inputs modulate RNN activity along specific directions, shaped through gradient descent, to select the relevant task for each input. In contrast, this study only considers inputs aligned with the m-direction (left loading vector in the low-rank connectivity for a task). Such input cannot enhance the corresponding task's activity through recurrent amplification (Fig. S4). Yet, low-rank RNN theory predicts that inputs aligned with the n-direction (task's right loading vector) are selectively amplified by the recurrent dynamics. Why inputs aligned with the n-direction were not examined? More broadly, focusing only on inputs aligned with either m- or n-direction appears too narrow, as trained multi-task RNNs indicate that task selection through external inputs is possible, but may require input directions different from m and n.

      (3) It is unclear what the reason is for the task interference: overlaps between task loading vectors for different tasks or other nonlinear effects? This issue is especially prominent in the analysis of task capacity (Fig. 2D and Fig. 4A). The text states that the loading vectors are independent across tasks, i.e. there is zero expected overlap between loading vectors of different tasks (line 126). For a network of N neurons, a rank-R task requires 2R loading vectors. Under the assumption of independence, the maximum number of tasks is P = N/(2R). Then α=P/N can be at most 1/(2R). In the simulations, it is chosen R=2, such that alpha can be at most 1/4 for the loading vectors to remain independent. However, the range of alpha reaches up to 1 in Fig. 2D and Fig. 4A, suggesting that loading vectors are no longer independent across tasks for larger alpha. Is the linear dependence between loading vectors (i.e. overlap across tasks) the reason for the task-1 component norm to drop sharply around alpha=1/4 in Fig. 2D? More broadly, is it possible to isolate the contribution of overlap in task loading vectors versus other nonlinear effects?

      (4) In the current version of the paper, it is difficult to understand how the analysis based on the dynamical mean-field theory in Methods explains the key observations from the numerical simulations. For example, the mechanism underlying the transition from the spontaneous state to chaotic task-selected state, and to the non-chaotic task-selected state remain opaque. Further, it would be helpful to specify which section in Methods is being referred to at each mention throughout the main text. Finally, Fig. 7 and Eqs. (11-13) clearly state the input sources that drive single unit activity, non-dominant task latent states and dominant task latent state. The decomposition is potentially very informative, but its implications are not discussed in sufficient detail. It would be helpful to provide an intuitive explanation of how contributions of these input sources evolve as connectivity strength of one task increases, and which of them eventually leads to the loss of stability of the previous network state.

      (5) The text states that the results can be easily extended to include task-specific inputs and outputs (lines 117-119). However, such an extension does not seem to be straightforward and requires additional explanations. The dynamical mean-field theory analyses here are stationary and describe the steady-state of network dynamics. In contrast, common input-output tasks typically involve transient dynamics, in which time-dependent external inputs keep changing the RNN flow field, and steady-state is never reached. Under these transient conditions, it is unclear whether the same conclusions apply. For example, if a network is at a low-activity baseline when a task-input begins to drive activity in the corresponding task subspace, it is unclear whether the activity in other task-subspaces would grow sufficiently fast to cause interference, or whether such interference would not be observed. Thus, a more detail analyses are necessary to support the extension of the results to input-driven transient tasks beyond autonomous dynamical systems.

      (6) When task strength is the same for all tasks, what determines which task will win the competition? Is it frozen noise in connectivity such that one task always wins, or do initial conditions determine which task wins, based on which task's activity grows faster?

      (7) Does the theory require the activity of all neurons to operate in the saturating part of nonlinearity? For example, the text states "increased activity reduces the gain factor <Φ'(t)>" (line 169). This statement is only true when Φ'<0. For sigmoid nonlinearity used in the paper, Φ'>0 when firing rate is small. If a substantial fraction of neurons in the network is near the rest state, would the theory still apply? Similarly, this statement does not hold for ReLU nonlinearity, and it is unclear how the theory applies to ReLU networks in Fig. S3. The paper states that the results are not specific to the choice of nonlinearity (lines 176-178). However, the dynamics being studied (bistability and limit cycles) both operate on the saturating part of the nonlinearity. Could the authors clarify the assumptions on the nonlinearity for the theory to apply?

      (8) Is it possible to interpret the results in Fig. 4? What does this dependence on the overlap matrix mean? Is there an intuition for this particular dependence, or is it just an observation without general interpretation?

      (9) The results in Fig. 5E appear underdeveloped and somewhat arbitrary. It is unclear how the dependence of dimensionality on the recording time would change as a function of time spent in a task. If this time is long, then the curve grows slowly and total dimensionality is high. If this time is very short, then the network may not have sufficient time for all task variables to grow sufficiently large to contribute significantly to the total variance. Thus, the grows may be faster and the total variance may saturate at a lower value. Hence, it is unclear whether there will be always a qualitative difference from the spontaneous activity curve. Furthermore, since only one of two curves is measured, what quantitative criteria should be used to determine whether it is consistent with task switching or spontaneous state?

      (10) On line 368: "For sufficiently large number of tasks, the dimensionality associated with sequential task selection can greatly exceed that of the spontaneous state (Fig. 5E inset)" - it seems that Fig. 5E inset shows the opposite that the dimensionality of spontaneous state can saturate at a very high value for large N, exceeding the dimension of task-switching network in the main plot. Although it is hard to say, since the inset has many lines with only two labels and no ticks on axis, so it is unclear what exactly does it show.

      (11) Related to discussion on line 368-370: In a task-switching state, would the switching between tasks also be reflected in behavior? In addition, the time-correlation functions would not be stationary in task-switching state, i.e. they would change over time, whereas they will be stationary in the spontaneous state. Thus, could the two mechanisms be dissociated in experiments using behavior or metrics beyond dimensionality?

      (12) On line 385, could the authors provide more details on how they envision the two potential mechanisms-synaptic plasticity and targeted neuromodulation-to reinforce a task-specific low-rank connectivity pattern? If neuromodulation changes the gain of individual neurons, this modulation corresponds to scaling of connectivity by a diagonal matrix, not strengthening of a specific low-rank component. Short-term facilitation or depression also modulates synaptic strength depending on the activity of the presynaptic neuron, thus also scaling connectivity by a diagonal matrix. It is unclear how these two mechanisms could produce a specific low-rank modulation.

    3. Reviewer #2 (Public review):

      Summary

      The authors ask what recurrent connectivity supports many distinct task-related manifolds when the associated dynamics interfere, how a circuit engages one task while suppressing others, and what produces high-dimensional activity. Extending previous theoretical studies on low-dimensional dynamics in large networks, they use a solvable model whose weight matrix is a weighted sum of many low-rank, task-specific components and develop a dynamical mean-field theory that relates connectivity, dynamics, and measurable population signatures of multi-tasking.

      Strengths

      (1) The question is timely. Low-rank networks are a leading model for low-dimensional latent dynamics, and the composition of dynamical systems has been proposed as a mechanism allowing for rapid, flexible learning; the paper connects these two ideas under a single theoretical framework.

      (2) The proposal that sequential transitions between low-dimensional, low-rank dynamics can account for the _apparent growth of dimensionality with recording time_ is novel and is the paper's most valuable conceptual contribution.

      (3) The mathematical analysis is rigorous, and the spontaneous-state theory is convincingly validated against simulation.

      (4) The model produces concrete, falsifiable predictions - heterogeneous, syllable-dependent single-neuron tuning, low within-state dimensionality despite single-neuron variability, and distinct dimensionality-versus-recording-time signatures for the spontaneous versus task-switching accounts.

      Weaknesses - whether the claims are supported by the data

      (1) Chaos is named but not demonstrated._ The large-P and intermediate task-selected regimes are labeled "chaotic," but the manuscript does not establish chaos. In a homogeneous network, it is known that once the fixed point loses stability, the surviving solution is chaotic (Sompolinsky, Crisanti & Sommers 1988); that guarantee does not transfer here. The DMFT noise term is not computed analytically, and the single-neuron correlation functions (Fig. 5) show disorder, not a demonstrated decay of the fluctuation autocorrelation to zero, nor a positive largest Lyapunov exponent. The concern is sharpened by the possibility of _transient_ chaos: orthogonal to a dominant limit cycle, fluctuations may be locally unstable only at certain amplitudes or phases, so the global attractor could remain a stable cycle visited with chaotic excursions. As it stands, the claim of chaos in the intermediate regime is unsupported; it may well hold for some range of the selected-task strength, but this is neither shown numerically nor proven.

      (2) "Analytical theory" overstates what is solved in closed form._ For the task-selected state, the kernels are non-stationary: The DMFT is entrained to the dominant task's dynamics, with an O(1) time-dependent quantity inside the nonlinearity. To my knowledge, there is no closed-form DMFT solution under these conditions. The Methods section supports this, explaining that the general scheme is solved by iterative numerical self-consistency (and described there as prohibitively expensive), and tractability is recovered only in a special block-Haar ensemble with Gaussian currents. This is entirely reasonable, but the main text presents it as an analytical theory; the reliance on numerical solutions of the self-consistency equations should be stated plainly.

      (3) The spontaneous-state transition is the classical critical-gain transition, only reparametrized._ The onset of the no-task-dominant state is governed by $g_{eff}^2 = \alpha R\langle D^2\rangle$. It appears to depend on the number of tasks only because per-task strength D is held fixed as tasks accumulate; under a normalization that holds g_eff fixed, the transition reduces to a critical-gain point independent of P, as in extensive random networks. Relatedly, the result that chaos "arises solely from learning many tasks" is, mechanistically, random-network chaos: the random task components raise the weight variance and play the role of effective disorder. This is a legitimate and appealing reframing, but it is not a new transition, and the manuscript should make the relationship to the standard criterion explicit.

      (4) Significance of the selection mechanism._ That boosting a task's gain selects it is intuitive, and the authors note the extreme (one $D^\mu$ dominating) is trivial. The non-trivial and genuinely useful contribution is quantitative - that only a small, O(1/P) modulation near criticality is required. This deserves to be foregrounded rather than left to the Discussion.

  2. Sep 2026
    1. eLife Assessment

      The manuscript presents a compelling application of the microfluidic mother machine as a "selection" machine. While individual elements of the work have existed for many years across various research communities, the simplicity and elegance of this particular application should make it readily implementable and therefore a useful contribution. The proposed plans for revision are sound, particularly in light of the senior author's stated logistical constraints.

      [Editors' note: this paper was reviewed by Review Commons.]

    2. Reviewer #1 (Public review):

      Summary:

      The authors present MiPS, a platform combining DMD-based patterned illumination, automated microscopy, retrained DeLTA segmentation, and mother-machine microfluidics to selectively inhibit or eliminate cells based on dynamic phenotypes. The system enables targeted UV or red-light illumination in real time using segmentation-informed projection masks, allowing selective enrichment directly within mother-machine devices. The manuscript demonstrates proof-of-concept enrichment of mCherry cells from mixed GFP/mCherry populations, characterizes off-target effects, and performs computational simulations of iterative enrichment rounds. Overall, the engineering and systems integration are impressive, and the platform has strong potential for applications in directed evolution, biosensor optimization, and dynamic phenotype-based selection workflows.

      Overall, I believe the work is suitable for publication after minor revisions and clarification of several aspects of the manuscript. In particular, the paper would benefit from additional context in the Introduction and Methods sections, clearer positioning relative to existing platforms, improved figure readability/captions, and a more careful revision of the English throughout the manuscript.

      Major comments:

      (1) The manuscript should better position MiPS relative to recent microscopy-based and DMD-enabled selection/control systems, particularly Lugagne et al., Nature Communications (2024), DOI: 10.1038/s41467-024-46361-1. That work also combines mother-machine microfluidics, DeLTA-based real-time image analysis, and DMD projection. The key distinction here appears to be physical selection/enrichment through targeted killing rather than optogenetic control, and this difference should be stated more explicitly.

      (2) The manuscript currently compares MiPS mostly to FACS/MACS. However, the more relevant comparison may be recent image-based and microfluidic photoselection systems. A dedicated comparison table discussing throughput, temporal phenotyping, iterative selection, dynamic phenotype tracking, and enrichment capabilities would strengthen the paper.

      (3) The enrichment experiment in Figure 4 represents a relatively simple classification problem (GFP vs mCherry). Since the proposed applications involve subtle continuous phenotypes, it would considerably strengthen the manuscript to include at least one experiment selecting for high vs. low expressors within a single fluorescent reporter population.

      (4) The strongest enrichment result (~170-fold enrichment in Figure 5) is entirely simulation-based. Since the manuscript already states that ~45 min is sufficient between rounds for growth evaluation, a real 2-3-round enrichment experiment seems feasible and would substantially strengthen the platform's practical relevance. This experiment appears realistic within a relatively short time investment.

      (5) The bimodal distributions in Figure 2 suggest that a fraction of cells may be stress-resistant rather than simply surviving randomly. It would be useful to discuss whether repeated rounds could progressively enrich UV-resistant subpopulations.

      (6) The manuscript repeatedly uses the term "killed," although the data shown in Figures 2 and 4 mostly demonstrate strong growth arrest/inhibition. Please clarify how the cutoff of division rate <0.4 h⁻¹ was selected and whether an independent viability assay was performed.

      (7) The off-target analysis in Figure 3 is one of the strongest parts of the paper and should probably be emphasized more. The conclusion that the dominant effects are global rather than local is interesting, but additional discussion about optical scattering, ROS diffusion, or device-wide coupling effects would strengthen the interpretation.

      (8) UV exposure is inherently mutagenic in E. coli, and untargeted cells still receive a substantial fraction of the UV dose at high targeting fractions. Please discuss whether the MB/red-light modality may be preferable in applications where preserving genotype integrity is important.

      (9) The manuscript discusses that methylene blue (MB) improves the on:off target ratio, but MB also appears to reduce baseline growth by ~40% even without red-light exposure. This is potentially important for iterative selection workflows. Please discuss whether this effect is reversible after washout and how rapidly cells recover.

      (10) The manuscript states that the retrained DeLTA model used ~3,000 annotated fluorescence images, but no train/validation/test split or segmentation performance metrics are reported. Since segmentation directly impacts phenotype classification and projection targeting, these details are important for reproducibility.

      (11) The manuscript would benefit from a stronger Methods description regarding DMD calibration, alignment procedures, projection accuracy validation, and computational timing requirements for the real-time analysis pipeline.

      Significance:

      General assessment: This is a creative and technically impressive study that combines mother-machine microfluidics, automated microscopy, real-time image analysis, and DMD-based photoselection into a unified platform for dynamic, phenotype-based enrichment. The strongest aspects of the work are the systems integration, the quantitative characterization of off-target effects, and the conceptual demonstration that dynamic microscopy-derived phenotypes can be linked to physical enrichment workflows.

      The main limitations are that the biological validation remains largely proof-of-concept and the most compelling enrichment results are currently simulation-based rather than experimentally demonstrated across multiple rounds. In addition, the manuscript would benefit from stronger positioning relative to recent image-based and DMD-enabled microfluidic control systems.

      Advance: The study extends the field of single-cell microfluidics and image-based selection by introducing a platform that links longitudinal microscopy measurements directly to physical enrichment decisions within mother-machine devices. To my knowledge, the combination of iterative feedback-driven selection, DMD-based targeted elimination, and dynamic phenotype tracking in this context is novel.

      The closest related systems appear to be recent DMD-enabled mother-machine platforms for real-time optogenetic control, particularly those reported by Lugagne et al. (Nature Communications 2024, DOI: 10.1038/s41467-024-46361-1). However, MiPS introduces a distinct conceptual advance by using patterned illumination for selective enrichment/elimination rather than gene-expression modulation alone.

      The advance is primarily technical and conceptual, with potential downstream applications in directed evolution, synthetic biology, biosensor engineering, and dynamic phenotype screening workflows that are difficult or impossible to implement using FACS alone.

      Audience: The work will likely be of strongest interest to researchers working in synthetic biology, microfluidics, single-cell analysis, systems biology, bioengineering, and automated microscopy. It may also be of broader interest to communities developing dynamic phenotype screening technologies, closed-loop biological control systems, and next-generation directed evolution platforms.

      The audience is likely specialized but multidisciplinary, spanning both engineering-oriented and biology-oriented researchers. The methods and conceptual framework may also influence future development of automated selection systems beyond the specific mother-machine context.

      Expertise - My expertise includes: Microfluidics, Synthetic biology, Single-cell systems, Automated microscopy, Real-time image analysis, Bioengineering platforms, Dynamic phenotype characterization.

    3. Reviewer #2 (Public review):

      Summary:

      In this manuscript, the authors reported Microscopic PhotoSelection (MiPS), a closed-loop automated robotic platform designed to link time-resolved imaging with physical sample recovery in mother machine microfluidic devices. By pairing a standard mother machine layout with a custom DMD optical path, an LED array, and an optimized DeLTA deep-learning model, the system tracks dynamic single-cell phenotypes and isolates specific cells via automated, targeted phototoxicity, i.e. selection by elimination. This is a novel technical development that addresses a clear limitation of snapshot sorting methods like FACS or MACS when screening for time-resolved, lineage-dependent traits. However, several methodological limitations and presentation errors must be addressed before publication.

      Major Comments:

      (1) Definition of 'Optimal' Dose (Figure 2D): The authors identify 8.0 W*cm-2 UV light for 300s as the optimal condition. However, this data point lies at the absolute boundary of the tested parameter space. In classical dose-response characterization, an optimum is defined by a local peak or a plateau followed by a decline in performance (typically due to rising off-target toxicity or scatter). Because the performance curve has not rolled over, this represents a boundary condition rather than a demonstrated mathematical optimum. The authors should either extend the parameter sweep to locate the true peak or soften their language to reflect that this is simply the highest performing condition tested.

      (2) UV Exposure Time Gap: The exposure time sweep skips directly from 60s to 300s. While the closely spaced early timepoints are appropriate for capturing initial cell-death kinetics, the large gap to 300s leaves a significant engineering blind spot. Figure 3D demonstrates that off-target scattering damage scales linearly with cumulative light energy. If complete target cell arrest can be achieved at an intermediate exposure (e.g., 120s, 180s or 240s), operating the system at 300s unnecessarily subjects neighboring "surviving" cells to secondary global UV stress via device-wide scattering. An intermediate temporal sweep is recommended to optimize the selection window and properly balance target lethality with background library viability.

      (3) Baseline Chemical Toxicity of Methylene Blue (MB): The photosensitizer workflow shows a clear improvement in contrast at lower power densities and exposure times. However, lines 151-153 note that the addition of 2 uM MB alone, even without light activation, stunts the baseline bacterial growth rate by ~40%. This is a major biological confounder. For applications like directed evolution or dynamic physiological screening, introducing a chemical stressor that nearly halves fitness imposes an unintended selective pressure. This baseline stress may activate pathways that mask or alter the phenotypes of interest. The authors must expand their discussion on how this baseline toxicity impacts multi-round iterative selections, and should ideally evaluate lower concentrations (e.g., 0.5uM or 1uM) or alternative photosensitizers to identify a more viable operational window.

      (4) Negative Selection Framework and Search Space Scale: The MiPS platform relies entirely on negative selection by destroying unwanted variants. While effective for the demonstrated 1:1 binary proof-of-concept mixture, negative selection scales poorly when screening for rare variants within large libraries. For instance, isolating a single high performer from a library of 105 cells requires the system to successfully target and kill 99,999 individual cells; any statistical leak or failure in killing efficiency directly leads to heavy contamination of the recovered sample. The Discussion section requires a quantitative evaluation of these search space constraints, outlining how they limit the system's utility compared to positive selection mechanisms (such as optical tweezers or droplet sorters) when scaling to rare mutations (<1 in 104).

      Significance:

      This study presents a significant methodological advance in single-cell analysis and microfluidics by integrating long-term live-cell imaging, automated image analysis, and phenotype-guided cell recovery into a closed-loop platform. Existing approaches such as FACS and MACS are largely limited to endpoint or snapshot measurements, whereas MiPS enables selection based on dynamic and lineage-dependent cellular behaviors, thereby addressing an important gap in current single-cell screening technologies.

      A key strength is the effective integration of mother machine microfluidics, custom optics, and deep-learning-based tracking into an automated and functional system. While the individual components are established, their combination into a phenotype-driven selection platform is innovative and expands the utility of live-cell microscopy from passive observation to active cell selection. The advance is therefore primarily methodological and technological, with potential to enable future conceptual discoveries in cellular heterogeneity and lineage dynamics.

      However, limitations remain regarding scalability, robustness, selection accuracy, and generalizability across biological systems. Additional benchmarking and validation would strengthen the work further.

      Overall, the study will be of interest to researchers in microfluidics, single-cell biology, microbial systems biology, bioengineering, quantitative imaging, and synthetic biology.

      My expertise is in microfluidics, cell sorting and disease mechanobiology.

    4. Reviewer #3 (Public review):

      Summary:

      The work describes an optofluidic automation setup to optically inhibit and enrich selected bacterial populations in confined microchannels through negative selection using light stimulation. The work is well described and the manuscript is well constructed.

      Major comment:

      The authors reported that methylene blue with 2uM incubation has superior performance than UV light. But it's also noted on line 152 there is an inhibition effect from the chemical affecting ~40% of the growth rate.

      It will be noteworthy what is the growth curve or at least the MIC of methylene blue used on the MG1655 E. coli by the authors.

      Significance:

      The optics part of the work is well described, however the materials and methods details of the biological and microfluidic part can be extended.

      Overall, the system demonstrated the practical use of combining microfluidics for enrichment of microbial population as a novel alternative method, despite that the efficiency is currently subpar to conventional methods.

      But combining further with deep learning phenotype or growth rate monitoring, the technology represents a new path for phenotypic selection which is also novel that conventional methods cannot offer. The work will benefit readers in applied science seeking for new target enrichment based on optofluidics.

    1. eLife Assessment

      This study uses convincing analyses of publicly available connectomic and transcriptomic datasets to survey the anatomy and connectivity of neurosecretory cells in the Drosophila brain. The systems-level description of the neurosecretory system is complemented by selected functional experiments and predictions for circuit function and paracrine signaling networks that can now be tested in future studies. This important work will be of interest to neuroscientists working on hormonal signaling in Drosophila and other animals.

    2. Reviewer #1 (Public review):

      Summary:

      The study by McKim et al (eLife-RP-RA-2024-102684) seeks to provide a comprehensive description of the connectivity of neurosecretory cells (NSCs) using a high-resolution electron microscopy dataset of the fly brain and several single cell RNA seq transcriptomic datasets from the brain and peripheral tissues of the fly. They use connectomic analyses to identify discrete functional subgroups of NSCs and describe both the broad architecture of the synaptic inputs to these subgroups as well as some of the specific inputs including from chemosensory pathways. They then demonstrate that NSCs have very few traditional presynapses consistent with their known function as providing paracrine release of neuropeptides. Acknowledging that EM datasets can't account for paracrine release, the authors use several scRNAseq datasets to explore signaling between NSCs and characterize widespread patterns of neuropeptide receptor expression across the brain and several body tissues. The thoroughness of this study allows it to largely achieve its goal and provides a useful resource for anyone studying neurohormonal signaling.

      Strengths:

      The strengths of this study are the thorough nature of the approach and the integration of several large-scale datasets to address shortcomings of individual datasets. The study also acknowledges the limitations that inherent to studying hormonal signaling and provide interpretations within the context of these limitations.

    3. Reviewer #2 (Public review):

      Summary:

      The authors provide a comprehensive description of the neurosecretory network in the adult Drosophila brain. They assigned and verified the types of neurosecretory cells (NSCs) found in three publicly available drosophila brain connectomes. They then describe the organization of synaptic inputs and outputs for across NSC types. They show that NSCs are regulated by multiple sensory modalities, including enteric neurons. The authors then focus on a concise pathway from corazonin-expressing NSCs to a set of descending neurons, DNg27 and demonstrate that this pathway has the capacity to regulate egg-laying in female flies. Leveraging existing transcriptomic data, they also describe the hormone and receptor expressions in the NSCs and show putative paracrine signaling between NSCs. Taken together, this study provides a framework for future functional experiments, which may demonstrate whether and how NSCs, and the circuits to which they belong, shape physiological function and behavior.

      Strengths:

      This study uses three Drosophila brain connectomes to assign cell types to ten classes of neurosecretory cells (NSCs), based on clustering of synaptic connectivity and morphological features. The authors then verify type assignments for selected populations by matching cluster sizes to anatomical localization and cell counts using immunohistochemistry of neuropeptide expression and markers with known co-expression.

      The authors compare their findings to previous work describing the synaptic connectivity of the neurosecretory network in larval Drosophila (Huckesfeld et al., 2021), finding that there are some differences between these developmental stages. Direct comparisons between adult and larvae are made possible through direct comparison in Table 1, as well as the authors' choice to adopt similar (or equivalent) analyses and data visualizations in the present paper's figures.

      The authors extract core themes in NSC synaptic connectivity and generate predictions regarding sensory inputs and downstream physiological and behavioral functions. They test one newly identified NSC-premotor pathway, from corazonin-expressing NSCs to the descending neuron DNg27, with loss-of-function experiments and demonstrate that this pathway has the capacity to regulate female egg-laying.

      The authors illustrate expression patterns of neuropeptides and receptors across NSC cell types from existing transcriptomic data and present a putative paracrine signaling network among NSCs. The authors also catalog hormone receptor expression across tissues.

      Taken together, this study provides a comprehensive account of the neurosecretory system of the adult fly.

      Weaknesses:

      In Figure 6 authors use a linear dynamical modeling approach (described in Bates et al. 2026) to quantify the influence of different sensory source neuron types on the different NSC classes. The authors should discuss the two main assumptions baked into this approach: 1) all path segments (connections) from sources to targets are given the same sign and therefore result in activation, despite likely biological variation in their synaptic valences. 2) Each connection is given the same time constant for the response kinetics. Therefore, the model assumes uniform intrinsic "biophysical" properties.

      Although the actual intrinsic properties (e.g. complements of voltage-gated ion channels) of the intermediate and target neurons are unknown, they are likely heterogenous. Such heterogeneity would have consequences on the steady-state responses. Thus, the response magnitudes measured in this model are unlikely to provide an accurate representation of feedforward "influences" in this circuit.

      Although the intrinsic properties of all nodes in these paths will remain unknown in the absence of electrophysiological recordings, one could still consider the signs of connections using neurotransmitter predictions in the connectome (Eckstein et al. 2024). It would then be useful to compare the relative influences calculated with the Bates et al. approach to 1) simple weight propagation methods which are agnostic to time (as in Hoeller et al. 2026; doi: https://doi.org/10.64898/2025.12.22.696097) and 2) this Bates et al. approach and weight propagation methods that conserve the signs of the connections.

      In Figure 8 and associated supplements, the authors probe the function of CRZ-expressing NSCs > DNg27 pathways in female and male flies. Although the authors test the effects of silencing both CRZ-expressing cells and DNg27 on feeding, egg-laying, and flight behaviors in females. They recapitulate a previous finding that CRZ-expressing cells regulate feeding behavior and then identify potential regulatory roles for this pathway in egg-laying. However, the authors did not test this full palette of behaviors in males. The authors do not test feeding or flight behaviors in males. They do, however, confirm previously reported activation phenotypes (copulation-like behaviors), via optogenetic activation of CRZ-expressing cells in males. These experiments would be more ethological if executed in freely walking male flies, rather than males that were glued, on their backs. It is unclear why the authors did not also test for activation or loss-of-function phenotypes for DNg27 in males. Taken together: the authors show compelling loss-of-function phenotypes for feeding and egg-laying for the CRZ-expressing NSC > DNg27 pathway in females, but evaluation in males remains incomplete.

    4. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      This valuable study uses the analysis of connectomic and transcriptomic datasets to survey the anatomy and connectivity of neurosecretory cells in the Drosophila brain. While the connectivity analyses are convincing, the anatomical and functional data provided to verify cell type identity and paracrine signaling is incomplete. Once these aspects are improved, this study would be of interest to neuroscientists working on hormonal signaling in Drosophila and other animals.

      We thank the editor and reviewers for their assessment of our manuscript. We hope that the additional results in the revised manuscript addresses all of the concerns.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The study by McKim et al seeks to provide a comprehensive description of the connectivity of neurosecretory cells (NSCs) using a high-resolution electron microscopy dataset of the fly brain and several single-cell RNA seq transcriptomic datasets from the brain and peripheral tissues of the fly. They use connectomic analyses to identify discrete functional subgroups of NSCs and describe both the broad architecture of the synaptic inputs to these subgroups as well as some of the specific inputs including from chemosensory pathways. They then demonstrate that NSCs have very few traditional presynapses consistent with their known function as providing paracrine release of neuropeptides. Acknowledging that EM datasets can't account for paracrine release, the authors use several scRNAseq datasets to explore signaling between NSCs and characterize widespread patterns of neuropeptide receptor expression across the brain and several body tissues. The thoroughness of this study allows it to largely achieve it's goal and provides a useful resource for anyone studying neurohormonal signaling.

      Strengths:

      The strengths of this study are the thorough nature of the approach and the integration of several large-scale datasets to address short-comings of individual datasets. The study also acknowledges the limitations that are inherent to studying hormonal signaling and provides interpretations within the context of these limitations.

      We thank this reviewer for the thorough assessment and highlighting the strengths of our manuscript. Based on comments from the other reviewer, we now include additional analyses of NSCs from two new recent datasets – the brain and nerve cord (BANC) connectome and the male central nervous system (maleCNS) connectome. Our original conclusions based on the FlyWire connectome remain unchanged, further validating our analyses.

      Weaknesses:

      Overall, the framing of this paper needs to be shifted from statements of what was done to what was found. Each subsection, and the narrative within each, is framed on topics such as "synaptic output pathways from NSC" when there are clear and impactful findings such as "NSCs have sparse synaptic output". Framing the manuscript in this way allows the reader to identify broad takeaways that are applicable to other model system. Otherwise, the manuscript risks being encyclopedic in nature. An overall synthesis of the results would help provide the larger context within which this study falls.

      We agree with the reviewer and have modified the subsection titles to highlight the main findings within those sections.

      We have also included a figure (new Figure 10) which summarizes the main findings from our manuscript and places them within the larger context of neuroendocrine signaling in adult Drosophila in relation to other studies.

      The cartoon schematic in Figure 5A (which is adapted from a 2020 review) has an error. This schematic depicts uniglomerular projection neurons of the antennal lobe projecting directly to the lateral horn (without synapsing in the mushroom bodies) and multiglomerular projection neurons projecting to the mushroom bodies and then lateral horn. This should be reversed (uniglomerular PNs synapse in the calyx and then further project to the LH and multiglomerular PNs project along the mlACT directly to the LH) and is nicely depicted in a Strutz et al 2014 publication in eLife.

      We thank the reviewer for spotting this error. We have now modified the schematic as suggested.

      Reviewer #2 (Public review):

      Summary:

      The authors aim to provide a comprehensive description of the neurosecretory network in the adult Drosophila brain. They sought to assign and verify the types of 80 neurosecretory cells (NSCs) found in the publicly available FlyWire female brain connectome. They then describe the organization of synaptic inputs and outputs across NSC types and outline circuits by which olfaction may regulate NSCs, and by which Corazon-producing NSCs may regulate flight behavior. Leveraging existing transcriptomic data, they also describe the hormone and receptor expressions in the NSCs and suggest putative paracrine signaling between NSCs. Taken together, these analyses provide a framework for future experiments, which may demonstrate whether and how NSCs, and the circuits to which they belong, may shape physiological function or animal behavior.

      Strengths:

      This study uses the FlyWire female brain connectome (Dorkenwald et al. 2023) to assign putative cell types to the 80 neurosecretory cells (NSCs) based on clustering of synaptic connectivity and morphological features. The authors then verify type assignments for selected populations by matching cluster sizes to anatomical localization and cell counts using immunohistochemistry of neuropeptide expression and markers with known co-expression.

      The authors compare their findings to previous work describing the synaptic connectivity of the neurosecretory network in larval Drosophila (Huckesfeld et al., 2021), finding that there are some differences between these developmental stages. Direct comparisons between adults and larvae are made possible through direct comparison in Table 1, as well as the authors' choice to adopt similar (or equivalent) analyses and data visualizations in the present paper's figures.

      The authors extract core themes in NSC synaptic connectivity that speak to their function: different NSC types are downstream of shared presynaptic outputs, suggesting the possibility of joint or coordinated activation, depending on upstream activity. NSCs receive some but not all modalities of sensory input. NSCs have more synaptic inputs than outputs, suggesting they predominantly influence neuronal and whole-body physiology through paracrine and endocrine signaling.

      The authors outline synaptic pathways by which olfactory inputs may influence NSC activity and by which Corazonin-releasing NSCs may regulate flight. These analyses provide a basis for future experiments, which may demonstrate whether and how such circuits shape physiological function or animal behavior.

      The authors extract expression patterns of neuropeptides and receptors across NSC cell types from existing transcriptomic data (Davie et al., 2018) and present the hypothesis that NSCs could be interconnected via paracrine signaling. The authors also catalog hormone receptor expression across tissues, drawing from the Fly Cell Atlas (Li et al., 2022).

      We thank this reviewer for the thorough assessment and for highlighting the strengths of our manuscript. Based on comments from the other reviewer, we now include additional analyses of NSCs from two new recent datasets – the brain and nerve cord (BANC) connectome and the male central nervous system (maleCNS) connectome. Our original conclusions based on the FlyWire connectome remain unchanged, further validating our analyses.

      Weaknesses:

      The clustering of NSCs by their presynaptic inputs and morphological features, along with corroboration with their anatomical locations, distinguished some, but not all cell types. The authors attempt to distinguish cell types using additional methodologies: immunohistochemistry (Figure 2), retrograde trans-synaptic labeling, and characterization of dense core vesicle characteristics in the FlyWire dataset (Figure 1, Supplement 1). However, these corroborating experiments often lacked experimental replicates, were not rigorously quantified, and/or were presented as singular images from individual animals or even individual cells of interest. The assignments of DH44 and DMS types remain particularly unconvincing.

      We thank the reviewer for this comment. We would like to clarify that all immunohistochemical images presented in this manuscript are representative images based on at least 5 independent samples. We have now clarified this in the methods.

      Additionally, we show DH44 > retro-Tango signal across five samples (new Figure 2 Supplement 3) to highlight the consistency of retrograde trans-synaptic labeling. We also show the neurons providing inputs to putative m-NSC<sup>DH44</sup> and putative m-NSC<sup>DMS</sup> in both FAFB and maleCNS connectomes (new Figure 2 Supplement 2B-C). In both the FAFB and maleCNS datasets, we see a group of neurons (marked by black arrows) providing inputs to m-NSC<sup>DMS</sup> but not m-NSC<sup>DH44</sup>. Importantly, these input neurons are not labelled in DH44 > retro-Tango samples, lending further support to our assignment of DH44 and DMS cell types.

      The electron micrographs showing dense core vesicle (DCV) characteristics (new Figure 2 Supplement 2E-G) are also representative images based on examination of multiple neurons. However, we agree with the reviewer that a rigorous quantification would be useful to showcase the differences between DCVs from NSC subtypes. Therefore, we have now performed a quantitative analysis of the DCVs in putative m-NSC<sup>DH44</sup> (n=6), putative m-NSC<sup>DMS</sup> (n=6) and descending neurons (n=2) known to express DMS across three datasets (FlyWire, BANC and maleCNS connectomes). For consistency, we examined the cross section of each cell where the diameter of nuclei was the largest. We quantified the mean gray value of at least 50 DCVs per cell. The individual who performed these analyses was blind to the neuron identity. Our analysis (new Figure 2 Supplement 2H-J) shows that mean gray values of putative m-NSC<sup>DMS</sup> and DMS descending neurons in FAFB and maleCNS are not significantly different, whereas the mean gray values of m-NSC<sup>DH44</sup> are significantly higher. This analysis agrees with our initial DH44 and DMS NSC subtype assignments. Nonetheless, given the similarity in morphology and synaptic connectivity of DH44 and DMS neurons, we have included the limitation on cell type assignment in the absence of molecular markers in the connectome datasets.

      The authors present connectivity diagrams for visualization of putative paracrine signaling between NSCs based on their peptide and receptor expression patterns. These transcriptomic data alone are inadequate for drawing these conclusions, and these connectivity diagrams are untested hypotheses rather than results. The authors do discuss this in the Discussion section.

      We agree with the reviewer that the novel paracrine pathways presented are untested hypotheses. However, there is a very high likelihood that a given NSC subtype can signal to another NSC subtype using a neuropeptide if its receptor is expressed in the target NSC. This is due to the fact that all NSC axons are part of the same nerve bundle (nervi corpora cardiaca) which exits the brain. The axons of different NSCs form release sites that are extremely close to each other. While the release sites in NSCs cannot be visualized in adult Drosophila connectomes (since these regions were not included in the sample prep), these have been mapped in the larvae and shown to be in close proximity (Hückesfeld et al., 2021: https://doi.org/10.7554/eLife.65745). Neuropeptides from these release sites can easily diffuse via the hemolymph to peripheral tissues (e.g. fat body and ovaries) that are much further away from the release sites on neighboring NSCs. We believe that neuropeptide receptors are expressed in NSCs near these release sites where they can receive inputs, not just from the adjacent NSCs, but also from other sources such as the gut enteroendocrine cells. Hence, neuropeptide diffusion is not a limiting factor preventing paracrine signaling between NSCs, and receptor expression is a good indicator for putative paracrine signaling. Consistent with this, several pathways highlighted in the plot (CRZ to CAPA, DH44 to Hugin and Hugin to DH44) have been anatomically and/or functionally validated previously (Zandawala et al., 2021: https://doi.org/10.1371/journal.pgen.1009425; King et al., 2017: https://doi.org/10.1016/j.cub.2017.05.089; Mizuno et al., 2021: https://doi.org/10.1111/dgd.12733). Additionally, a similar analysis was also employed to depict putative interactions between NSCs in larval Drosophila (Hückesfeld et al., 2021). We have now modified the caption for this figure to explicitly state these connections are putative. We hope that the putative pathways presented here will inspire future functional studies, and have also highlighted this outstanding question in the summary Figure 10.

      Reviewer #3 (Public review):

      Summary:

      The manuscript presents an ambitious and comprehensive synaptic connectome of neurosecretory cells (NSC) in the Drosophila brain, which highlights the neural circuits underlying hormonal regulation of physiology and behaviour. The authors use EM-based connectomics, retrograde tracing, and previously characterised single-cell transcriptomic data. The goal was to map the inputs to and outputs from NSCs, revealing novel interactions between sensory, motor, and neurosecretory systems. The results are of great value for the field of neuroendocrinology, with implications for understanding how hormonal signals integrate with brain function to coordinate physiology.

      The manuscript is well-written and provides novel insights into the neurosecretory connectome in the adult Drosophila brain. Some, additional behavioural experiments will significantly strengthen the conclusions.

      Strengths:

      (1) Rigorous anatomical analysis

      (2) Novel insights on the wiring logic of the neurosecretory cells.

      We thank this reviewer for the thorough assessment and highlighting the strengths of our manuscript.

      Weaknesses:

      (1) Functional validation of findings would greatly improve the manuscript.

      We agree with this reviewer that assessing the functional output from NSCs would improve the manuscript. Given that we currently lack genetic tools to measure hormone levels and that behaviors and physiology are modulated by NSCs on slow timescales, it is difficult to assess the immediate functional impact of the sensory inputs to NSC using approaches such as optogenetics. However, since l-NSC<sup>CRZ</sup> are the only known cell type that provide output to descending neurons, we have functionally tested this output pathway using different behavioral assays (new Figure 8 and Supplements). Our analysis identifies a novel role for l-NSC<sup>CRZ</sup> and DNg27 neurons in female reproduction (based on the number of eggs laid).

      Recommendations for the authors:

      Reviewing Editor Comments:

      You will see that the reviewers found your work interesting and valuable, but had some suggestions for how revision could improve the manuscript. A common thread in the reviews is that functional speculations about the extracted circuits and paracrine signaling would benefit from revision, and would fit better in the Discussion, not Results section. Caveats could be more explicitly stated and language asserting functionality could be tempered. The reviewers were unanimous in their desire for a summary diagram or model.

      We thank the editor for these suggestions to improve the manuscript. We have now functionally validated some output pathways from l-NSC<sup>CRZ</sup>. We have also toned down the language regarding functionality where appropriate. Finally, we included a figure (new Figure 10) which summarizes the main findings from our manuscript and places them within the larger context of neuroendocrine signaling in adult Drosophila in relation to other studies.

      Reviewer #2 (Recommendations for the authors):

      Suggestions for improved or additional experiments, data, or analyses:

      The authors present connectomic analyses for NSCs identified in the FlyWire dataset. All of their connectomic findings would be strengthened by executing these same analyses in the freely available female hemibrain connectome (Scheffer et al. 2020; Plaza et al. 2022), thereby effectively increasing their sample size from one whole brain to three hemispheres. It is unclear why the authors chose only to focus on the FlyWire dataset.

      We thank the reviewer for this suggestion. We had performed a preliminary analysis using the hemibrain dataset. However, out of the 80 endocrine cells that we found in FlyWire, the hemibrain dataset lacks both the NSC subtypes in the SEZ (SEZ-NSC<sup>CAPA</sup> and SEZ-NSC<sup>Hugin</sup>) as well as l-NSC subtypes in the other hemisphere (l-NSC<sup>ITP</sup>, l-NSC<sup>DH31</sup>, l-NSC<sup>CRZ</sup>). In addition, a majority of the input synapses for all NSC are in the SEZ region which allowed us to classify the different NSC subtypes in FlyWire. Since this information is missing in the hemibrain dataset, we are unable to classify the m-NSC into the different subtypes (not shown). Therefore, we cannot perform a comprehensive analysis of input and output pathways of different NSC subtypes using the hemibrain dataset. To address this concern, we have repeated several analyses with two new recent datasets – the brain and nerve cord (BANC) connectome and the male central nervous system (maleCNS) connectome (Table 1, new Figure 1 Supplement 1, new Figure 2 Supplement 2, new Figure 3 Supplement 3, new Figure 6 Supplement 1, new Figure 7 Supplement 3). Our original conclusions based on the FlyWire connectome remain unchanged, further validating our analyses.

      The authors initially map assign NSC types based on anatomical locations and clustering of presynaptic connections and morphological features. Due to matching cell counts and similar soma locations, they find that DMS and DH44 types cannot be easily distinguished. The authors attempt to assign cell types to these two populations using two methods, neither of which are convincing as executed:

      (1) The authors attempt to distinguish the identities of the two populations by anatomically comparing presynaptic inputs in FlyWire to those observed with light microscopy using retrograde trans-synaptic labeling. Due to the lack of a genetic driver line for the DMS population, the authors could complete this only for the DH44 population. The authors present only one animal, at inadequate magnification to see the absence of distinguishing presynaptic neurons. The results would be strengthened by the presentation and quantification of multiple samples; without more than one sample, it is not possible to know how robust this finding is in this genetic driver line. The authors might also consider taking advantage of the widely-used template brain (Bogovic, 2020) to align their light micrographs of presynaptic inputs from the retrograde tracing, with the presynaptic skeletons from FlyWire and compare in a more quantitative and precise manner. The authors might also consider taking a similar approach using anterograde tracing (Talay et al. 2017) to label postsynaptic outputs. Given that postsynaptic outputs are fewer, so long as there are identifiable, distinct postsynaptic partners, it may be easier to distinguish the two populations with anterograde tracing.

      We thank the reviewer for this comment. We would like to clarify that all immunohistochemical images presented in this manuscript are representative images based on at least 5 independent samples. We have now clarified this in the methods.

      Additionally, we show DH44 > retro-Tango signal across five samples (new Figure 2 Supplement 3) to highlight the consistency of retrograde trans-synaptic labeling. We also provide a magnified image in this figure to highlight the absence of presynaptic neurons that distinguish m-NSC<sup>DH44</sup> and m-NSC<sup>DMS</sup>.

      We also show the neurons providing inputs to putative m-NSC<sup>DH44</sup> and putative m-NSC<sup>DMS</sup> in both FAFB and maleCNS connectomes (new Figure 2 Supplement 2B-C). In both the FAFB and maleCNS datasets, we see a group of neurons (marked by black arrows) providing inputs to mNSC<sup>DMS</sup> but not m-NSC<sup>DH44</sup>. Importantly, these input neurons are not labelled in DH44 > retroTango samples, lending further support to our assignment of DH44 and DMS cell types.

      As per this reviewer’s suggestion, we also aligned our retrograde tracing light micrographs to a template brain (Author response image 1). However, we were unable to quantitatively compare neurons in our light micrographs with neuronal skeletons from the connectome. This is because retroTango labels several neurons in the SEZ which obscures morphology of individual neurons needed for such comparisons. Additional experiments, where retro-Tango output is restricted to sparse populations of neurons using a Flp-out strategy, are needed to perform such quantitative analyses. These experiments are beyond the scope of this study since we now provide additional lines of evidence for cell assignments.

      Author response image 1.

      DH44 > retro-Tango presynaptic signal aligned to JRC2018 unisex template brain

      We appreciate the suggestion to use the anterograde tracing tool trans-Tango to distinguish mNSC<sup>DH44</sup> and m-NSC<sup>DMS</sup>. There is very little synaptic output from m-NSC<sup>DMS</sup> and m-NSC<sup>DH44</sup> based on the FlyWire connectome. There is no synaptic output from both of these cell types if we use a threshold of 5 synapses for significant connections (new Figure 7). Using a threshold of 2 synapses for significant synaptic connections, 3 neurons are downstream of m-NSC<sup>DH441</sup> and 5 neurons are downstream of m-NSC<sup>DMS</sup> (not shown). Since these postsynaptic neurons are not bilaterally paired (we do not anticipate unilateral pathways), we don’t think that these connections are significant. Consistent with our analysis with the FlyWire connectome, we did not observe any significant post-synaptic signal with DH44 > trans-Tango (Author response image 2) even using flies raised at 21ºC which increases the synaptic strength during development. Since we do not have a GAL4 driver to specifically target m-NSC<sup>DMS</sup> , we could not perform similar trans-Tango analysis of m-NSC<sup>DMS</sup>.

      Author response image 2.

      DH44 > trans-Tango (left) and w<sup>1118</sup> > trans-Tango (right; control). Presynaptic neurons are labelled in green and post-synaptic neurons are in red. Representative images based on 5 samples.

      Our connectome analyses revealed that putative m-NSC<sup>DMS</sup> receive direct synaptic inputs from enteric neurons but m-NSC<sup>DH44</sup> do not. We used this information to perform another trans-Tango analysis using Gr43a-Gal4 which labels a subpopulation of enteric neurons (Miyamoto and Amrein, 2013: https://doi.org/10.4161/fly.27241) (Author response image 3).

      Author response image 3.

      Initiating trans-Tango from Gr43a neurons (green) does not label any postsynaptic neurons (magenta) in the pars intercerebralis (white arrow head), including those labelled by the DMS antibody (cyan).

      Unfortunately, initiating trans-Tango from Gr43a neurons did not label any post-synaptic neurons in the pars intercerebralis where m-NSC<sup>DH44</sup> and m-NSC<sup>DMS</sup> soma are located. This could be due to a) low trans-Tango sensitivity or b) m-NSC<sup>DMS</sup> are downstream from other enteric neurons not captured by Gr43a-GAL4. In the absence of other broad enteric neuron drivers, we are unable to perform additional analyses.

      (2) The authors attempt to assign cell types by qualitatively assessing the darkness of dense core vesicles in these two populations. However, there is a presentation of only single planar images through three selected cells (a DMS-expressing descending neuron, DMS-expressing NSC, and DH44-expressing NSC) without any quantitative analyses of vesicle characteristics within or across NSC cell types. It is not possible for the reader to assess whether the darker vesicles constitute a real trend, or if these images are hand-selected to support their point. This piece of evidence would be more convincing if the authors demonstrate consistent vesicle characteristics within NSC type and differences across type. Moreover, such analysis of dense core vesicle features in cell types with distinct and known peptide expression would be broadly interesting.

      Given that NSC type assignment is a major contribution of the present paper, it is critical that the authors are clear about the remaining uncertainty in assigning cell types, so as not to propagate false certainty into future work.

      This comment has been addressed above, and we refer the reviewer to the new Figure 2 Supplement 2E-J.

      The authors suggest larger peptide release capacity from CAPA-producing NSCs based on their larger morphological features (Figure 1, Supplement 2), which is more speculative than certain. In Figure 1, Supplement 1 the authors demonstrate the capacity to visualize vesicles number and size in individual NSCs. Rather than speculate over larger peptide release capacity based on cell size, the authors could quantify these vesicle features, which are surely a better indication of peptide release capacities.

      We thank the reviewer for this comment. We agree that number of dense core vesicles within these and other neurons would be a better indicator of their peptide release capacity. We are performing these analyses on a brain-wide scale as part of another project. Therefore, we have removed the following speculative statement from the present manuscript:

      “But given their location, large size, and presumed large release capacity, we speculate that SEZ-NSC<sup>CAPA</sup> participate in global modulation of post-feeding physiology.”

      The authors provide an analysis of NSCs' synaptic inputs and outputs, but never mention whether NSCs are synaptically connected to each other. If connected, it would be very sensible to provide some analysis of synaptic connectivity between NSCs. If they are not connected, the authors should explicitly mention this in the main text, as it is relevant to the overall aim of this study.

      All NSCs are classified as endocrine cells in the FlyWire connectome. Hence, as shown in new Figures 3B and 7B, NSCs do not provide output to any endocrine cells (NSCs) using a threshold of 5 synapses for significant connections. Similarly, we do not see any synaptic connectivity between NSCs in the BANC dataset (new Figure 7 Supplement 3A-B). We do observe sparse connectivity between NSCs in the maleCNS dataset with a threshold of 5 synapses (new Figure 7 Supplement 3C-D), as well as in the Flywire connectome when the threshold is reduced to 2 synapses (new Figure 7 Supplement 2). However, we refrain from emphasizing on these connections because additional validation is required to rule out false positives in synapse predictions. Dense-core vesicles in NSCs can frequently be mistaken for synaptic T-bars during the prediction (unpublished observation).

      Although unlikely, NSCs could also form synapses with each other near their release sites and outside the brain volumes captured in all three datasets examined in this study. This limitation has been included in the discussion.

      There is no substitute for a good circuit wiring diagram; the motifs that are extracted in Figure 3H might be better appreciated if the reader was first presented with a well-formatted complete circuit diagram, which may then foreshadow the points made in the main text and in Figure 3H.

      We appreciate this suggestion. We now include a circuit diagram (new Figure 3G) to highlight the connectivity between NSCs and their presynaptic partners. The proportion plot (old Figure 3G) has now been moved to new Figure 3 Supplement 5.

      The authors provide extensive bar graphs showing synaptic input body IDs in Figure 3 Supplement 2, however they don't complete the same analysis for synaptic outputs (likely due to low numbers). Even so, it would be useful to the reader if the body IDs and cell types for both synaptic inputs and outputs were documented in a supplemental table. Providing such an inventory is aligned with the goals of this study.

      Only l-NSC<sup>unknown</sup> and l-NSC<sup>CRZ</sup> provide synaptic outputs in the FlyWire connectome (new Figure 7B-F). We have included bar graphs showing output from both these cell types at a single-cell level (new Figure 7G). Additionally, we have annotated all the NSC subtypes in FlyWire and BANC on Codex. Further exploring the inputs and outputs of NSC subtypes can be done interactively on Codex. For example, the search command “{upstream_union} cell_type == SEZ_NSC_CAPA” will retrieve all the neurons providing inputs to SEZ-NSC<sup>CAPA</sup>. As a quick search, this is more convenient than pasting individual body IDs from a supplementary table into Codex. All code outputs (csv files) containing this information are also available on Zenodo.

      In describing possible paracrine signaling, the authors write "Given the proximity of NSC axon terminations, it is extremely likely that a hormone released from a given NSC will influence the activity of other NSC types if its receptor is expressed in those cells." In the absence of functional experiments and/or information about spatial localization and/or peptide diffusion and the proximity of receptors to release sites, the expression patterns alone are insufficient to support this conclusion. Thus, the authors might consider removing the circular connectivity plots in Figure 7C, and Figure 7, Supplement 1A-H, and instead emphasize what can be concluded with certainty from the transcriptomic data (which are expression patterns of the hormones and receptors across NSCs and other tissue types). The authors might instead speculate over potential paracrine signaling between NSCs in the Discussion. Given that the authors describe paracrine signaling between NSCs as 'putative' in the abstract and main text, the authors will likely agree the legend for Figure 7 is misleading.

      This comment has been addressed above. We agree with the reviewer and have modified the figure legends (new Figure 9 and Figure 9 Supplement 1) to emphasize that the connections are putative.

      Should the authors keep these connectivity diagrams, it is important to reconsider their threshold wherein 50% of cells in a cluster must express a given hormone for it to be considered present in their analysis. It is entirely conceivable that there is real heterogeneity in hormone expression within the cluster, so it is surprising the authors have applied this artificial criterion.

      We thank the reviewer for presenting us with this option.

      We also apologize for the oversight in explaining our thresholding carefully. To minimize false positives, neuropeptides were subjected to a two-step filtering process. First, only those expressed in at least 50% of the cells within a given cluster were retained. Second, a composite expression score was calculated for each neuropeptide by multiplying its average expression by its percent detection. These values were normalized to the maximum observed signal across the dataset, and only hormone-cluster pairs maintaining a relative score of 0.25 or higher were included in the final analysis. This stringent filtering approach was implemented to focus the analysis on dominant neuropeptides and to exclude contamination from ambient RNA, which is common for neuropeptides (Allen et al., 2020: https://doi.org/10.7554/eLife.54074). To account for lower abundance of receptor transcripts, we used a more permissive threshold for receptors by retaining those expressed in at least 5% of the cells within a cluster. Unlike the neuropeptides, no secondary relative-score filtering was applied to the receptors to ensure that biologically relevant signaling targets were not prematurely excluded due to low transcript density. We have now revised our methods to explain these details.

      Importantly, we used this thresholding criteria to align previous anatomical studies with our single-cell expression analysis and filter out neuropeptides that likely represent contamination: 1) Transcript for leucokinin (Lk) neuropeptide is expressed in l-NSC<sup>ITP</sup> (but previous studies have not been able to detect this peptide in l-NSC<sup>ITP</sup> (Zandawala et al., 2018: https://doi.org/10.1371/journal.pgen.1007767). 2) Hugin is not expressed in m-NSC<sup>DMS</sup> (Oh et al., 2021: https://doi.org/10.1016/j.neuron.2021.04.028). 3) Ilp2 is not expressed in SEZNSC<sup>CAPA</sup> and m-NSC<sup>DMS</sup> (this study and various others). 4) ITP is not expressed in any m-NSC (Gera et al., 2025: https://doi.org/10.7554/eLife.97043.3). Based on these and other examples, we feel that our stringent criteria recover putative pathways that are likely functional while filtering obvious false positive. Nonetheless, we agree with the reviewer that some of these NSCs could represent heterogeneous clusters as has been shown recently for m-NSC<sup>DILP</sup> (Held, Bisen, Zandawala et al., 2025: https://doi.org/10.7554/eLife.99548.3). This heterogeneity could result in some authentic connections to drop out. However, our goal for this analysis was to not identify all the putative paracrine connections, but rather the strongest ones with the hope that it can inspire future functional studies.

      The data shown in Figure 2 would be easier to interpret and therefore more convincing with better use of insets, appropriate overlays of multiple markers, higher image magnifications, and quantification across samples. Specifically: In Figure 2C, authors show mCherry expression but it isn't clear where these cell bodies are located with respect to the image in Figure 2B. This is also true for Figure 2D. Insets in Figure 2B that correspond with regions shown in 2C and 2D would be helpful. In Figure 2, the authors do not provide cell counts across samples for all markers. Thus, it isn't clear how consistent these cell counts are across samples. In Figure 2E, authors claim that no Gr64f-positive cells innervate the NCC, yet there is clearly a GFP signal in the NCC region in the merged image. The authors should provide an additional marker or a higher magnification image to convince the reader that these projections are not in the NCC region.

      We thank the reviewer for these suggestions. To improve clarity, we have made the following changes:

      Figures 2C and 2D are based on different samples than the one shown in Figure 2B. But we have added dashed boxes in Figure 2B to indicate the regions shown in Figure 2C and 2D.

      Included sample sizes in Figure 2A and 2C. The rest of the panels are representative images based on at least 5 samples. This has been included in the methods.

      The cell count has been provided for m-NSC<sup>DILP</sup> for both markers in Figure 2A. The cell counts for m-NSC<sup>DMS</sup>, labelled using mCherry alone, has been provided in Figure 2C. We did not perform cell counting when using the membrane GFP reporter as it is difficult to accurately count overlapping cells (see dashed box labelled C in Figure 2B).

      We also provide a new supplementary file (new Figure 2 Supplement 1) showing cell counts for m-NSC<sup>DILP</sup> using different markers. Based on this, we can confidently conclude that adult Drosophila typically have more than 14 m-NSC<sup>DILP</sup>.

      We have corrected a typo in our label for Figure 2E: it should be Gr64a instead of Gr64f.

      We have modified the Figure 2E inset to show that the four pairs of Gr64a > myrGFP expressing corazonin cells do not project via the NCC (labelled with an arrow). We have also identified the four pairs of Gr64a neurons (Author response image 4 left panel) in the FlyWire connectome, which shows that these neurons do not exit the brain via the NCC.

      Author response image 4.

      Corazonin-expressing Gr64a neurons in the FlyWire connectome (left) and a light micrograph (right, same as in Figure 2E) showing Gr64a neurons (green) and corazonin neurons (magenta).

      Recommendations for improving the writing and presentation:

      Throughout the paper, the authors provide scant or, at times, no citations. Inadequate citation is as much an issue in the introduction as it is in the results and discussion sections. As such, the authors often do not provide a well-supported premise for the present work and/or do not place their findings and interpretations into the context of existing literature. Related, there is a predominance of references to the work of the authors themselves, often in place of citing earlier foundational work. Citations are nearly exclusive to the Drosophila literature, with the exception of the second paragraph of the introduction. This paper would be greatly improved with references to a broader literature.

      We have now added additional references to give credit to foundational work where appropriate. We have also included citations to non-Drosophila literature for more general statements in the introduction and discussion; however, we refrain from citing such studies in the results section to keep it focused.

      Figure 1 Supplement 1 is referenced after Figure 2 in the text. The authors might consider reassigning it as a supplement to Figure 2, which also uses imaging methodologies to distinguish NSC cell types.

      We agree with the reviewer and have reassigned the figures accordingly.

      Figure 3G is difficult to interpret, and its figure legend is brief and inadequate.

      As suggested by this reviewer, we have replaced this panel with a circuit diagram. The proportion plot (old Figure 3G) has now been moved to Figure 3 Supplement 5, and we have expanded the figure legend.

      The bar graphs in Figure 3 Supplements 2 and 3 would best benefit the reader if the x-axis labels are not simply body IDs, but also cell types or instances (if assigned in FlyWire).

      We appreciate this suggestion. Cell types are routinely updated on Codex while the root IDs remain static for v783. Therefore, we chose root IDs for these plots as they can be used to query Codex easily and reliably. We now provide all raw data as csv files on Zenodo used to make these plots. This includes cell types and other classifications.

      Minor corrections to the text and figures:

      Table 1 compares the observed numbers of NSC types in adult flies to those in larvae and those expected based on previous literature. The authors should cite the previous studies that support each of the expected or larval numbers, either within the table or in the table legend. It would also be appreciated if the expected numbers were cited in the main text.

      References for NSC numbers in larvae and expected numbers in adults are now included in Table 1.

      In describing the author's approach to analyzing synaptic connectivity by cosine similarity, authors cite their own previous work rather than the foundational study describing this approach or earlier studies that use it.

      We have now also cited Schlegel et al., 2021 (https://doi.org/10.7554/eLife.66018) who used a similar approach in the olfactory system.

      Reviewer #3 (Recommendations for the authors):

      (1) The observation that most gustatory inputs to NSCs are indirect (particularly for feedingrelated NSCs) is very interesting but lacks functional validation. I suggest that the authors conduct behavioural assays where specific sensory inputs are activated or silenced while monitoring outputs from NSCs. This could include optogenetics to stimulate or inhibit sensory neurons, or alternative feeding assays.

      We thank the reviewer for this insightful suggestion. We agree that the functional validation of gustatory-to-NSC pathways is a highly compelling direction for future research. However, we believe that behavioral assays, as suggested, pose significant interpretive challenges for the following reasons:

      NSCs primarily function by releasing hormones into the systemic circulation. Unlike classical neurotransmission, hormonal modulation typically operates on much slower timescales (minutes to hours). Consequently, acute activation of sensory inputs is unlikely to elicit immediate, quantifiable behavioral changes that can be specifically attributed to NSC activity.

      Most NSC classes are known to influence multiple physiological and behavioral processes simultaneously. Attributing a specific behavioral phenotype to a single NSC class following sensory stimulation would be confounded by these overlapping roles.

      Activating or silencing taste neurons will directly impact feeding behavior through canonical motor circuits, independent of the neuroendocrine system. In such a paradigm, it would be nearly impossible to isolate the specific "indirect" contribution of the NSCs to the observed behavior.

      While we agree that functional connectivity, such as optogenetic activation of taste neurons paired with calcium imaging (e.g., GCaMP) in NSCs, would be the ideal way to validate these inputs, we consider these extensive physiological experiments to be beyond the scope of this anatomical and connectomic study.

      Nonetheless, to address this important question, we now use a recently developed approach (Bates et al., 2026: https://doi.org/10.1101/2025.07.31.667571) based on linear dynamical modeling to estimate the influence of various sensory neurons (gustatory, olfactory, enteric, hygrosensory, etc.) on different NSC classes. Our analysis (new Figure 6 and Figure 6 Supplement 1) reveals that contents of consumed food (detected by enteric neurons) have a stronger influence on NSCs compared to inputs from external taste receptors.

      (2) Descending neurons appear to play a crucial role in regulating both motor and endocrine output. However, their functional contribution is only inferred from the connectomic data. The authors could perform functional activity manipulations (silencing or activating) of these descending neurons (for instance dMS descending neurons) to explore their role in behaviour. This could be tested with simple behavioural assays such as feeding or reproduction (i.e egg laying).

      We believe that there might be some confusion. DMS descending neurons (DNp32 cell type) used for dense-core vesicle comparisons with m-NSC<sup>DMS</sup> and m-NSC<sup>DH44</sup> (new Figure 2 Supplement 2) are different from the descending neurons (DNg27 cell type) that receive inputs from l-NSC<sup>CRZ</sup> (new Figure 7). We have indicated the cell type of DMS descending neurons in the text to clarify this. We have also functionally tested DNg27 (instead of DMS descending neurons suggested by the reviewer) using optogenetic and chemogenetic approaches for effects in feeding, food preference, starvation survival, egg laying and flight (new Figure 8 and Figure 8 Supplement 1). While we expected DNg27 to influence flight based on our connectome analyses, we do not see any phenotype in our free flight setup following DNg27 activation (new Figure 8 Supplement 1). However, this could be due to the split GAL4 driver used being very weak (new Figure 8 Supplement 2). This is also supported by the egg-laying assay where DNg27 inactivation only produces a phenotype after day 8 (Figure 8). Since we currently don’t have access to another driver to specifically target DNg27, we are unable to validate our results in the free flight setup using an independent driver.

      (3) The authors describe a sparse olfactory input pathway to NSCs, with emphasis on odours playing major roles. However, the physiological consequences of these connections are not explored in detail. Authors should use ORN/AL stimulation (e.g., using optogenetics) to explore how odour sensory pathways affect hormonal secretion in NSCs.

      We acknowledge the reviewer’s interest in the physiological consequences of the olfactory to NSC pathways identified in our study. While we agree that exploring how specific odors modulate neuroendocrine output is a logical next step, we believe that such experiments are currently unfeasible due to significant technical and biological constraints as highlighted above for taste neurons. Hence, we calculated the influence of olfactory receptor neuron activation on different NSC classes using an approach based on linear dynamical modeling (new Figure 6 and Figure 6 Supplement 1). Our analysis reveals that smell has a weaker influence on NSC compared to taste.

      (4) The authors present a large amount of nice yet complex data, which can be difficult to navigate through and is sometimes hard to follow. Consider adding more schematic diagrams to summarize the key pathways and interactions between NSC types and their inputs/outputs.

      We thank the reviewer for this suggestion. We have now included a figure (new Figure 10) which summarizes the main findings from our manuscript and places them within the larger context of neuroendocrine signaling in adult Drosophila in relation to other studies.

    1. eLife Assessment

      This study reports results characterizing two boundaries (homie and nhomie) flanking the TAD of a developmental gene. The study's thorough and mechanistic characterization of these boundaries provides useful information to the field, and the conclusions are well supported by convincing evidence. The novelty of the work is somewhat limited because the major conclusion is confirmatory in nature.

    2. Reviewer #1 (Public review):

      Summary:

      This study extends the authors' prior work on transgenic nhomie/homie boundary pairing (Fujioka et al. 2016 PLoS Genetics), which showed that these elements - corresponding to the eve TAD's left and right boundaries - can pair with endogenous copies over large genomic distances (142 kb here), bridging a linked reporter gene to endogenous eve enhancers for long-range gene activation (shown again here in Figure 1). Physical pairing was previously confirmed (Chen et al. 2018 Nat Genet) and further resolved by Micro-C (Bing et al. 2024 eLife), supporting the hypothesized "stem-loop" or "circle loop" topologies used to explain homie/nhomie directional pairing (shown here for nhomie in Figures 2-3). The authors recently showed that a Su(Hw) binding site is required for homie-mediated reporter gene activation by eve enhancers (Fujioka et al. 2025 Genetics); here, they extend this finding to nhomie (Figures 4-6), further showing that Su(Hw) motifs are required for Micro-C-detectable looping between transgenic and endogenous eve boundaries (Figures 7-9). Finally, they show that while cis-pairing over 142 kb is highly specific to homie/nhomie elements, transvection between homologous transgene insertions is more permissive (functioning with the Su(Hw)-bound gypsy insulator) but still shows some specificity (failing with the CTCF-bound Fab8 boundary, Figures 10-11).

      Strengths:

      The question of how pairs of loci can specifically physically pair over relatively long genomic distances is an interesting fundamental question. The study's strengths are the clarity and meticulous interpretation of the results, and the authors' conclusions are compelling.

      Weaknesses:

      A major weakness is that some figures reproduce previously published findings; in some cases it is unclear whether the same fly lines were used, and in others, the lines differ only slightly from those used previously (e.g., a shorter version of the Homie transgene than the one used previously). Most conclusions in this manuscript have already been published elsewhere. As a result, the paper does not report a genuine new discovery, and only incrementally advances our understanding of boundary pairing.

    3. Reviewer #2 (Public review):

      The results in Ke et al., build on 15 years of work focused on dissecting the pairing properties of the Drosophila Homie insulator. Here, the authors use similar methods to those shown in Fujioka et al., 2016, Ke et al., 2024, and Fujioka et al., 2025, but with a focus on nHomie pairing and the role of Su(Hw) in both Homie and nHomie long-range interactions. The main question the authors hope to address is what the mechanisms are behind the physical interactions involved in boundary:boundary pairing. They attempt to answer this question through mutating the Su(Hw) binding sites located within the nHomie and Homie transgenic sequences and observing how pairing is altered.

      The work presented is thorough and thought out; however, some of the conclusions that the authors focus on are not what makes the work interesting and could be reprioritized. For example, the authors spend several paragraphs in the discussion (lines 531-595) addressing how the data presented does not support an argument for cohesion-mediated loop extrusion. While the interactions shown throughout the manuscript do not support cohesion-mediated loop extrusion occurring at the Homie locus, the authors have already made this point in both Bing et al., 2024 and Ke et al., 2024 and thus do not need to expound on this point.

      Instead, the authors have a more compelling story in their specificity vs promiscuity arguments. Homie is a unique insulator in Drosophila and even when located 142kb away will still find its unique pairing partners (itself and nHomie). The authors have shown this several times prior, yet here they show that some level of this long-distance homing interaction is dependent upon the Su(Hw) binding site. Additionally, the authors show in this study that addition of gypsy sequence, in a less demanding assay, is sufficient for transvection pairing with Homie. This transvection result is a novel finding, as gypsy was previously shown to be insufficient for long-distance pairing with Homie based on the authors' prior studies. It is likely different architectural proteins that bind within the Homie sequence and allow it to pair specifically with itself, regardless of assay type, and these elements are likely absent from the gypsy sequence, leading to pairing that is more situational (see point 8 in recommendations).

      Finally, to no fault of the authors, the art of visualizing complex 3D pairing configurations is difficult. Unfortunately, that can at times mask the ultimate points that the authors are trying to make about pairing early in the manuscript.

      Overall, the work mainly supports the authors' claims, and the findings are a useful addition to the insulator and Drosophila 3D genome organization field.

    4. Reviewer #3 (Public review):

      Summary:

      This manuscript investigates the function of Su(Hw) binding sites found in two boundaries/insulators, homie and nhomie, in TAD formation that encompasses the eve gene. They tested the hypothesis that Su(Hw) binds homie and nhomie, thereby forming a stem-loop TAD. The authors used transgene reporters with various mutations, and the results support the hypothesis strongly.

      Strengths:

      They combine reporter assays (GFP and LacZ expression) with MicroC contact profiling to robustly support their conclusion. Overall, they propose how Su(Hw) mediates physical interaction between boundary elements (homie and nhomie).

      Weaknesses:

      The writing is quite dense and not easily accessible to outside readers.

    1. eLife Assessment

      This study presents important evidence showing that the high susceptibility to sepsis of Kit-mutant mice is not due to mast cell deficiency. The revised manuscript, with solid data, focuses on refuting the notion that mast cells play a major role in sepsis. This paper would be of interest to researchers in mast cell biology and mucosal immunology.

    2. Reviewer #1 (Public review):

      Summary:

      Mast cells have previously been reported to play an important role in bacterial immune defence and act protectively in sepsis. However, many of these findings were based on studies using Kit mutant mice. In this study, the authors conducted a detailed investigation using mast cell-deficient Cpa3 Cre-Master mice. As a result, the authors found that the Cpa3 Cre-Master mice exhibited responses similar to wild-type mice in terms of bacterial immune defense. This suggests that the observed phenotype is not due to mast cell-dependent bacterial immune defense, but rather is associated with dysbiosis of the gut microbiota.

      Strengths:

      Mast cells have long been reported to play an important role in the protective response against sepsis, and their function in infection defense has been demonstrated. However, Kit mutant mice have been reported to exhibit impaired peristalsis, and several mast cell-specific genetically modified mouse lines have since been developed and examined in detail. This study presents an important finding by logically demonstrating that the exacerbation of sepsis in Kit mice is due to alterations in the gut microbiota, and that the phenotype previously thought to be mast cell-dependent was, in fact, not.

      In addition, the experiments were carefully designed using mice with matched genetic backgrounds. These findings underscore the importance of microbiota composition in interpreting immune phenotypes and highlight the need for co-housing controls in mutant mouse studies.

      A major strength of this work is the robustness of the CLP data, generated over eight years by three independent researchers across two institutions with large sample sizes, lending strong support to the conclusions.

      Weaknesses:

      The study assesses only a limited subset of gut bacterial species, leaving the extent to which E. coli expansion contributes to the observed phenotype unclear. Moreover, in the cohousing experiments, there is no evidence provided to confirm successful microbiota normalization between groups. A more detailed analysis of the microbial composition would be necessary to strengthen the reliability of the findings.

      It is also important to note that Cpa3-deficient mice exhibit not only mast cell depletion but also defects in basophils and T cells. These additional immunological alterations may counterbalance one another, potentially masking phenotypic changes and complicating interpretation.

      Furthermore, it remains to be determined whether the altered gut microbiota observed in KitW/Wv mice is a consequence of impaired intestinal motility, whether a similar phenotype is observed in KitW-sh/W-sh mice, and whether comparable results occur in SCF-deficient models. Addressing these questions would provide greater clarity on the contribution of mast cells versus secondary factors in the observed phenotypes.

      Given that KitW/Wv mice exhibit impaired peristalsis, is the observed increase in E. coli a consequence of this dysfunction?

      Previous studies with BMMC reconstitution experiments have indicated that mast cells are a source of TNF-how does this align with the current findings?

      Comments on revised version.

      The authors have made substantial efforts to address the concerns raised in the previous review, and the revised manuscript has been substantially strengthened by the additional microbiome analyses. In particular, the new data provide a more comprehensive characterization of the intestinal microbiota in both Kit mutant and mast cell-deficient mice and strengthen the interpretation that the microbiota alterations observed in Kit mutant mice are associated with Kit deficiency rather than mast cell deficiency per se. Although Enterobacteriaceae were significantly increased based on the unadjusted Welch's t-test, this difference did not remain statistically significant after correction for multiple comparisons. The authors appropriately acknowledge this limitation in the revised manuscript. Overall, I consider the major concerns regarding the microbiome analysis to have been adequately addressed.

    3. Reviewer #2 (Public review):

      Summary:

      The authors showed that the high susceptibility to CLP sepsis of Kit-mutant mice is not due to mast cell deficiency, but to dysbiosis.

      Recommendations:

      (1) The authors showed that E. coli increases in the cecum of Kit-mutant mice, which causes high CLP susceptibility. However, they did not provide any evidence E. coli is responsible for the high susceptibility. In the Figure 3 experiments, the authors administered the same number of cecal bacteria and did not show the number of E. coli after the administration. The authors should provide evidence showing that depletion of E. coli decreases susceptibility.

      (2) The author should provide direct evidence of dysbiosis by, for example, shotgun sequencing of cecal and fecal contents.

      (3) In case the authors find dysbiosis, they should analyze the mechanisms by which Kit mutation causes dysbiosis.

      Comments on revised version.

      The revised manuscript focuses on refuting the notion that mast cells play important roles in sepsis. The reviewer agrees with this claim.

    4. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      Mast cells have previously been reported to play an important role in bacterial immune defense and act protectively in sepsis. However, many of these findings were based on studies using Kit mutant mice. In this study, the authors conducted a detailed investigation using mast cell-deficient Cpa3 Cre-Master mice. As a result, the authors found that the Cpa3 Cre-Master mice exhibited responses similar to wildtype mice in terms of bacterial immune defense. This suggests that the observed phenotype is not due to mast cell-dependent bacterial immune defense, but rather is associated with dysbiosis of the gut microbiota.

      Strengths:

      Mast cells have long been reported to play an important role in the protective response against sepsis, and their function in infection defense has been demonstrated. However, Kit mutant mice have been reported to exhibit impaired peristalsis, and several mast cell-specific genetically modified mouse lines have since been developed and examined in detail. This study presents an important finding by logically demonstrating that the exacerbation of sepsis in Kit mice is due to alterations in the gut microbiota, and that the phenotype previously thought to be mast cell-dependent was, in fact, not.

      In addition, the experiments were carefully designed using mice with matched genetic backgrounds. These findings underscore the importance of microbiota composition in interpreting immune phenotypes and highlight the need for cohousing controls in mutant mouse studies.

      A major strength of this work is the robustness of the CLP data, generated over eight years by three independent researchers across two institutions with large sample sizes, lending strong support to the conclusions.

      Weaknesses:

      The study assesses only a limited subset of gut bacterial species, leaving the extent to which E. coli expansion contributes to the observed phenotype unclear.

      We now performed 16S rRNA sequencing of cecal samples isolated from Kit<sup>W/Wv</sup> and Cpa3<sup>Cre/+</sup> mice and their respective littermates. Results are display in a new Figure 4. Our comparative analysis of the cecal microbial communities in Kit<sup>W/Wv</sup> and Kit<sup>+/+</sup> mice (Figure 4A+B) confirmed the expansion of E. coli (Enterobacteriaceae) that we had observed by CFU counts (Figure 3D). Furthermore, it revealed a dysbiotic shift marked by increased abundance of Peptostreptococcaceae, Verrucomicrobiaceae, Coriobacteriaceae, and Erysipelotrichaceae in KitW/Wv mice.

      None of these changes was observed when comparing the cecal microbiomes of Cpa3Cre/+ and Cpa3+/+ mice (Figure 4C+D), indicating that the compositional shift in Kit<sup>W/Wv</sup> mice is due to the deficiency in Kit but not mast cells. Of note, as stated on page 14, the microbial changes that we observed in Kit<sup>W/Wv</sup> mice resemble dysbiotic patterns reported in chronic intestinal inflammation, experimental colitis, and impaired barrier function. These new findings fully align with and further support our earlier conclusion that Kit<sup>W/Wv</sup> mice harbour pro-pathogenic microbiota.

      The new results are display in a new Figure 4, and described on pages 9-10 and discussed on pages 13-14.

      Moreover, in the cohousing experiments, there is no evidence provided to confirm successful microbiota normalization between groups.

      It is correct that we have no direct data to confirm microbiota normalization between groups after co-housing. We note, however, that co-housing is a generally accepted method for microbiota equalization or conversion (Caruso et al., Cell Rep. 2019, Ridaura et al., Science 2013, and reviewed in Moore et al., Clin. Transl. Immunol. 2016). In any case, Kit<sup>W/Wv</sup> mutants were made resistant to CLP by co-housing. Similar microbiota sequencing results between groups, while useful, would again only be correlative.

      A more detailed analysis of the microbial composition would be necessary to strengthen the reliability of the findings.

      See above the new data from 16S rRNA sequencing.

      It is also important to note that Cpa3-deficient mice exhibit not only mast cell depletion but also defects in basophils and T cells. These additional immunological alterations may counterbalance one another, potentially masking phenotypic changes and complicating interpretation.

      Regarding basophils in Cpa3<sup>Cre/+</sup> mice, compared to wild-type mice, basophils are reduced to about 40% of normal (Feyerabend et al., Immunity 2011). In Kit<sup>W/Wv</sup> mice, compared to wild-type mice, basophils are reduced to about 10% of normal. To our knowledge, there has been no phenotype reported in which a reduction in basophils compensates for the loss for mast cells. Given that Kit<sup>W/Wv</sup> mice have about threefold lower numbers of basophils and are highly susceptible to sepsis, there is no evidence that a reduction in basophils is protective in mast cell-deficient mice. On the contrary, mice that were normal for mast cells but had their basophils depleted were more susceptible to sepsis (Piliponsky et al., Nat. Immunol. 2019). Hence, basophils appear to be protective, and their reduction increases susceptibility. In light of these data and considerations, there is no evidence for a reduction in basophils to counterbalance the loss of mast cells in Cpa3<sup>Cre/+</sup> mice.

      Regarding T cells, there is no evidence, and there are no reports, that Cpa3<sup>Cre/+</sup> mice have defects in T cells (Feyerabend et al., Immunity 2011, Feyerabend et al., Cell Metabolism 2016). Cpa3 is weakly and transiently expressed early in the T cell lineage (Feyerabend et al., Immunity 2009; for expression levels in T cells versus mast cells, see Author response image 1). In summary, in contrast to the reviewer's claim, there are no known defects in T cell development or T cell functions in Cpa3<sup>Cre/+</sup> mice. We think the reviewer needs to provide published evidence for his/her claim that Cpa3-deficient mice exhibit defects in T cells. We as authors are also obliged to support our claims scientifically, and rightfully so.

      Author response image 1.

      Generated from the Immgen database. Shown are RNAseq gene expression levels of diverse T-cell and mast cell populations.

      Furthermore, it remains to be determined whether the altered gut microbiota observed in KitW/Wv mice is a consequence of impaired intestinal motility, whether a similar phenotype is observed in KitW-sh/W-sh mice, and whether comparable results occur in SCF-deficient models. Addressing these questions would provide greater clarity on the contribution of mast cells versus secondary factors in the observed phenotypes.

      The purpose of our study was to verify or refute the key claim dating back to two 1996 Nature papers that mast cells play important roles against sepsis. We demonstrate here that this is not the case because mice without mast cells (Cpa3<sup>Cre/+</sup> mice) were as resistant to sepsis as wild-type mice. Hence, mast cells are not involved in the immunity against sepsis, and 'secondary factors' are not involved in this simple experiment (both groups of mice, wild-type and Cpa3<sup>Cre/+</sup> mice, were on the identical genetic background). Second, Kit<sup>W/Wv</sup> mice are also as resistant to sepsis as wild-type mice when confronted with the identical intestinal slurry. Therefore, Kit<sup>W/Wv</sup> mice have no immune deficit in response to sepsis. Hence, in our view, the underlying immunological question regarding the role of mast cells in sepsis has been conclusively addressed and answered by our data. We have changed the title to emphasize this central question.

      The reviewer now asks us to delve even deeper into Kit biology and in particular intestinal pathophysiology in this and other Kit or steel mutants. While we share his/her interest in such questions, we fully disagree with the statement that 'addressing these questions would provide greater clarity on the contribution of mast cells versus secondary factors in the observed phenotypes.' We do not intend to enter the field of gut physiology or its link to microbiota, all the more because any results would not affect the central conclusion of our manuscript.

      Given that KitW/Wv mice exhibit impaired peristalsis, is the observed increase in E. coli a consequence of this dysfunction?

      See above

      Previous studies with BMMC reconstitution experiments have indicated that mast cells are a source of TNF - how does this align with the current findings?

      It does not align well. It is possible that cultured and transplanted mast cells (BMMC) produce TNF. Given that we did not find a reduction in TNF levels in the peritoneal lavage or serum in mice without mast cells undergoing sepsis, under physiological conditions mast cell-derived TNF does not seem to have a measurable impact on total TNF levels.

      Reviewer #2 (Public review):

      Summary:

      This study presents a useful finding that the high susceptibility to CLP sepsis of Kitmutant mice is not due to mast cell deficiency, but to dysbiosis.

      However, the present data are insufficient and incomplete to support the conclusion, and would benefit from more rigorous approaches. With the mechanism part strengthened, this paper would be of interest to researchers on mast cell biology and mucosal immunology.

      We disagree with the view that our data are insufficient and incomplete. Our results demonstrate that mice lacking mast cells (Cpa3<sup>Cre/+</sup> mice) are as resistant to sepsis as wild-type mice, demonstrating that mast cells do not play a detectable role in immunity against sepsis. Additionally, we show that Kit<sup>W/Wv</sup> mice exhibit the same resistance to sepsis as wild-type mice when confronted with the identical intestinal slurry. This finding demonstrates that Kit<sup>W/Wv</sup> mice have no immune deficit in response to sepsis. These central data are both sufficient and complete, given that our data fully address the potential role of mast cells in sepsis. Our study aimed to investigate the role of mast cells in sepsis, not to examine the mechanisms of dysbiosis or associated pathological phenotypes in Kit-mutant controls. We have changed the title to make this point.

      Recommendations:

      (1) The authors showed that E. coli increases in the cecum of Kit-mutant mice, which causes high CLP susceptibility. However, they did not provide any evidence E. coli is responsible for the high susceptibility.

      We showed that E. coli CFUs were increased in the cecum of Kit-mutant mice, but we did not state that this causes CLP susceptibility. We wrote: 'Hence, Kit<sup>W/Wv</sup> microbiota contains high levels of E. coli, which may underlie the observed pathogenicity'. We demonstrated that intestinal slurry from Kit<sup>W/Wv</sup> mice is more pathogenic compared to intestinal slurry from wild-type mice. However, we did not search for or identify the bacterial species that causes this increased pathogenicity because we were addressing the role of mast cells in sepsis. We demonstrate an association of pathogenicity in sepsis experiments with cecal content of pathogenic bacteria (see also the new data on 16S rRNA sequencing). The same argument could be made for each bacterial species identified but this would be very complex experiments (both microbiologically and immunologically) given requirements for bacterial isolation, titration, and considerations of synergism. We therefore refrain from this undertaking.

      In the Figure 3 experiments, the authors administered the same number of cecal bacteria and did not show the number of E. coli after the administration.

      The samples were split and one aliquot was analysed by microbiology and the other aliquot was injected intraperitoneally. Fig. 3d shows the colony-forming units (for Lactobacilli and E coli) from aliquots of cecal slurry used in the intraperitoneal injection experiments shown in Fig. 3a-c. Hence, our data show the colony-forming units that were injected into the mice. It is unclear to us why this is not the key information rather than 'the number of E. coli after the administration'.

      The authors should provide evidence showing that depletion of E. coli decreases susceptibility.

      See response to point 1 above.

      (2) The author should provide direct evidence of dysbiosis by, for example, shotgun sequencing of cecal and fecal contents.

      We performed 16S rRNA sequencing of cecal contents and observed a dysbiotic shift towards an increase of Peptostreptococcaceae, Verrucomicrobiaceae, Coriobacteriaceae, Enterobacteriaceae, and Erysipelotrichaceae in Kit<sup>W/Wv</sup> mice compared to Kit<sup>+/+</sup> controls. None of these changes was observed when comparing the cecal microbiomes of Cpa3<sup>Cre/+</sup> and Cpa3<sup>+/+</sup> mice, indicating that the compositional shift in Kit<sup>W/Wv</sup> mice is due to deficiency in Kit but not mast cells. Of note, as stated on page 14, the microbial changes we observed in Kit<sup>W/Wv</sup> mice resemble dysbiotic patterns reported in chronic intestinal inflammation, experimental colitis, and impaired barrier function.

      These new findings fully align and further support with our earlier conclusion that Kit<sup>W/Wv</sup> mice harbour pro-pathogenic microbiota.

      The new results are display in a new Figure 4, and described on pages 10-11 and discussed on pages 13-14.

      (3) In case the authors find dysbiosis, they should analyze the mechanisms by which Kit mutation causes dysbiosis.

      We have no intention to further explore Kit biology and in particular the intestinal pathophysiology caused by the Kit mutation because any results would not affect the central conclusion of our manuscript (see title). The review process and the revision shall center on making the core of a paper as conclusive as possible, and not widen a paper by requests 'tangential to the main conclusion' (Kaelin Jr. Nature 2017).

      References:

      Caruso, R., Ono, M., Bunker, M. E., Núñez, G. & Inohara, N. Dynamic and Asymmetric Changes of the Microbial Communities after Cohousing in Laboratory Mice. Cell Rep. 27, 3401-3412.e3 (2019).

      Feyerabend, T. B. et al. Deletion of Notch1 Converts Pro-T Cells to Dendritic Cells and Promotes Thymic B Cells by Cell-Extrinsic and Cell-Intrinsic Mechanisms. Immunity 30, 67–79 (2009).

      Feyerabend, T. B. et al. Cre-Mediated Cell Ablation Contests Mast Cell Contribution in Models of Antibody- and T Cell-Mediated Autoimmunity. Immunity 35, 832–844 (2011).

      Feyerabend, T. B., Gutierrez, D. A. & Rodewald, H.-R. Of Mouse Models of Mast Cell Deficiency and Metabolic Syndrome. Cell Metab 24, 1–2 (2016).

      Kaelin Jr, W. G. Publish houses of brick, not mansions of straw. Nature 545, 387– 387 (2017).

      Moore, R. J. & Stanley, D. Experimental design considerations in microbiota/inflammation studies. Clin. Transl. Immunol. 5, e92 (2016).

      Piliponsky, A. M. et al. Basophil-derived tumor necrosis factor can enhance survival in a sepsis model in mice. Nat. Immunol. 20, 129–140 (2019).

      Ridaura, V. K. et al. Gut Microbiota from Twins Discordant for Obesity Modulate Metabolism in Mice. Science 341, 1241214 (2013).

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Suggestions for improved or additional experiments, data, or analyses:

      (1) The study examines only a limited range of gut bacterial species, making it difficult to determine the specific contribution of E. coli expansion to the observed phenotype. A more comprehensive microbial profiling (e.g., 16S rRNA sequencing or metagenomics) would significantly strengthen the conclusions (e.g., Cpa3-mast cell deficient mice, Kit WWv, Kit W-sh/W-sh mice).

      As mentioned above, we now performed 16S rRNA sequencing of cecal contents from Kit<sup>W/Wv</sup> and Cpa3<sup>Cre/+</sup> mice and their respective littermates. Addressing microbiota in Kit<sup>W-sh/W-sh</sup> mice would not add information relevant for our paper.

      (2) The role of impaired peristalsis in KitW/Wv mice as a contributor to microbial dysbiosis and increased E. coli burden should be further explored. Complementary studies using KitW-sh/W-sh or SCF-deficient mice could clarify whether the observed microbiota changes are unique to the W/Wv model.

      We changed the title of the manuscript to emphasize our (unchanged) focus on the immunological role of mast cells in protecting against bacterial sepsis. The responses of Cpa3<sup>Cre</sup> mice clearly ruled out a role of mast cells to these infectious conditions. Our observation of altered microbiota in Kit<sup>W/Wv</sup> mice is consistent with their increased CLP susceptibility. We cite the known peristalsis deficit of Kit<sup>W/Wv</sup> mice as a possible explanation for the microbiota alterations. In the future, other investigators may find it interesting to elucidate the link between Kit mutations and dysbiosis. As stated further above, in our view, these additional questions and potential data have no bearings on the conclusions of our paper.

      (3) In the cohousing experiments, no data are provided to confirm whether microbiota normalization was achieved between groups. Including microbial composition data pre- and post-cohousing would improve the reliability of the interpretation.

      Cohousing made the susceptibility of Kit<sup>W/Wv</sup> and Kit<sup>+/+</sup> mice comparable. Detailed analysis of the extent of microbiota normalization would only make sense to ultimately determine specific taxa or combinations thereof that are responsible for the increased susceptibility of Kit<sup>W/Wv</sup> mice, a question that was never the goal of this study.

      (4) The use of Cpa3 Cre/+ mice introduces potential confounders, as these mice also have defects in basophils and T cells. Functional validation or additional models (e.g., Mas-TRECK or Mcpt5-Cre mice) could help isolate the mast cell-specific effects.

      See our detailed explanation above (Reviewer #1 Public review). It is incorrect to claim that Cpa3<sup>Cre/+</sup> mice have defects in T cells. The reviewer needs to provide published evidence for his/her claim that Cpa3-deficient mice exhibit defects in T cells. We as authors are also obliged to support our claims scientifically, and rightfully so.

      (5) Clarification is needed regarding the role of mast cell-derived TNF. Given previous reports using BMMC reconstitution that implicate mast cells as a source of TNF, reconciling these findings with the current study's results would strengthen interpretation.

      We also disagree here. Experiments in normal unmanipulated mice are inevitably superior to Kit mutants after BMMC reconstitution which is an artificial system. The transplanted cells do not mature normally and don’t settle in their natural niches. We mentioned in the discussion that there is conflicting literature derived from different models and mice.

      But we do not share the expectation that results obtained with Kit-independent models, that differ from previous studies using Kit mutants, necessarily require reconciliation. The experiments are simply incomparable and the most physiologically relevant experiment will pave the way. Research on mast cell functions based on BMMC-reconstituted Kit mutants has meanwhile proven to be unreliable.

      (6) A clearer delineation between mast cell-dependent and microbiota-mediated mechanisms in the discussion sections would enhance readability and impact.

      We restructured the discussion and distinguished between Kit-dependent and mast cell-dependent phenotypes. We also discussed in detail the observed microbiota differences and their influences for the outcomes in the different sepsis experiments.

    1. eLife Assessment

      This study presents a systematic analysis of interactions involving components of the outer membrane protein biogenesis machinery in Escherichia coli. The authors identify interactions between Bam-associated proteins and several pathways. The evidence supporting the principal conclusions is convincing, including several newly identified interactions as well as findings that reinforce previously proposed connections described in the literature. The resulting dataset provides a useful resource for future investigations of bacterial envelope homeostasis and the functions of poorly characterized genes.

    2. Reviewer #3 (Public review):

      In this work, Bryant, et al. investigate genetic interactions between non-essential members of the outer membrane protein biogenesis pathway and other genes in the genome using a transposon-directed insertion sequencing (TraDIS) approach in E. coli K-12. The authors identify interactions with other components of the envelope including LPS, peptidoglycan, and enterobacterial common antigen biogenesis, and they tie these interactions to specific members of the outer membrane biogenesis pathway. Although many of these interactions are known and have been previously investigated in the field, the study provides several synthetic phenotypes that could be useful for further investigations.

      The strengths of the paper include the unbiased, TraDIS approach, and follow up on the interactions observed. The interactions with genes of unknown function also are of interest as they may suggest experiments to find the functions of these genes. Although the paper could better address the relation of its findings to existing literature, the work in the paper is well controlled and the findings will be of interest to the field.

    3. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public Review):

      Weaknesses:

      (1) The cutoffs the authors used to define "conditionally essential" mutants are not reported. The results also lack validation for lethality using a titratable system. It would be ideal to validate several genes in each dataset to determine cutoffs (i.e. 5-fold decrease in insertion mutants) for conditional lethality. It was not done (or described) here.

      We acknowledge that independent validation using targeted mutants would further strengthen the assignment of synthetic lethality. As the primary aim of this study was the genome-wide identification of genetic interactions associated with loss of fitness in several mutant backgrounds, such validation was beyond the scope of the current work. Our experiments identified hundreds of lethal combinations and we have six datasets; therefore, validation of these interactions is not feasible and is indeed not common for a publication using TraDIS to generate leads for the community to follow up on. However, we already validated some of the hits in our original submission and compared our TraDIS data to some known synthetic-lethal interactions. We have also revised the manuscript to describe all other loci as candidate synthetic-lethal interactions and have highlighted the need for future validation studies in the Discussion.

      Regarding the reviewer’s query on thresholds, candidate synthetic-lethal interactions were identified using the tradis_essentiality.R script within the BioTraDIS analytical framework independently on each library: the six mutant backgrounds (DbamB, DbamC, DbamE, DsurA, Dskp, DdegP) and two E. coli BW25113 WT reference sets (an "internal" WT replicate sequenced as part of this study, and an "external" WT dataset from a previous study). This classifies each gene as essential, ambiguous, or non-essential for each library based on the bimodal distribution of insertion indices. Synthetic-lethal gene lists were then built by comparing essentiality classifications between each mutant and the WT sets, which were then flagged as shared/not shared with the internal or external WT essential gene lists in Supplementary Table 1. Therefore, a gene was treated as synthetic-lethal in a given mutant when it was called essential in that mutant but not shared with the WT essentiality call. We have clarified this point on line 923 in the Methods section as follows:

      “We ran the tradis_essentiality.R script within the BioTraDIS package independently on each library: the six mutant backgrounds (DbamB, DbamC, DbamE, DsurA, Dskp, DdegP) and two E. coli BW25113 WT reference sets (an "internal" WT replicate sequenced as part of this study, and an "external" WT dataset from a previous study[95]). This classifies each gene as essential, ambiguous, or non-essential for each library based on the bimodal distribution of insertion indices [30, 34]. Synthetic-lethal gene lists were then built by comparing essentiality classifications between each mutant and the WT sets, which were then flagged as shared/not shared with the internal or external WT essential gene lists in Supplementary Table 1. Therefore, a gene was treated as synthetic-lethal in a given mutant when it was called essential in that mutant but not shared with the WT essentiality call.”

      (2) Also, two mutations that both make the cells sick could provide an additive effect (i.e. dapF and BamB), which doesn't necessarily mean the pathways are linked. The authors should revise their wording. They have not shown genetic linkage in some cases.

      We revised the text to address this on line 693. However, the bamC mutant demonstrates no significant fitness cost under any of the conditions tested in the manuscript. Therefore, if this is simply an additive effect then it is not clear how this occurs, especially in the case of the dapF, bamC double mutant, and we offer an alternative explanation in the Discussion based on interpretation of the literature.

      (3) Mutations throughout the manuscript are not complemented. It would be ideal to add complementation data to show the gene-phenotype relationship is specific.

      We thank the reviewers for highlighting this and have complemented the experiments for the bamB-DNA replication link observation as described in response to reviewer 3.

      (4) Also, I would argue the term "conditionally essential genes" should be replaced with "synthetically lethal". Strains were compared in the same conditions but with different genetic backgrounds.

      We take the reviewer’s point and revised the text throughout.

      Reviewer #2 (Public Review):

      Weaknesses:

      (1) An important control in any genetic interaction study is to do complementation tests to demonstrate that the phenotype observed is indeed due to the missing gene under analysis. Although the Keio library was designed to avoid polar effects, it is impossible to predict other undesirable effects of the deletions (hitting of a non-annotated sRNA or RNA stability effects, for example). Thus, before one can safely conclude that a proposed genetic interaction is real, complementation tests should be carried out. This seems particularly important in the case of a new and surprising interaction, such as that between bamB and DNA replication and repair genes.

      We thank the reviewers for highlighting this and have provided the complementation experiments for the bamB-DNA replication link observations.

      (2) Why not include the suppressor interactions in the work? There are probably plenty, and in principle, they should be as informative as the conditional essential (or synthetic lethal) ones. The only one highlighted in the paper is that between bamB and diaA, since it nicely fits with the synthetic lethal effects with initiation inhibitors seqA and hda. Even if the authors cannot make sense of the suppressor interactions, their inclusion in the paper should make the dataset richer and more valuable to the community.

      Due to the nature of the BioTraDIS pipeline, we focused on gene essentiality and so only picked up potential genes that are essential in the parent but become non-essential in the mutants. This misses observations such as that made for diaA, which we hypothesised and checked manually. The data are publicly available for readers to use for their own studies and we have included some notes in Supplementary Table 1 to explain the filtering process along with another tab including the filtered essential gene lists.

      (3) The enrichment analysis in Figure 2B deserves some clarification. What is the meaning of gene ratio? How can single genes of a pathway yield an enrichment signal? Why weren’t seqA and hda included in the DNA replication class in 2B?

      We thank the reviewer for highlighting this point and realise we did not include a section on this analysis in the Methods section. As such we have included a section on line 935. KEGG pathway enrichment analysis was performed on the conditionally essential gene sets for each mutant background using the enrichKEGG function from the clusterProfiler R package [PMCID: PMC3339379], with the whole E. coli K-12 BW25113 genome used as the background gene set. Gene ratio is defined as the proportion of genes within a given conditionally essential gene set that are annotated to a specific KEGG pathway. Enrichment significance was assessed using a hypergeometric test comparing pathway representation within each query gene set to the whole-genome background.

      SeqA and Hda were not included in the DNA replication enrichment category because the KEGG enrichment analysis was based on existing KEGG pathway annotations, in which these genes are not assigned to the DNA replication pathway despite their well-established roles in replication initiation control. Considering the revision of the results regarding DNA replication, we feel this does not warrant further changes.

      (4) The writing puts too much emphasis on demonstrating that bam lipoproteins and chaperones are specialized instead of fully redundant. However, I have the impression this is a long-settled conclusion in the field, as the manuscript itself describes at several points when reviewing the literature.

      We revised the manuscript throughout to reduce this emphasis.

      Reviewer #3 (Public Review):

      In this work, Bryant, et al. investigate genetic interactions between non-essential members of the outer membrane protein biogenesis pathway and other genes in the genome using a transposon-directed insertion sequencing (TraDIS) approach in E. coli K-12. The authors identify interactions with other components of the envelope including LPS, peptidoglycan, and enterobacterial common antigen biogenesis, and they tie these interactions to specific members of the outer membrane biogenesis pathway. Although many of these interactions are known and have been previously investigated in the field, the study provides several synthetic phenotypes that could be useful for further investigations.

      The strengths of the paper include their unbiased, TraDIS approach, and follow up on the interactions they observe. The interactions with genes of unknown function also are of interest as they may suggest experiments to find the functions of these genes. The largest weakness of this paper is the use of a gene deletion allele for bamB that is known to be polar leading to decreased expression of an essential gene. This largely invalidates all results related to DNA replication. In addition, it is a weakness that the paper does not adequately address its place in the field through discussion of existing results on the interactions they investigate.

      The bamB mutant used here has been widely used in several previous studies (Cox et al., 2017, Gunasinghe et al., 2018, Psonis et al., 2019, Storek et al., 2019, Ranava et al. 2021, Steenhuis et al., 2021, Thewasano et al., 2023) with no concern raised and so we appreciate the reviewer’s expertise here and that they highlighted this issue for us to address.

      We thank the reviewer for highlighting this issue, as we have now completed complementation experiments for the CRISPRi depletion experiments and found that expression of bamB from a pBAD plasmid does not complement the DbamB strain in which seqA or hda is depleted, but expression of der in this system does complement the phenotype. Therefore, we have revised the title and the text to remove discussion of this potential link to DNA replication. We have included the new results and revised the existing DNA replication related figures as new figures S6-S8 and included a brief discussion of this polar effect in lines 283-314. We are very grateful to the reviewer.

    1. eLife Assessment

      This valuable study presents findings on how prokaryotic antibiotics affect translation in mitochondrial ribosomes. Using mitoribosome profiling, the authors provide solid evidence that most tested antibiotics act similarly on bacterial and mitochondrial translation. Additionally, this work shows that alternative translation initiation events exist in two specific mt-mRNAs (MT-ND1 and MT-ND5).

    2. Reviewer #1 (Public review):

      Summary:

      This study aimed to determine whether bacterial translation inhibitors affect mitochondria through the same mechanisms. Using mitoribosome profiling, the authors found that most antibiotics, except telithromycin, act similarly in both systems. These insights could help in the development of antibiotics with reduced mitochondrial toxicity.

      They also identified potential novel mitochondrial translation events, proposing new initiation sites for MT-ND1 and MT-ND5. These insights not only challenge existing annotations but also open new avenues for research on mitochondrial function.

      Strengths:

      Ribosome profiling is a state-of-the-art method for monitoring the translatome at very high resolution. Using mitoribosome profiling, the authors convincingly demonstrate that most of the analyzed antibiotics act in the same way on both bacterial and mitochondrial ribosomes, except for telithromycin. Additionally, the authors report possible alternative translation events, raising new questions about the mechanisms behind mitochondrial initiation and start codon recognition in mammals.

      Weaknesses:

      All the weaknesses I previously highlighted were adequately addressed.