10,000 Matching Annotations
  1. Aug 2026
    1. Reviewer #1 (Public review):

      Summary:

      This manuscript addresses an important question in cardiac biology: whether distinct cardiomyocyte (CM) subpopulations play specialized roles during heart development and regeneration. Using single-cell RNA sequencing and newly generated genetic tools, the authors identify phlda2 as a specific marker of primordial cardiomyocytes in the adult zebrafish heart. They further show that these primordial CMs function are essential for myocardial morphogenesis and coronary vascularization but are dispensable for myocardial regeneration or revascularization after injury. These findings indicate that heart regeneration doesn't simply recapitulate developmental processes.

      Strengths:

      A major strength of the study is the generation of a phlda2 BAC reporter, which provides a specific and reliable marker for primordial cardiomyocytes. The lack of genetic tools has previously limited functional analysis of this CM population. By using phlda2 regulatory elements to generate reporter and NTR-based ablation lines, the authors can visualize and selectively manipulate primordial CMs in vivo. This enables a direct functional interrogation rather than relying on lineage tracing or correlative evidence. Through genetic ablation, the authors convincingly demonstrate that primordial CMs are essential for myocardial morphogenesis and coronary vascular organization during development but are not necessary for heart regeneration.

      Weaknesses:

      (1) The manuscript would benefit from clarifying whether the primordial cardiomyocytes ablation affects epicardial cell behaviors during heart development, given that the well-established role of the epicardium in supporting coronary vessel growth, it is possible that the vascular phenotypes observed after primordial CM ablation may be affected, at least in part, by altered epicardial cells.

      (2) Because primordial cardiomyocytes form a dense, single-cell-thick layer covering the ventricular surface, it would be informative to determine whether their loss alters the spatial distribution or inward migration of coronary endothelial cells or epicardial cells.

      (3) The manuscript carefully examines the relationship between primordial CMs and gata4⁺ cardiomyocytes during regeneration. However, their relationship during heart development should be more fully addressed.

      (4) As loss of cardiomyocytes is known to induce gata4:GFP activation during regeneration, it would be important to determine whether ablation of primordial cardiomyocytes alone triggers gata4:GFP expression in neighboring cardiomyocytes. This analysis would further support the conclusion that primordial cardiomyocytes are not required for regenerative responses.

    2. Reviewer #2 (Public review):

      Summary:

      In the manuscript "Primordial Cardiomyocytes orchestrate myocardial morphogenesis and vascularization but are dispensable for regeneration", Sun et al. identify a novel marker of primordial cardiomyocytes and use it to visualize and ablate the population during development and regeneration. The role of the primordial layer has not been investigated because the tools to manipulate this population have not existed. The manuscript is straightforward, easy to understand, and addresses an important question that has not been explored.

      While the manuscript provides important insights into the role of primordial CMs, backed by a convincing methodology, the authors should clarify their requirements for heart development and maturation. Specifically, is the primordial layer required for the fish to survive? Do primordial CMs regenerate when ablated during development, and do the defects observed (in trabecular and compact CMs and coronary vessels) resolve after 10 days post-treatment when they were detected?

      Strengths:

      The major strengths are the identification of a marker that enables manipulation of primordial cardiomyocytes and the tools generated by the team.

      Weaknesses:

      The major weakness is not considering the longer-term consequences of primordial layer ablation during development, as it is unclear whether the animals succumb to the acute cardiac defects observed or fully recover.

    3. Reviewer #3 (Public review):

      Summary:

      The authors performed single-cell RNA sequencing of adult zebrafish hearts and identified markers for distinct cardiomyocyte subpopulations. One marker, phlda2, marks primordial cardiomyocytes. They generated transgenic reporter lines to characterize phlda2 expression patterns and a phlda2-NTR ablation line to determine the functional requirement of primordial cardiomyocytes during heart regeneration. They found that phlda2+ primordial cardiomyocytes are essential for myocardial morphogenesis and coronary vessel development. Interestingly, when phlda2+ primordial cardiomyocytes are ablated during heart regeneration, gata4+ cortical cardiomyocytes, coronary vessel revascularization, and scar tissue formation are not affected.

      Strengths:

      The authors identified a new primordial cardiomyocyte marker, phlda2. They further demonstrated that primordial cardiomyocytes are important for heart morphogenesis but dispensable for heart regeneration. Their findings reveal a potential difference between heart development and regeneration programs.

      Weakness:

      Despite the interesting findings, the authors did not provide supplemental data for their scRNAseq to demonstrate the data quality and support their conclusions, and some results are not well described.

    4. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This manuscript addresses an important question in cardiac biology: whether distinct cardiomyocyte (CM) subpopulations play specialized roles during heart development and regeneration. Using single-cell RNA sequencing and newly generated genetic tools, the authors identify phlda2 as a specific marker of primordial cardiomyocytes in the adult zebrafish heart. They further show that these primordial CMs function are essential for myocardial morphogenesis and coronary vascularization but are dispensable for myocardial regeneration or revascularization after injury. These findings indicate that heart regeneration doesn't simply recapitulate developmental processes.

      Strengths:

      A major strength of the study is the generation of a phlda2 BAC reporter, which provides a specific and reliable marker for primordial cardiomyocytes. The lack of genetic tools has previously limited functional analysis of this CM population. By using phlda2 regulatory elements to generate reporter and NTR-based ablation lines, the authors can visualize and selectively manipulate primordial CMs in vivo. This enables a direct functional interrogation rather than relying on lineage tracing or correlative evidence. Through genetic ablation, the authors convincingly demonstrate that primordial CMs are essential for myocardial morphogenesis and coronary vascular organization during development but are not necessary for heart regeneration.

      Weaknesses:

      (1) The manuscript would benefit from clarifying whether the primordial cardiomyocytes ablation affects epicardial cell behaviors during heart development, given that the wellestablished role of the epicardium in supporting coronary vessel growth, it is possible that the vascular phenotypes observed after primordial CM ablation may be affected, at least in part, by altered epicardial cells.

      We thank the reviewer for this important suggestion. To address this possibility, we examined epicardial cells in primordial CM-ablated hearts using the epicardial marker tcf21. Surprisingly, we found that epicardial cells rapidly expanded following primordial CM ablation, with an increase already detected at 5 days post-treatment. This increase persisted through at least 30 dpt, when epicardial cell abundance remained higher than that in control hearts. These findings suggest that the coronary vessel defects are unlikely to result from a reduction in epicardial cell number. In addition, we investigated whether loss of the primordial CM layer permits abnormal inward migration of epicardial cells or coronary vessels into the myocardium. However, we observed no evidence of ectopic localization of either cell population following primordial CM ablation. While we cannot exclude the possibility that altered epicardial function or signaling contributes to the vascular phenotype, our results indicate that the vascular defects are not attributable to reduced epicardial cell abundance or abnormal epicardial invasion. We have included these experiments in Page 10 and 11, and Fig. 3I-3J of revised manuscript.

      “Because epicardial cells play critical roles in coronary vessel development, we next asked whether the vascular defects observed following primordial CM ablation were secondary to alterations in the epicardium. We found that tcf21+ epicardial cells revealed an increase rather than a decrease in epicardial cell abundance in primordial CM-ablated hearts (Fig. 3I3J). These findings suggest that the impaired coronary vessel organization is unlikely to result from a reduction in epicardial cell number, although we cannot exclude the possibility that altered epicardial function or signaling contributes to the vascular phenotype.”

      (2) Because primordial cardiomyocytes form a dense, single-cell-thick layer covering the ventricular surface, it would be informative to determine whether their loss alters the spatial distribution or inward migration of coronary endothelial cells or epicardial cells.

      We appreciate the reviewer for this insightful suggestion. Because primordial cardiomyocytes form a continuous single-cell-thick layer at the ventricular surface, we examined whether their ablation affects the spatial distribution or promotes inward migration of epicardial cells or coronary endothelial cells. Using tcf21 and deltaC reporters, we carefully analyzed the localization of these cell populations within the myocardium. We did not observe any evidence of abnormal inward migration or ectopic localization of either epicardial cells or coronary endothelial cells following primordial CM ablation (Fig. 3J–3K). These results indicate that loss of the primordial CM layer does not disrupt tissue compartmentalization or lead to inappropriate cellular invasion into the myocardial interior. We have included these experiments in Page 11, and Fig. 3J-3K of revised manuscript.

      “In addition, because primordial cardiomyocytes form a continuous layer at the ventricular surface, we investigated whether their ablation permits abnormal invasion of epicardial cells or coronary vessels into the myocardium. However, we observed no evidence of ectopic localization of either cell population following primordial CM ablation (Fig. 3J-3K). Thus, loss of the primordial CM layer does not appear to disrupt tissue compartmentalization or permit abnormal cellular invasion into the myocardium.”

      (3) The manuscript carefully examines the relationship between primordial CMs and gata4<sup>+</sup> cardiomyocytes during regeneration. However, their relationship during heart development should be more fully addressed.

      We thank the reviewer for this important suggestion. To further address the relationship between primordial cardiomyocytes and gata4<sup>+</sup> cardiomyocytes during heart development, we examined their spatial and cellular relationship in juvenile zebrafish hearts. Consistent with our observations during regeneration, we did not detect any overlap between phlda2<sup>+</sup> cardiomyocytes and gata4<sup>+</sup> cardiomyocytes in the juvenile heart (7-8 wpf). These results indicate that primordial CMs and gata4<sup>+</sup> proliferative CMs represent distinct cardiomyocyte populations during both heart development and regeneration. We have included these experiments in Page 13, and Fig.5E of revised manuscript.

      “Consistently, no overlap between phlda2<sup>+</sup> and gata4<sup>+</sup> cardiomyocytes was observed in the wpf juvenile zebrafish heart (Fig. 5E).”

      (4) As loss of cardiomyocytes is known to induce gata4:GFP activation during regeneration, it would be important to determine whether ablation of primordial cardiomyocytes alone triggers gata4:GFP expression in neighboring cardiomyocytes. This analysis would further support the conclusion that primordial cardiomyocytes are not required for regenerative responses.

      We appreciate the reviewer for this important suggestion. To determine whether ablation of primordial cardiomyocytes alone is sufficient to activate regenerative signaling, we treated adult phlda2:mCherry-NTR;gata4:EGFP fish and gata4:EGFP siblings with Mtz for 12 hours per day over three consecutive days without ventricular resection. Under these conditions, we did not observe any induction of gata4 expression following primordial CM ablation (Fig. S5). These results indicate that loss of phlda2<sup>+</sup> cardiomyocytes alone is not sufficient to trigger regenerative gata4 activation in the absence of injury, further supporting that primordial CMs are dispensable for activation of the regenerative response. We have included these experiments in Page 11 and Fig. S5 of revised manuscript.

      “To determine whether ablation of primordial cardiomyocytes is sufficient to activate regenerative signaling, we first treated adult phlda2:mCherry-NTR;gata4:EGFP fish and control gata4:EGFP siblings with Mtz for 12 hours per day over three consecutive days without ventricular resection. We did not observe any induction of gata4:EGFP expression following primordial CM ablation (Fig. S5), indicating that loss of phlda2<sup>+</sup> cardiomyocytes is not sufficient to trigger regenerative gata4 activation in the absence of injury.”

      Reviewer #2 (Public review):

      Summary:

      In the manuscript "Primordial Cardiomyocytes orchestrate myocardial morphogenesis and vascularization but are dispensable for regeneration", Sun et al. identify a novel marker of primordial cardiomyocytes and use it to visualize and ablate the population during development and regeneration. The role of the primordial layer has not been investigated because the tools to manipulate this population have not existed. The manuscript is straightforward, easy to understand, and addresses an important question that has not been explored.

      While the manuscript provides important insights into the role of primordial CMs, backed by a convincing methodology, the authors should clarify their requirements for heart development and maturation. Specifically, is the primordial layer required for the fish to survive?

      We thank the reviewer for this important question. We found that efficient ablation of phlda2<sup>+</sup> primordial cardiomyocytes does not affect overall survival of zebrafish under standard laboratory conditions. Although these animals exhibit clear defects in cardiac structure and coronary vascular organization, they remain viable during the experimental period, indicating that the primordial CM layer is not essential for survival. While we did not assess detailed physiological parameters such as cardiac function, swimming behavior, or long-term fitness in this study, the observed structural abnormalities suggest that subtle functional consequences may exist. These aspects will be important directions for future investigation. We have included the description on page 14, paragraph 2.

      “Although primordial CM ablation does not affect survival under laboratory conditions, the observed structural defects may have functional consequences on cardiac performance and overall physiological fitness, which warrant further investigation.”

      Do primordial CMs regenerate when ablated during development, and do the defects observed (in trabecular and compact CMs and coronary vessels) resolve after 10 days posttreatment when they were detected?

      We appreciate the reviewer for this important question. To determine whether primordial cardiomyocytes regenerate following ablation during development, we performed Mtz-mediated ablation in juvenile zebrafish (7–8 wpf) and examined the hearts at extended time points after treatment. We found that phlda2<sup>+</sup> cardiomyocytes did not recover even at 90 days post-treatment, indicating a persistent loss of this population following developmental-stage ablation (Fig. S6). Importantly, we further assessed whether the cardiac defects observed at earlier time points resolve over time. We found that the abnormalities in trabecular and compact myocardium, as well as coronary vessel organization, persisted at 90 days post-treatment and did not show evidence of recovery (Fig. S4C). These findings demonstrate that the observed defects are not transient developmental delays but represent long-lasting structural alterations of the heart following primordial CM ablation.

      Major Comments:

      (1) Figure 1: A more detailed characterization of the three CM populations would be helpful in the text as well as a new Supplemental Excel Data Sheet with the top unique genes expressed in each.

      We thank the reviewer for this helpful suggestion. In the revised manuscript, we have expanded the description of the three cardiomyocyte populations in the Results section to provide a more detailed characterization. We have revised the description on page 8. In addition, we have added a new Supplemental Excel Data Sheet (Table S1) listing the top differentially expressed genes for each CM cluster.

      “Notably, Cluster 2 showed additional enrichment for pathways involved in mitochondrial respiratory chain assembly, ATP synthesis, TCA cycle, and ribosome biogenesis, suggesting a relatively higher metabolic and biosynthetic activity state. In contrast, Cluster 1 was enriched for GO terms associated with mitochondrial stress responses, protein degradation, and cytoprotective pathways, indicating a stress-adapted cardiomyocyte state. Cluster 3 showed reduced enrichment of metabolic pathways, consistent with an immature metabolic profile, and was further enriched for genes involved in muscle development and epithelial morphogenesis, suggesting a role in cardiac morphogenesis and tissue organization.”

      (2) Figure 3: Lower magnification views of control and MTZ-treated hearts are needed for all reporters shown (cmlc2, gata4, deltaC) at 10 days post-treatment. These "whole heart" views will enable the reader to get a gross sense of how disrupted heart development is following ablation of the primordial layer.

      We appreciate the reviewer for this helpful suggestion. We have added low-magnification whole-heart images for all reported markers (cmlc2, gata4, and deltaC) to better illustrate the overall cardiac morphology following ablation of the primordial layer. We have included these data in Fig. S3 of revised manuscript.

      (3) It should also be stated clearly in the Results section text what day post-fertilization Mtx treatment began. A section on Mtx treatment, including timing (what day was it applied and what day was it washed out) and dose, should be added to the methods section.

      We thank the reviewer for this helpful suggestion. In response, we have now clearly stated the timing of Mtz treatment in the Results section in page 9, 10 and 11 of revised manuscript.

      “To address the role of phlda2<sup>+</sup> cells during heart development, we performed the following experiments using a standardized Mtz treatment protocol (see Methods), with identical treatment conditions applied to 7-8 wpf juvenile zebrafish.”

      “To address the role of primordial cells during heart regeneration, we perform the below experiments using standardized Mtz treatment protocol (see Methods), with identical treatment conditions applied to adult zebrafish (4–6 months post-fertilization).”

      In addition, we have added a Mtz treatment section in the Methods, which now includes detailed information on dosage, duration, and washout schedule to ensure full reproducibility of the experiments.

      “Mtz treatment

      For conditional ablation of phlda2<sup>+</sup> cardiomyocytes, zebrafish expressing phlda2:mCherryNTR were treated with 10 mM metronidazole (Mtz) for 12 hours per day for three consecutive days. Fish were washed out and maintained in fresh system water after each daily treatment. For developmental analyses, juvenile zebrafish (7–8 weeks post-fertilization) were subjected to Mtz treatment as described above. Following completion of Mtz exposure, fish were maintained under standard conditions. Hearts were collected at 10 days post-treatment for assessment of gata4 activation, and at 30 days post-treatment for analysis of myocardial structure and coronary vessel development. For regeneration experiments, adult zebrafish (4–6 months old) were similarly treated with Mtz for three consecutive days with daily washout. Three days after the final Mtz treatment, ventricular apex resection was performed. Hearts were harvested at 7 days post-amputation (dpa) for analysis of gata4 activation and early regenerative responses, and at 30 dpa for evaluation of myocardial regeneration and coronary vessel revascularization.”

      (4) What happens to the heart 2 and 6 months post-treatment? Are there long-term consequences to primordial layer ablation or do the defects seen at 10 days post-treatment eventually resolve?

      We thank the reviewer for this important question. To determine whether the developmental defects observed following primordial CM ablation are transient or persist long term, we performed additional analyses at later time points after Mtz washout. First, we found that phlda2<sup>+</sup> cardiomyocytes failed to recover following Mtz-mediated ablation. In adult phlda2:mCherry-NTR fish, phlda2<sup>+</sup> cells remained absent 30 days after Mtz washout (Fig. 5B). Similarly, when juvenile fish (7–8 wpf) were treated with Mtz and subsequently allowed to recover, phlda2<sup>+</sup> cells were still not restored at 90 days post-treatment (Fig. S6). Second, juvenile zebrafish (7–8 wpf) were subjected to Mtz-mediated primordial CM ablation and analyzed 90 days after Mtz washout. We found that vascular abnormalities persisted long after ablation of primordial CMs (Fig. S4C). Coronary vessels remained disorganized and fragmented, indicating that the vascular phenotype does not resolve over time. Together, these findings demonstrate that primordial CM ablation causes long-lasting defects and that the abnormalities observed are not transient developmental delays. Instead, loss of primordial CMs results in persistent cellular and vascular defects that remain evident months after Mtz treatment. We have included these experiments in Fig. S4C and Fig. S6 of revised manuscript.

      “Notably, these vascular abnormalities persisted at 90 days post-treatment, indicating that the defects do not resolve during subsequent cardiac growth and maturation (Fig. S4C).”

      “Similarly, when juvenile zebrafish (7–8 wpf) were subjected to Mtz-mediated ablation and analyzed 90 days after treatment, phlda2<sup>+</sup> cells remained absent, demonstrating a persistent failure of primordial CM recovery (Fig. S6).”

      (5) Also, does the primordial layer come back in these animals where the lineage is ablated during development? Or is it permanently lost as shown in Figure 5 when it is ablated during adulthood?

      We appreciate the reviewer for raising this important question. To determine whether the primordial layer can be reestablished following ablation during development, we treated juvenile phlda2:mCherry-NTR fish (7–8 wpf) with Mtz and examined hearts 90 days after treatment. We found that phlda2<sup>+</sup> cardiomyocytes remained absent at this late time point (Fig. S6), indicating that the primordial layer does not recover following developmental-stage ablation. We have included these experiments in Page 12, and Fig. S6 of revised manuscript.

      “Similarly, when juvenile zebrafish (7–8 wpf) were subjected to Mtz-mediated ablation and analyzed 90 days after treatment, phlda2<sup>+</sup> cells remained absent, demonstrating a persistent failure of primordial CM recovery (Fig. S6).”

      (6) Figure 4: Need to show that phlda2 reporter fluorescence in lost/reduced following Mtz treatment during adulthood before apex amputation.

      We thank the reviewer for this important suggestion. We agree that confirming efficient ablation of phlda2<sup>+</sup> cardiomyocytes in adult fish prior to regeneration analysis is essential.

      In our study, we have already demonstrated in Fig.5B that Mtz treatment in adult phlda2:mCherry-NTR fish results in efficient and sustained loss of phlda2<sup>+</sup> cells, with no detectable recovery at 7 and 30 days post-treatment. These data confirm robust ablation of the primordial CM population following Mtz treatment in adults. Therefore, additional redundant imaging prior to apex resection was not performed in Fig 4.

      (7) Figure 5: It is interesting that primordial CMs do not regenerate following apex amputation or genetic ablation. This result suggests that primordial CMs are only important during development and dispensable during adulthood? This result also makes me question whether primordial CMs are actually required for heart development, which is why it is important to address whether the fish recovers.

      We appreciate the reviewer for this insightful comment. Our data indicate that primordial cardiomyocytes are essential for proper heart development, as their ablation during juvenile stages leads to significant structural and vascular abnormalities. Importantly, we further examined long-term outcomes and found that these defects do not resolve over time. Juvenile zebrafish subjected to primordial CM ablation failed to recover phlda2<sup>+</sup> cardiomyocytes even at 90 days post-treatment, and coronary vascular abnormalities also persisted at this late stage (Fig. S4C). These findings indicate that the observed developmental defects are not transient delays but instead reflect long-lasting structural alterations of the heart. In addition, we found that adult zebrafish similarly fail to regenerate primordial cardiomyocytes following genetic ablation (Fig. 5B), further supporting the limited regenerative capacity of this population. Together, these data demonstrate that primordial cardiomyocytes are required for proper cardiac development, and their loss leads to persistent defects that are not reversed during subsequent growth or regeneration.

      Minor:

      Line 234: Did the authors mean to write Cluster 3 (instead of Cluster 2)?

      We thank the reviewer for pointing out this error. We confirm that this was a labeling mistake, and “Cluster 3” is correct. The text has been corrected in the revised manuscript.

      Line 265: There is a typo of some sort in the phrase, "54.7% reduction closed to the ventricular wall".

      We thank the reviewer for pointing out this error. We have changed the description on page 10 of the revised manuscript.

      “The compact myocardium was disorganized compared with controls, and trabecular muscle formation was severely impaired, with an approximately 54.7% reduction in trabecular area, predominantly observed in regions adjacent to the ventricular wall (Fig. 3A, 3B and S3A).”

      Is there a corollary lineage in mammals? This should be addressed in the Introduction or Discussion.

      We thank the reviewer for this insightful suggestion. At present, a direct corollary lineage to zebrafish phlda2<sup>+</sup> primordial cardiomyocytes have not been clearly defined in mammals. However, mammalian hearts also contain heterogeneous cardiomyocyte populations with distinct developmental states and metabolic profiles, including immature or embryonic-like cardiomyocytes that persist in specific regions during development and early postnatal stages. These populations may share functional similarities with the zebrafish primordial CMs in terms of developmental organization and maturation roles. We have now discussed this point in the Discussion and emphasized that whether a comparable lineage exists in mammals remains an important open question for future studies. We have included the discussion on page 15 of the revised manuscript.

      “Although a direct corollary of phlda2<sup>+</sup> primordial cardiomyocytes has not yet been identified in mammals, mammalian hearts contain heterogeneous cardiomyocyte populations with immature states. Whether these populations represent a functional equivalent of zebrafish primordial CMs remains an open question and needs further investigation.”

      Reviewer #3 (Public review):

      Summary:

      The authors performed single-cell RNA sequencing of adult zebrafish hearts and identified markers for distinct cardiomyocyte subpopulations. One marker, phlda2, marks primordial cardiomyocytes. They generated transgenic reporter lines to characterize phlda2 expression patterns and a phlda2-NTR ablation line to determine the functional requirement of primordial cardiomyocytes during heart regeneration. They found that phlda2+ primordial cardiomyocytes are essential for myocardial morphogenesis and coronary vessel development. Interestingly, when phlda2+ primordial cardiomyocytes are ablated during heart regeneration, gata4+ cortical cardiomyocytes, coronary vessel revascularization, and scar tissue formation are not affected.

      Strengths:

      The authors identified a new primordial cardiomyocyte marker, phlda2. They further demonstrated that primordial cardiomyocytes are important for heart morphogenesis but dispensable for heart regeneration. Their findings reveal a potential difference between heart development and regeneration programs.

      Weakness:

      Despite the interesting findings, the authors did not provide supplemental data for their scRNAseq to demonstrate the data quality and support their conclusions, and some results are not well described.

      We appreciate the reviewer for this important suggestion. In the revised manuscript, we have added supplemental data to support the scRNA-seq analysis, including gene expression tables for each cardiomyocyte cluster (Table S1), and full GO-term enrichment results (Table S2). In addition, we have revised the Results section to improve the clarity and description of the scRNA-seq findings. Please see detailed responses below for point-by-point clarification.

      Reviewer #3 (Recommendations for the authors):

      (1) The authors did not provide enough data to demonstrate the quality of their scRNAseq. They only mentioned that they obtained "high-quality" transcriptomics. Specific parameters such as how many total reads and reads per cell should be provided.

      We thank the reviewer for this important suggestion. In the revised manuscript, we have added detailed sequencing quality metrics to the Methods section in Page 6 of the revised manuscript. The dataset contains 136,174,297 total reads with an average sequencing depth of 36,168 reads per cell.

      “The newly generated scRNA-seq data yielded 136,174,297 total reads with an average sequencing depth of 36,168 reads per cell.”

      (2) The authors utilized cmlc2:EGFP fish to perform scRNASeq. It will be helpful to include feature plots of cmlc2 and EGFP transcripts.

      We thank the reviewer for this suggestion. We have now included the description in Page 8, and feature plots of cmlc2 transcripts and EGFP reporter expression in the scRNA-seq dataset as a supplementary figure (Fig. S1A and S1B).

      “The expression of cmlc2 transcripts and EGFP reporter signal in the single-cell dataset further confirmed the enrichment of cardiomyocytes (Fig. S1A and S1B)”

      (3) The authors show that notch 3 is in cluster 3 of cardiomyocytes and suggest that this reflects elevated NOTCH signaling activity. The authors might consider using RNAScope to further validate that Notch 3 is expressed in cardiomyocytes. It will be also helpful to confirm phlda2 expression patterns during zebrafish heart development and regeneration by RNAScope.

      We appreciate the reviewer for this helpful suggestion. To further examine the spatial expression patterns of these genes, we analyzed publicly available spatial transcriptomic data from zebrafish hearts. We found that phlda2 is enriched in the outer region of the heart during both uninjured and regenerating conditions. Similarly, notch3 and actn1 were also predominantly localized to the outer heart region in the uninjured heart. These findings are consistent with our scRNA-seq results and support the spatially restricted signature of Cluster 3 cardiomyocytes. We have included these data in Page 8, and Fig. S2A-S2C of revised manuscript.

      “To further validate their spatial distribution, analysis of previously published spatial transcriptomic data revealed that phlda2, notch3, and actn1 were predominantly expressed in the outer region of the heart (Fig.S2A-S2C).”

      (4) The authors did not provide any data as supplemental tables to support their analyses of scRNAseq and GO-term analysis.

      We thank the reviewer for this suggestion. In the revised manuscript, we have added new supplemental tables providing full support for the scRNA-seq and GO-term analyses (Table S1 and S2), including lists of differentially expressed genes for each cardiomyocyte cluster and the corresponding GO enrichment results.

      (5) The description of the phenotype in Fig. 3A and B is not clear, especially for the sentence "trabecular muscle formation was severely impaired showing an approximately 54.7% reduction close to the ventricular wall.". The authors might consider using a bracket to show the compact muscle and trabecular muscle and the distance to the ventricular wall.

      We appreciate the reviewer for this suggestion. We have added brackets in Fig. 3A to label the compact and trabecular myocardium for improved clarity. Regarding “distance to the ventricular wall,” we found that this measurement varies substantially across different regions within the same heart, making a single distance-based metric unreliable. Therefore, we quantified trabecular muscle using the trabecular area fraction (trabecular area/total ventricular area) within a defined region of interest as a robust and unbiased indicator. The reported 54.7% reduction refers to this area fraction, and we have revised the text accordingly for clarity in Page 10 of the revised manuscript.

      “The compact myocardium was disorganized compared with controls, and trabecular muscle formation was severely impaired, with an approximately 54.7% reduction in trabecular area, predominantly observed in regions adjacent to the ventricular wall (Fig. 3A, 3B and S3A).”

      (6) It is not clear how the authors quantify the vessel "length"/ventricular area and found that there is no difference (Fig. 3G). The vessel length is significantly shorter in the images (Fig. 3F) as the author also indicated that the vessels are fragmented.

      We thank the reviewer for this important comment. Coronary vessel “length” was quantified by selecting a fixed region of interest (ROI) within the ventricular area, followed by skeletonization of deltaC:EGFP<sup>+</sup> vessels using ImageJ. The total vessel length within the ROI was measured and normalized to the ROI area to obtain vessel length density. Although the representative images (Fig. 3F) show a more fragmented vascular pattern, the total summed vessel length within the defined region was not reduced. This indicates that primordial CM ablation primarily affects vascular organization rather than overall vessel length within the ventricular area. We have updated the figure legend for clarity of the revised manuscript

      “Vessel length was measured within a fixed region of interest (ROI) after skeletonization of deltaC:EGFP<sup>+</sup> vessels in ImageJ.”

    1. eLife Assessment

      This important study identifies PRRT2 as an auxiliary regulator of Nav channel slow inactivation in vitro and in vivo, demonstrating that PRRT2 facilitates entry into, and delays recovery from, the slow-inactivated state. The revised manuscript has been substantially strengthened, providing compelling evidence that PRRT2 is relevant to normal brain physiology and disease pathophysiology, providing a mechanistic link between PRRT2 mutations and episodic neurological phenotypes. Overall, this study will be of interest to ion channel biophysicists and neurophysiologists, particularly those studying channelopathies.

    2. Reviewer #1 (Public review):

      [Editors' note: this version has been assessed by the Reviewing and Senior Editor without further input from the original reviewers. The authors have addressed the comments raised in the previous round of review.]

      Summary:

      The manuscript by Lu and colleagues demonstrate convincingly that PRRT2 interacts with brain voltage-gated sodium channels to enhance slow inactivation in vitro and in vivo. The work is interesting and rigorously conducted. The relevance to normal physiology and disease pathophysiology (e.g., PRRT2-related genetic neurodevelopmental disorders) seems high. Some simple additional experiments could elevate the impact and make the study more complete.

      Strengths:

      Experiments are conducted rigorously including experimenter blinding and appropriate controls. Data presentation is excellent and logical. The paper is well written for a general scientific audience.

      Comments on revised version.

      The manuscript by Lu and colleagues has been revised sufficiently to address all my prior concerns.

      Experiments are conducted rigorously including experimenter blinding and appropriate controls. Data presentation is excellent and logical. The paper is well written for a general scientific audience.

    3. Reviewer #2 (Public review):

      Summary:

      As a member of DspB subfamily, PRRT2 is predominantly expressed in CNS and has been associated with various paroxysmal neurological disorders. Previous studies have shown that PRRT2 interacts with Nav and Cav channels, modulating channel properties and neuronal excitability.

      In this manuscript, Lu et al. demonstrate that PRRT2 is a potent regulator of Nav channel slow inactivation, promoting the development of Nav slow inactivation and impeding the recovery from slow inactivation. This effect is highly conserved in PRRT2s across species as well as among DspB family members (TRARG1 and TMEM233). The authors further confirmed the interaction between Nav channels and PRRT2 in heterologous expression systems as well as in Prrt2-V5 knock-in mice. Prrt2-mutant mice, which lack PRRT2 expression, require lower stimulation thresholds for evoking after-discharges when compared with WT mice.

      Overall, this is a well-executed and methodologically comprehensive study. This work offers valuable insight into the physiological functions of PRRT2 and reveals a potential pathogenic mechanism underlying PRRT2-associated neurological disorders.

      The revised manuscript has addressed most of the concerns raised by the reviewers and has been substantially strengthened, although I still have several concerns regarding the discussion section.

      Strengths:

      (1) Overall, this is a well-executed and methodologically comprehensive study. The electrophysiological data strongly support the conclusion that PRRT2 is a potent regulator of Nav channel slow inactivation. The observation that this regulation is conserved in PRRT2 across species and among DspB family members raises the possibility that altered regulation of Nav channels may also contribute to the pathogenesis of TRARG1- or TMEM233-associated disorders.

      (2) Co-immunoprecipitation assay performed using brain tissue from genetically modified Prrt2-V5 knock-in mice provides convincing in vivo evidence for the interaction between PRRT2 and Nav1.2 channels.

      (3) Prrt2-V5 KI mice show markedly reduced PRRT2 protein expression and display phenotypes similar to those observed in Prrt2-mutant mice, supporting an important role of PRRT2 in regulating neuronal and network excitability.

      Weaknesses:

      (1) Nav1.6 is also highly expressed in cortical neurons and is widely regarded as a major contributor to action potential initiation and sustained high-frequency firing. Given that PRRT2 similarly regulates the fast and slow inactivation of Nav1.6 and Nav1.2 channels, the potential contribution of Nav1.6 regulation to neuronal and network excitability should be discussed.

      (2) Slow inactivation is generally considered to develop over timescales ranging from hundreds of milliseconds to seconds or longer. Therefore, the statement in Discussion (Page 13, line 381-382) that "slow inactivation develops on a timescale of tens of milliseconds to seconds" may not accurately reflect the conventional kinetic definition of slow inactivation and should be clarified.

      (3) Page 14, line 417-430: "question about how Nav channel slow inactivation is regulated in cells that do not express PRRT2".<br /> PRRT2 is unlikely to be the sole regulator of Nav channel slow inactivation. Other molecules and signaling pathways may regulate Nav channel and contribute to neuronal excitability. In addition, neuronal excitability can also be regulated through modulating other Nav properties, such as long-term inactivation or slow recovery from inactivation, as well as through modulating the activity of other ion channels, for example, Kv7.2 and Kv7.3 channels. Therefore, PRRT2-negative cells may utilize alternative mechanisms to fine-tune neuronal excitability. In its current form, this paragraph somewhat overstates the role of PRRT2 and would benefit from a more balanced discussion.

      (4) Page 50, Figure 7-figure supplement 2: It would be helpful to include representative traces of the 1st and the last (20th) compound APs in panels B and C.

    4. Reviewer #3 (Public review):

      This paper reveals that the neuronal protein PRRT2, previously known for its association with paroxysmal dyskinesia and infantile seizures, modulates the slow inactivation of voltage-gated sodium ion (Nav) channels, a gating process that limits excitability during prolonged activity. Using electrophysiology, molecular biology, and mouse models, the authors show that PRRT2 accelerates entry of Nav channels into the slow-inactivated state and slows their recovery, effectively dampening excessive excitability. The effect seems evolutionarily conserved, requires the C-terminal region of PRRT2, and is recapitulated in cortical neurons, where PRRT2 deficiency leads to hyper-responsiveness and reduced cortical resilience in vivo. These findings extend the functional repertoire of PRRT2, identifying it as a physiological brake on neuronal excitability. The work provides a mechanistic link between PRRT2 mutations and episodic neurological phenotypes.

      Comments:

      (1) The precise structural interface and the molecular basis of gating modulation remain inferred rather than demonstrated.

      (2) The in vivo phenotype reflects a complex circuit outcome and does not isolate slow-inactivation defects per se.

      (3) Expression of PRRT2 in muscle or heart is low, so the cross-isoform claims are likely of limited physiological significance.

      (4) The mechanistic separation between trafficking of PRRT2 and its gating effects is not clearly resolved.

      (5) Additional studies with Nav1.6 should be carried out.

      Comments on revised version.

      These comments have been addressed in the revised version.

    5. Author response:

      The following is the authors’ response to the original reviews.

      Reviewer #2 (Public review):

      Summary:

      As a member of DspB subfamily, PRRT2 is predominantly expressed in CNS and has been associated with various paroxysmal neurological disorders. Previous studies have shown that PRRT2 interacts with Nav and Cav channels, modulating channel properties and neuronal excitability.

      In this manuscript, Lu et al. demonstrate that PRRT2 is a potent regulator of Nav channel slow inactivation, promoting the development of Nav slow inactivation and impeding the recovery from slow inactivation. This effect is highly conserved in PRRT2s across species as well as among DspB family members (TRARG1 and TMEM233). The authors further confirmed the interaction between Nav channels and PRRT2 in heterologous expression systems as well as in Prrt2-V5 knock-in mice. Prrt2-mutant mice, which lack PRRT2 expression, require lower stimulation thresholds for evoking after-discharges when compared with WT mice.

      Overall, this is a well-executed and methodologically comprehensive study. This work offers valuable insight into the physiological functions of PRRT2 and reveals a potential pathogenic mechanism underlying PRRT2-associated neurological disorders.

      The revised manuscript has addressed most of the concerns raised by the reviewers and has been substantially strengthened, although I still have several concerns regarding the discussion section.

      Strengths:

      (1) Overall, this is a well-executed and methodologically comprehensive study. The electrophysiological data strongly support the conclusion that PRRT2 is a potent regulator of Nav channel slow inactivation. The observation that this regulation is conserved in PRRT2 across species and among DspB family members raises the possibility that altered regulation of Nav channels may also contribute to the pathogenesis of TRARG1- or TMEM233-associated disorders.

      (2) Co-immunoprecipitation assay performed using brain tissue from genetically modified Prrt2-V5 knock-in mice provides convincing in vivo evidence for the interaction between PRRT2 and Nav1.2 channels.

      (3) Prrt2-V5 KI mice show markedly reduced PRRT2 protein expression and display phenotypes similar to those observed in Prrt2-mutant mice, supporting an important role of PRRT2 in regulating neuronal and network excitability.

      We sincerely thank the reviewer for the meticulous evaluation of our revised manuscript and for the constructive and insightful comments.

      Weaknesses:

      (1) Nav1.6 is also highly expressed in cortical neurons and is widely regarded as a major contributor to action potential initiation and sustained high-frequency firing. Given that PRRT2 similarly regulates the fast and slow inactivation of Nav1.6 and Nav1.2 channels, the potential contribution of Nav1.6 regulation to neuronal and network excitability should be discussed.

      We appreciate the reviewer’s suggestion. In the revised manuscript, we have clarified that PRRT2-mediated regulation of Nav1.2, together with its regulation of Nav1.6, may contribute to neuronal and network excitability in the cortex. Please refer to Page 15, Line 439-440.

      (2) Slow inactivation is generally considered to develop over timescales ranging from hundreds of milliseconds to seconds or longer. Therefore, the statement in Discussion (Page 13, line 381-382) that "slow inactivation develops on a timescale of tens of milliseconds to seconds" may not accurately reflect the conventional kinetic definition of slow inactivation and should be clarified.

      We thank the reviewer for this comment. We have corrected the timescale description in the Discussion accordingly. Please refer to Page 13, Line 382.

      (3) Page 14, line 417-430: "question about how Nav channel slow inactivation is regulated in cells that do not express PRRT2".

      PRRT2 is unlikely to be the sole regulator of Nav channel slow inactivation. Other molecules and signaling pathways may regulate Nav channel and contribute to neuronal excitability. In addition, neuronal excitability can also be regulated through modulating other Nav properties, such as long-term inactivation or slow recovery from inactivation, as well as through modulating the activity of other ion channels, for example, Kv7.2 and Kv7.3 channels. Therefore, PRRT2-negative cells may utilize alternative mechanisms to fine-tune neuronal excitability. In its current form, this paragraph somewhat overstates the role of PRRT2 and would benefit from a more balanced discussion.

      We appreciate the reviewer’s constructive suggestion. We agree that PRRT2 is unlikely to be the sole regulator of neuronal excitability or Nav channel slow inactivation. In the revised Discussion, we have clarified that PRRT2-dependent regulation of Nav channel slow inactivation represents one mechanism among several that fine-tune neuronal excitability. Other mechanisms, including regulation of potassium channels such as Kv7.2/Kv7.3, and other ion channel- or signaling- dependent pathways, may also contribute to excitability control in both PRRT2-positive and PRRT2-negative neurons. We have revised relevant paragraph to provide a more balanced discussion of alternative mechanisms. Please refer to Page 15, Line 430-438.

      (4) Page 50, Figure 7-figure supplement 2: It would be helpful to include representative traces of the 1st and the last compound APs in panels B and C.

      We appreciate the reviewer’s valuable suggestion. We have now added representative traces of the first and last compound action potentials, corresponding to the 1st and the 100th responses, respectively, to panels B and C of Figure 7-figure supplement 2. Please refer to Page 51, Figure 7-figure supplement 2.

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      (1) Page 19, line 546: a sampling rate of 10kHz was used for recording Nav currents. Because the Nav channel isoforms examined in this manuscript (Nav1.1, Nav1.2, Nav1.4, Nav1.5, and Nav1.6) exhibit extremely rapid activation and inactivation kinetics, the temporal resolution provided by a 10 kHz sampling rate may not be optimal for detailed kinetic analysis. Although I am not requesting additional experiments, the authors may wish to choose a higher sampling rate (e.g., 50 kHz) in future Nav channel studies.

      We thank the reviewer for this helpful suggestion regarding the sampling rate for Nav current recordings. We will adopt higher sampling rates for rapid kinetic analyses of sodium currents in our future work.

      (2) Typo: Page 24, line 712-713: should "a Digidata (Molecular Devices, 1332A)" be "... 1322A"?

      We thank the reviewer for pointing out this typo. We have corrected “Digidata 1332A” to “Digidata 1322A” in the relevant Method section of revised manuscript. Please refer to Page 25, Line 722.

    1. eLife Assessment

      The authors describe the results of a longitudinal study of pertussis infection in mother/infant dyads in Lusaka, Zambia. This work represents an important contribution to our understanding of the global burden of pertussis. The evidence presented is compelling, with strengths including routine sampling irrespective of symptoms and rigorous qPCR methodology. The study is a significant and much-needed contribution, which sheds light on an often-overlooked dimension of pertussis transmission and opens avenues for future research and policy consideration.

    2. Reviewer #1 (Public review):

      Summary:

      The study investigates the role of asymptomatic pertussis carriage in transmission between mothers and their infants in particular. The authors use a longitudinal cohort study that involved 1,315 mother-infant dyads in Lusaka Zambia and they utilized qPCR based detection of IS481 to track Bordetella pertussis transmission over time. Insights from the study suggest that minimally symptomatic or asymptomatic mothers may act as a reservoir for B. pertussis transmission in the infants thus challenging the traditional surveillance methods that focus on symptomatic cases. Additionally, the study also identified a subgroup of persistently colonized individuals where mothers were majorly asymptomatic despite sustained bacterial presence.

      The authors aimed to improve comprehension of pertussis transmission dynamics in high burden low resource settings and they advocated for an enhanced molecular surveillance strategies to capture full pertussis infection including those that might have gone undetected.

      Strengths:

      The strength are the use of innovative study design especially the longitudinal approach and routine sampling rather than symptom driven testing that minimizes bias in the study. The methodology were also rigorous and transparent by evaluating IS481 signal strength to classify pertussis detection and conducts retesting to assess qPCR reliability. There was also important epidemiological insights and the findings challenge the traditional wisdom by suggesting that pertussis transmission may frequently occur outside of symptomatic cases. The findings also showed its relevance to global health and policy by arguing for the incorporation of molecular tools like qPCR for surveillance of pertussis in low resource setting.

      Weaknesses:

      These includes reliability on qPCR based detection without additional validation measures like confirmatory culture or serology. There are also potential alternate explanation for transmission patterns observed in the study such as shared environmental exposure or household transmission. Additionally, there are limited generalizability as the study was done in a single urban site in Zambia. There is also lack of functional immune data.

    3. Reviewer #2 (Public review):

      Summary:

      In this paper, the authors describe the results of a longitudinal study of pertussis infection in mother/infant dyads in Lusaka, Zambia. Unlike many past studies, the authors assessed the infection status of individuals independently of whether they were symptomatic for a respiratory infection. As a result, this work represents one of the first studies specifically designed to assess asymptomatic transmission of pertussis. Using qPCR, the authors find strong evidence for the role of asymptomatic transmission from mothers to infants and also evidence for long-term bacterial carriage. This work represents an important contribution to our understanding of the global burden of pertussis. Also, it highlights the still under-appreciated role of asymptomatic transmission across many infectious diseases (including vaccine-preventable ones).

      Strengths:

      Unlike many past studies, the authors assessed the infection status of individuals independently of whether they were symptomatic for a respiratory infection. As a result, this work represents one of the first studies specifically designed to assess asymptomatic transmission of pertussis. Using qPCR, the authors find strong evidence for the role of asymptomatic transmission from mothers to infants and also evidence for long-term bacterial carriage.

      Comments on revised version:

      I appreciate the authors' attention to my comments during the revision process and still believe that their work represents an important contribution to our understanding of pertussis epidemiology. In most cases, the authors have done a thorough job of either addressing or responding to my comments. However, I do not believe the authors engaged sufficiently with two of the queries raised in my previous round of comments. The two queries were about the vaccination status of the mothers and engagement with literature on asymptomatic transmission. I still think they matter and ask the authors to consider them again.

      I do not think the authors can rule out two alternative explanations: (1) recent introduction of pertussis and low vaccination coverage amongst study mothers, or (2) recent introduction of a breakthrough strain (either w.r.t. the vaccine or prior infection) and higher vaccination coverage/infection-derived immunity amongst study mothers. Depending on which mechanism was mostly driving the observed patterns in Zambia, i.e.,

      a. long-running, widespread, unreported transmission;<br /> b. transmission started recently, and vaccination was low amongst study mothers;<br /> c. a breakthrough strain is causing the current rise (here we'd still want to know about vaccination status); and<br /> d. something else that I have not considered

      would have implications for how the results are interpreted, and potentially far-reaching implications for the broader pertussis community. All of that is to say, I think the authors were too quick to dismiss these concerns (even if they disagree with my assertions).

      In their reply, the authors largely dismissed concerns about not knowing the mother's vaccination status, stating in their reply that, "our findings strongly suggest ongoing pertussis transmission in this population. Based on this, we expect that mothers in our study who were not vaccinated would likely have some degree of infection-derived immunity."

      However, they also stated that, "Zambia offers an evocative example of pertussis surveillance, where no cases have appeared in official WHO reports since 2009" and "As we noted above (and now address in our Discussion), widespread genomic surveillance and microbiological characterization of pertussis are sorely lacking across Africa."

      I don't disagree with the authors' conclusion that pertussis is clearly spreading in Zambia. I also don't disagree that there's clearly evidence for minimally symptomatic, infectious mothers spreading infections to children. Both of these findings matter for Zambia and for our broader understanding of pertussis. However, I don't see how the authors can so confidently conclude that low vaccination rates, coupled with a recent introduction, high vaccination rates, coupled with a breakthrough strain, or high infection-derived immunity, coupled with a breakthrough strain, couldn't be what's driving the increase. The authors do hedge in places and also state in the discussion that their findings don't line up with expectations related to WP/infection-derived immunity, "This corresponds to a mean return frequency of one infection per 14.8 years, which is much shorter than the presumed duration of immunity from natural infection or the whole-cell pertussis vaccination used in Zambia (70, 71)." But, my read of the paper is that the authors are pushing way to ward for a preferred hypothesis that is not more favored than other alternatives.

      Secondly, I asked about placing this work in the context of other studies on asymptomatic transmission, but realize that I did not list any specific papers. Two worth considering are Warfel et al. 2014 and Althouse and Scarpino 2015. Restating for the editor, the Warfel study found that WP facilitated rapid clearance in a non-human primate experimental infection study (admittedly with small sample sizes and many other caveats). Many took that as evidence that WP would also block transmission (admittedly experiments Warfel did not run). If the mechanism underlying the results in Zambia is that either WP or natural infection does not block transmission (in the absence of a breakthrough strain), that would upend many of the assumptions in pertussis research. While not incompatible with the Warfel et al. results, it would negate most of the importance of their finding that WP blocked transmission. From what I can see, the authors do not even cite Warfel et al. 2014, which is a serious gap regardless of whether the authors agree or disagree with the findings. A quick sidebar, the authors seem to duplicate Craig et al. 2020 10.1093/cid/ciz531, listing it as both citation 9 and 38.

      In Althouse and Scarpino, they found evidence of a rise in asymptomatic/underreported/subclinical transmission following the switch from WP to AP. While not as directly relevant to the current study as the Warfel paper (so I leave it to the authors to decide whether citing this paper is important), Althouse and Scarpino discuss asymptomatic transmission at length and also assume that WP conferred strong protection against transmission, so their results (along with dozens and dozens of other studies assuming similar WP/infection-induced immunity protection and durability) would also need to be reinterpreted in the context of this study. The authors should engage with the implication of their results in the context of past modeling studies and what we think we know about vaccine-/infection-derived immunity.

      Going back to my earlier points, unvaccinated mothers and the recent introduction of pertussis, or vaccinated/infection-induced immune mothers with a breakthrough strain, would both explain the current results and be compatible with Warfel et al., Althouse and Scarpino, and a sizable number of other studies. Instead, if transmission from WP- or naturally infected mothers is common (in the absence of a breakthrough strain), that would really change the landscape of pertussis epidemiology. The authors have not convinced me that they can make this conclusion. Hence, why I think it's important that the authors engage more actively with those hypotheses and with relevant debates in the pertussis literature on asymptomatic transmission. I think it's appropriate for the authors to present their preferred hypothesis, but, absent other data, they should also present plausible alternatives that are consistent with past publications.

      References:

      Althouse, B. M., & Scarpino, S. V. (2015). Asymptomatic transmission and the resurgence of Bordetella pertussis. BMC medicine, 13, 1-12.

      Warfel, J. M., Zimmerman, L. I., & Merkel, T. J. (2014). Acellular pertussis vaccines protect against disease but fail to prevent infection and transmission in a nonhuman primate model. Proceedings of the National Academy of Sciences, 111(2), 787-792.

    4. Author response:

      The following is the authors’ response to the current reviews.

      We again thank the editor and reviewers for their detailed attention to our work. In our previous revisions we endeavored to address the principal concerns raised by reviewers that we were capable of addressing. We recognize that asymptomatic pertussis transmission represents a particularly thorny area of epidemiology and public health, where multiple (and sometimes overlapping) mechanisms have been proffered even as empirical evidence remains thin, particularly in low-resource settings such as sub-Saharan Africa. A key finding of our work is that prospective surveillance in such a low-resource setting revealed abundant evidence of otherwise unobserved asymptomatic incidence. Moreover, as we note in our prior revisions, this finding is supported by recent work in South Africa and elsewhere (Kayina et al., 2015; Moosa et al., 2019, 2025). As such, we believe that further prospective surveillance in similar settings would be highly informative, a point that we have sought to emphasize in our present manuscript (and associated commentary).

      While we broadly agree with many of the concerns raised by the reviewers, we believe that we are unable to significantly strengthen the present work through further revisions. We do, however, wish to respond to several points raised in these reviews. Of note, a reviewer raises the possibility of a "breakthrough strain" without reference to existing literature. We agree that we cannot test this hypothesis, and though it is not incompatible with our own findings, it is, however, not consistent with recent molecular surveillance in South Africa (Moosa et al., 2023). The reviewer also raises the potential of low adult vaccination coupled with recent reintroduction. This hypothesis relies on our investigation looking "at the right place and the right time", and further does not explain how immunologically naive adults would have escaped morbidity. We have adopted what we believe is a more parsimonious interpretation of our results (i.e., that asymptomatic infection represents evidence of previous immune exposure), though we agree that a more thorough exploration of this particular issue is warranted, particularly in light of our persistently colonized mothers. In addition, we have noted similar studies in sub-Saharan Africa that also found widespread evidence of asymptomatic pertussis, which we believe is inconsistent with a “right time, right place” interpretation.

      A reviewer also pointed to the work of Warfel et al. (2014) and Althouse and Scarpino (2015). We are familiar with both of these studies and agree with their broad relevance to the field (our apologies for omitting Warfel et al.). The reviewer states that, "If the mechanism underlying the results in Zambia is that either WP or natural infection does not block transmission (in the absence of a breakthrough strain), that would upend many of the assumptions in pertussis research." Critically, we believe that our prospective field study of human patients in a real-world public health system complements previous research, including animal trials and simulation studies. Simply put, given that our study was unable to establish the prior vaccination or exposure status of participants, we do not claim to have shown evidence for transmission despite wP vaccination or prior infection.

      Regarding Althouse & Scarpino (2015), we believe that, for the majority of readers, the most compelling analysis in their paper was the examination of genome sequences that pointed to substantial asymptomatic transmission in the US. This conclusion emerged from their population model, which required that “births” (representing transmission events) exceeded “deaths” (representing recovery of infectious individuals) in order to be consistent with the sequence data. Unfortunately, this paper does not provide a detailed explanation of their methods and data sources, nor is this work directly reproducible through, for example, an open-access code/data repository. We have explored the availability of US genome sequences over the time period of their study and were able to find only 36 sequences: 2 from the pre-vaccine era, 8 from the wP vaccine era, and 26 from the aP vaccine era. Given this notable imbalance in the number of sequences (and thus sequence diversity) that was biased in favour of the most recent time period, is it then surprising that the “birth rate” in their model had to exceed the “death rate” in order to match the genetic diversity in the data? Based on a careful inspection of this work, we do not consider its conclusions to represent a gold standard against which all subsequent studies should be judged. We also note that genomic surveillance and analysis of pertussis remains sparse relative to other fields, though recent works have added dramatically to the corpus of available sequences (Bridel et al., 2022).

      Finally, we note that our previous revisions addressed several concerns raised in the present reviews. For example, we previously sought to address reviewers' about our presentation of the strength of our evidence. In this regard, we broadly agree with the reviewers, and we now state that our results "suggest that pertussis transmission occurs between minimally symptomatic mothers and their newborn infants." We believe this largely addresses a present reviewer's concern that 'the mother-to-infant transmission pathway should be framed as "highly suggestive" rather than "confirmed"'. We also note that our results examine three different threshold Ct values (survival analysis, Fig 4), a point that we believe partially addresses a reviewer's suggestion to "including a sensitivity analysis using a stricter cut-off" and concerns about "the decision to use a Ct<45 threshold, as this is higher than standard clinical cut-offs". Indeed, we discuss the issue of clinical cut-offs (and their appropriateness) at some length in the section, "Test reliability, disease surveillance, and public health where we state, "We recognize that such weak and potentially ambiguous signals may not be appropriate for clinical diagnosis. However, our results demonstrate that they nonetheless contain valuable information about pathogen presence and infection intensity that can (and should) be leveraged for disease surveillance." We have also included in the present work a detailed discussion of qPCR sensitivity and efficiency that we believe should interest others working in pertussis surveillance.

      We do not view our own research as the "last word" in this rather controversial subject. In this spirit, we have attempted to present our work transparently, state our claims carefully, and underscore future activities that we believe would benefit the pertussis research community going forward. For example, we agree that further attention to shared exposure and functional immune data among low-resource communities could provide valuable insights into epidemiology and ecology of pertussis. However, we also believe the trade-offs of including one set of activities over another should be clearly acknowledged by researchers, clinicians, and public health officials. To simply state that we must measure more fails to account for the very real resource constraints that we all face.

      Althouse, B. M., & Scarpino, S. V. (2015). Asymptomatic transmission and the resurgence of Bordetella pertussis. BMC Medicine, 1–12. https://doi.org/10.1186/s12916-015-0382-8

      Bridel, S., Bouchez, V., Brancotte, B., Hauck, S., Armatys, N., Landier, A., Mühle, E., Guillot, S., Toubiana, J., Maiden, M. C. J., Jolley, K. A., & Brisse, S. (2022). A comprehensive resource for Bordetella genomic epidemiology and biodiversity studies. Nature Communications, 13(1), 3807. https://doi.org/10.1038/s41467-022-31517-8

      Kayina, V., Kyobe, S., Katabazi, F. A., Kigozi, E., Okee, M., Odongkara, B., Babikako, H. M., Whalen, C. C., Joloba, M. L., Musoke, P. M., & others. (2015). Pertussis prevalence and its determinants among children with persistent cough in urban Uganda. PLoS One, 10(4), e0123240.

      Moosa, F., du Plessis, M., Weigand, M. R., Peng, Y., Mogale, D., de Gouveia, L., Nunes, M. C., Madhi, S. A., Zar, H. J., Reubenson, G., & others. (2023). Genomic characterization of Bordetella pertussis in South Africa, 2015–2019. Microbial Genomics, 9(12), 001162.

      Moosa, F., du Plessis, M., Wolter, N., Carrim, M., Cohen, C., von Mollendorf, C., Walaza, S., Tempia, S., Dawood, H., Variava, E., & others. (2019). Challenges and clinical relevance of molecular detection of Bordetella pertussis in South Africa. BMC Infectious Diseases, 19, 1–11.

      Moosa, F., Kleynhans, J., Makhathini, L., du Plessis, M., Tempia, S., McMorrow, M. L., Moyes, J., Buys, A., Maake, L., Smit, S., & others. (2025). Bordetella pertussis infection and antibody dynamics in household cohorts in two South African communities, 2016–2018: Findings from the PHIRST study. Journal of Infection, 106550.

      Warfel, J. M., Zimmerman, L. I., & Merkel, T. J. (2014). Acellular pertussis vaccines protect against disease but fail to prevent infection and transmission in a nonhuman primate model. Proceedings of the National Academy of Sciences, 111(2), 787–792. https://doi.org/10.1073/pnas.1314688110


      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The study investigates the role of asymptomatic pertussis carriage in transmission between mothers and their infants, in particular. The authors used a longitudinal cohort study that involved 1,315 mother-infant dyads in Lusaka, Zambia, and they utilized qPCR-based detection of IS481 to track Bordetella pertussis transmission over time. Insights from the study suggest that minimally symptomatic or asymptomatic mothers may act as a reservoir for B. pertussis transmission in the infants, thus challenging the traditional surveillance methods that focus on symptomatic cases. Additionally, the study also identified a subgroup of persistently colonized individuals where mothers were majorly asymptomatic despite sustained bacterial presence.

      The authors aimed to improve comprehension of pertussis transmission dynamics in high-burden low-resource settings, and they advocated for enhanced molecular surveillance strategies to capture full pertussis infection, including those that might have gone undetected.

      Strengths:

      The strengths are the use of innovative study design, especially the longitudinal approach and routine sampling, rather than symptom-driven testing that minimizes bias in the study. The methodology was also rigorous and transparent by evaluating the IS481 signal strength to classify pertussis detection and conducting retesting to assess qPCR reliability. There were also important epidemiological insights, and the findings challenge the traditional wisdom by suggesting that pertussis transmission may frequently occur outside of symptomatic cases. The findings also showed its relevance to global health and policy by arguing for the incorporation of molecular tools like qPCR for surveillance of pertussis in low-resource settings.

      Weaknesses:

      These include reliability on qPCR-based detection without additional validation measures like confirmatory culture or serology. There are also potential alternate explanations for transmission patterns observed in the study such as shared environmental exposure or household transmission. Additionally, there is limited generalizability as the study was done in a single urban site in Zambia. There is also a lack of functional immune data.

      Reviewer #2 (Public review):

      Summary:

      In this paper, the authors describe the results of a longitudinal study of pertussis infection in mother/infant dyads in Lusaka, Zambia. Unlike many past studies, the authors assessed the infection status of individuals independently of whether they were symptomatic for a respiratory infection. As a result, this work represents one of the first studies specifically designed to assess asymptomatic transmission of pertussis. Using qPCR, the authors find strong evidence for the role of asymptomatic transmission from mothers to infants and also evidence for long-term bacterial carriage. This work represents an important contribution to our understanding of the global burden of pertussis. Also, it highlights the still under-appreciated role of asymptomatic transmission across many infectious diseases (including vaccine-preventable ones).

      Strengths:

      Unlike many past studies, the authors assessed the infection status of individuals independently of whether they were symptomatic for a respiratory infection. As a result, this work represents one of the first studies specifically designed to assess asymptomatic transmission of pertussis. Using qPCR, the authors find strong evidence for the role of asymptomatic transmission from mothers to infants and also evidence for long-term bacterial carriage.

      Weaknesses:

      While I am quite enthusiastic about the work, I am concerned that a number of likely relevant confounders were not discussed and that the broader implications of their findings were not well grounded in the existing literature. For example, I could not find information on the vaccination status of the mothers in the study. Given the conclusions about asymptomatic transmission and the durability of immunity, it is important to know the vaccination status of the mothers. Moreover, did the authors have other metadata on the mother/infant dyads, e.g., household size, vaccination status of household members, etc.? Given the potential implications of more widespread asymptomatic transmission associated with pertussis infection, I believe the authors should better couch their results in the context of the broader debate around asymptomatic transmission.

      We appreciate the reviewers' detailed feedback. We provide an overview of our responses here and we address specific recommendations below. In light of reviewers’ comments, we have revised our manuscript in order to improve the clarity of our presentation and to better situate our results within the context of the existing literature. Unfortunately, as the field study has been concluded, many of the reviewers’ recommendations are not possible. These include additional testing (i.e., culture or serology) or sequencing. We have updated the manuscript to more clearly indicate our knowledge regarding maternal vaccine status and immunological immunity of study participants. We have also provided a more comprehensive overview of existing pertussis studies, including genomic surveillance and details regarding sub-Saharan Africa and Zambia in particular. Finally, we have revised the formatting of Figure 4 (survival analysis) to more clearly highlight differences between mothers and infants and to better align with the text, and note that the underlying results are unchanged.

      A particular concern raised in the reviews that we wish to address is the recommendation of culture- or serology-based tests as "confirmatory". We have revised the manuscript in light of this feedback to better reflect our own position on this matter. We believe these recommendations do not adequately account for important trade-offs between testing sensitivity and specificity that are widely recognized in both clinical practice and epidemiology (Enøe et al., 2000; Florkowski, 2008; Swift et al., 2020). When the results of different testing methodology disagree, rarely is one method, a priori, correct. Rather, the disagreement may point to specific test limitations or important biological questions about the study system.

      In the case of pertussis detection, cell culture is recognized for its very low sensitivity, while serological detection is complicated by debate around appropriate threshold levels and time horizons for seroconversion and subsequent decay (Lee et al., 2018; van der Zee et al., 2015). Furthermore, while anti-PT antibodies are a common target of serological detection, these are not reliable correlates of protection (Mills, 2001; Wilk et al., 2019), nor are they reliably generated in response to colonization (de Cellès & Rohani, 2024; Graaf et al., 2020). Overall, the detailed relationship between exposure, carriage, transmissible infection, and the dynamics of anti-PT serology remains poorly characterized (Craig et al., 2020; de Cellès et al., 2025).

      While we agree that these are important questions in epidemiology and public health, we nonetheless wish to highlight that there is no “free lunch": each additional test and protocol comes with additional cost and complexity that should be evaluated based on the specific goals of the intended surveillance. In our case, the repeated sampling of longitudinal surveillance serves as a low-cost "confirmatory" testing regime. We disagree that cell culture would have added value to the present study and would not recommend its addition to future studies (primarily due to low sensitivity). While we agree that before-and-after serology of mothers would have added important context to the present study, we nonetheless expect that significant ambiguity would have surrounded any such results (e.g., Moosa et al. (2025)).

      One area that we strongly agree warrants further attention is the household dynamics in pertussis transmission, particularly in low-resource settings where crowding is common. In our study we were not able to rule out environmental and/or shared transmission events, though our survival analysis did demonstrate a greater impact of mothers on infants than vice versa, results which suggest a causal mechanistic role. In previous studies we detailed the demographics of household size, number of children, and mothers' age (Gill et al., 2021; Gunning et al., 2020), though we have not conducted formal analyses of these important covariates here. We also note that the code and data are freely available, allowing for others to build on our work.

      We believe that future prospective studies are an invaluable tool for directly tracking pertussis disease transmission, including both community and household studies. We have argued here for the value of qPCR-based community surveillance, which could integrate into existing public health activities. Regarding household studies, we note that a key challenge in implementing these studies is selecting an appropriate sampling interval and duration to best capture epidemiological linkages. Our results suggest that qPCR-based real-time population-level surveillance could be used to initiate such a prospective household study during a pertussis outbreak so that a higher sampling frequency (e.g., weekly swabs) could be gainfully employed over a shorter time period.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) Enhance Validation of qPCR Findings:

      We address these comments above at greater length. We also note that the term “false positive” is rather ambiguous here, as we lack a clear distinction between carriage and transmissible infection. We note that our manuscript includes a considerable discussion of qPCR validation, including negative controls and sample retesting. We agree that test sensitivity and specificity remains an important question, and have endeavored to clearly indicate these concerns throughout the manuscript.

      (2) Clarify Transmission Dynamics:

      While we agree that this is an important question, we lack the relevant sequence data to test it. Notably, we suspect that any such phylodynamic linkage would require genomic sequencing due to the relatively low genetic diversity observed in pertussis (see, e.g., population-level estimates of time to most common ancestor in Lefranc et al., (2022). However, sequencing pertussis genomes remains resource-intensive, and we expect that the deployment of such sequencing at scale would be cost-prohibitive in low-resource settings.

      We have revised our manuscript to underscore uncertainties around shared/household exposure. We also direct the reviewer’s attention to our survival analysis, where a notable asymmetry exists between mothers and infants. Here, mothers’ prior qPCR signals exhibit a larger impact on their infants than infants on their mothers (Fig 4). If shared exposure was the principal cause of the observed increase in hazard in ego from alter, then we would expect no such asymmetry between mothers and infants.

      (3) Expand Discussion on Public Health Implications:

      As noted above, we have revised our introduction and expanded our discussion to better account for the existing literature. And, while we are hesitant to put forward specific recommendations based on our study, we feel confident in stating that active pertussis surveillance in low-resource settings is A) almost entirely absent B) possible to achieve, and C) necessary to resolve long-standing questions about pertussis epidemiology at regional and national levels. We have endeavored to clarify these points, particularly within the discussion.

      (4) Address the Role of Immunity More Directly:

      While we lack such immune data, we have pointed to recent work from the notable PHIRST study in South Africa, as well as highlighted ambiguities surrounding these data.

      Reviewer #1 (Recommendations for the authors):

      (1) Do we know the vaccination status of the mothers in the study? … If these data are not available, I think that the paper must be re-framed to acknowledge that all the conclusions are statistical in nature, based on publicly available vaccine coverage data from Zambia.

      We do not have information on the immunization status of mothers, though we cite national rates for Zambia across the relevant time period. We have clarified this point in the revised manuscript. We have endeavored to clearly acknowledge that many of our conclusions are statistical in nature and to clearly quantify the strength of evidence.

      We strongly disagree that anomalously low vaccination rates amongst mothers (i.e., relative to national averages) would materially alter the interpretation of our findings.

      Overall, our findings strongly suggest ongoing pertussis transmission in this population. Based on this, we expect that mothers in our study who were not vaccinated would likely have some degree of infection-derived immunity. Indeed, some have argued that the preponderance of mild/asymptomatic infections in mothers is, of itself, evidence of prior immunological exposure (Fine & Clarkson, 1982).

      (2) Do we know anything about rates of pertussis in Zambia, especially in the study site?

      We address this important question in the discussion. In particular, we state that: “As a populous, middle-income and primarily urban country, Zambia offers an evocative example of pertussis surveillance, where no cases have appeared in official WHO reports since 2009”.

      (3) I couldn't find information in the paper related to the severity of infection in the infants. It's mentioned in the section describing results in Figure 5, but I only saw analyses with symptoms (as opposed to severe symptoms). Do you have outcome data from infants testing positive?

      This question was addressed in more detail in our previous work (Gill et al., 2021), which we briefly summarize in the Introduction. We also show the frequency of severe symptoms in Fig 5 (bottom panel), and detail mild versus serious symptoms in our subgroup analysis (Fig 7D).

      (4) Do we know anything about vaccine-resistant strains of pertussis in Zambia?

      We are not aware of any such work. As we noted above (and now address in our Discussion), widespread genomic surveillance and microbiological characterization of pertussis are sorely lacking across Africa.

      (5) While I believe sequencing is beyond the scope of the current study, the authors should comment on the potential utility of sequencing elements of the pertussis genome and use that to demonstrate causality and direction of transmission more strongly.

      We believe that existing literature has addressed the potential of sequence data and phylodynamics to infer transmission, particularly for pathogens with high mutation rates such as RNA viruses. To date, research into the phylodynamics of pertussis has focused exclusively on population-level dynamics (Lefrancq et al., 2022), where estimates of time to most recent common ancestor (TMRCA) are long, indicating low genomic variability at the scale of countries and years. To our knowledge, no work on pertussis has directly inferred transmission chains from sequence data. Given the existing evidence, we expect that any such work would require genome-level sequencing, which would likely be cost-prohibitive in low-resource settings.

      (6) … However, it would be helpful to understand more about how your results fit into the broader story around pertussis resurgence. … if the infant cases were all mild, they might never have been captured in surveillance data sets.

      We believe that a key result of our study is the remarkable mismatch between country-level symptoms-based surveillance and prospective surveillance, which demonstrates that such mild cases have almost certainly not been captured. These findings are mirrored by recent work in South Africa (now addressed in our Discussion, see Moosa et al. (2025)). We believe that prospective surveillance, particularly in under-surveilled regions, is critical to understanding pertussis transmission writ large, which we have attempted to communicate throughout our discussion.

      (7) Relatedly, if there are still high rates of asymptomatic mother-to-infant transmission with whole cell vaccination, then why is there an observed drop in infant pertussis following whole vaccination in most countries?

      In previous work, we demonstrated that some infants in this cohort exhibited asymptomatic infection (Gill et al., 2021). We note that a drop in pertussis incidence amongst infants after the roll-out of the whole-cell vaccine is not contradictory with our findings. We want to clarify that our results, and evidence that mother-to-infant transmission can occur, does not imply that the whole-cell vaccine fails to protect against transmission.

      We have previously used epidemiological evidence to infer the population-level impacts following the roll-out of whole-cell pertussis infant immunization. For example, we observed an increase in the inter-epidemic period that, together with the drop in infant cases, are consistent with a reduction in transmissible infections (Broutin et al., 2010; Rohani et al., 2000).

      References

      Broutin, H., Viboud, C., Grenfell, B. T., Miller, M. A., & Rohani, P. (2010). Impact of vaccination and birth rate on the epidemiology of pertussis: A comparative study in 64 countries. Proceedings of the Royal Society B: Biological Sciences, 277(1698), 3239–3245.  https://doi.org/10.1098/rspb.2010.0994  

      Craig, R., Kunkel, E., Crowcroft, N. S., Fitzpatrick, M. C., Melker, H. de, Althouse, B. M., Merkel, T., Scarpino, S. V., Koelle, K., Friedman, L., Arnold, C., & Bolotin, S. (2020). Asymptomatic Infection and Transmission of Pertussis in Households: A Systematic Review. Clinical Infectious Diseases, 70(1), 152–161. https://doi.org/10.1093/cid/ciz531

      de Cellès, M. D., & Rohani, P. (2024). Pertussis vaccines, epidemiology and evolution. Nature Reviews Microbiology, 1–14. https://doi.org/10.1038/s41579-024-01064-8

      de Cellès, M. D., Wong, A., Dalby, T., & Rohani, P. (2025). Natural immune boosting biases pertussis infection estimates in seroprevalence studies. Nature Communications, 16(1), 8883. 

      Enøe, C., Georgiadis, M. P., & Johnson, W. O. (2000). Estimation of sensitivity and specificity of diagnostic tests and disease prevalence when the true disease state is unknown.  Preventive Veterinary Medicine, 45(1–2), 61–81.

      Fine, P. E. M., & Clarkson, JacquelineA. (1982). The recurrence of whooping cough: Possible implications for assessment of vaccine efficacy. The Lancet, 319(8273), 666–669.  https://doi.org/10.1016/S0140-6736(82)92214-0 

      Florkowski, C. M. (2008). Sensitivity, specificity, receiver-operating characteristic (ROC) curves and likelihood ratios: Communicating the performance of diagnostic tests. The Clinical Biochemist Reviews, 29(Suppl 1), S83.

      Gill, C. J., Gunning, C. E., MacLeod, W. B., Mwananyanda, L., Thea, D. M., Pieciak, R. C., Kwenda, G., Mupila, Z., & Rohani, P. (2021). Asymptomatic Bordetella pertussis infections in a longitudinal cohort of young African infants and their mothers. eLife, 10, e65663. https://doi.org/10.7554/elife.65663

      Graaf, H. de, Ibrahim, M., Hill, A. R., Gbesemete, D., Vaughan, A. T., Gorringe, A., Preston, A.,  Buisman, A. M., Faust, S. N., Kester, K. E., Berbers, G. A. M., Diavatopoulos, D. A., & Read, R. C. (2020). Controlled Human Infection With Bordetella pertussis Induces Asymptomatic, Immunizing Colonization. Clinical Infectious Diseases: An Official Publication of the Infectious Diseases Society of America, 71(2), 403–411.  https://doi.org/10.1093/cid/ciz840

      Gunning, C. E., Mwananyanda, L., MacLeod, W. B., Mwale, M., Thea, D. M., Pieciak, R. C., Rohani, P., & Gill, C. J. (2020). Implementation and adherence of routine pertussis vaccination (DTP) in a low-resource urban birth cohort. BMJ Open, 10(12), e041198.

      Lee, A. D., Cassiday, P. K., Pawloski, L. C., Tatti, K. M., Martin, M. D., Briere, E. C., Tondella, M. L., Martin, S. W., & Group, C. V. S. (2018). Clinical evaluation and validation of laboratory methods for the diagnosis of Bordetella pertussis infection: Culture, polymerase chain reaction (PCR) and anti-pertussis toxin IgG serology (IgG-PT). PLoS One, 13(4), e0195979.

      Lefrancq, N., Bouchez, V., Fernandes, N., Barkoff, A.-M., Bosch, T., Dalby, T., Åkerlund, T.,  Darenberg, J., Fabianova, K., Vestrheim, D. F., Fry, N. K., González-López, J. J.,  Gullsby, K., Habington, A., He, Q., Litt, D., Martini, H., Piérard, D., Stefanelli, P., … Brisse, S. (2022). Global spatial dynamics and vaccine-induced fitness changes of Bordetella pertussis. Science Translational Medicine, 14(642), eabn3253.  https://doi.org/10.1126/scitranslmed.abn3253 

      Mills, K. H. G. (2001). Immunity to Bordetella pertussis. Microbes and Infection, 3(8), 655–677. https://doi.org/10.1016/s1286-4579(01)01421-6 

      Moosa, F., Kleynhans, J., Makhathini, L., du Plessis, M., Tempia, S., McMorrow, M. L., Moyes, J., Buys, A., Maake, L., Smit, S., & others. (2025). Bordetella pertussis infection and antibody dynamics in household cohorts in two South African communities, 2016–2018:  Findings from the PHIRST study. Journal of Infection, 106550. 

      Rohani, P., Earn, D. J., & Grenfell, B. T. (2000). Impact of immunisation on pertussis transmission in England and Wales. The Lancet, 355(9200), 285–286. https://doi.org/10.1016/S0140-6736(99)04482-7 

      Swift, A., Heale, R., & Twycross, A. (2020). What are sensitivity and specificity?  Evidence-Based Nursing, 23(1), 2–4.

      van der Zee, A., Schellekens, J. F., & Mooi, F. R. (2015). Laboratory diagnosis of pertussis.  Clinical Microbiology Reviews, 28(4), 1005–1026.

      Wilk, M. M., Allen, A. C., Misiak, A., Borkner, L., & Mills, K. H. G. (2019). The immunology of Bordetella pertussis infection and vaccination. In Pertussis: Epidemiology, Immunology & Evolution. Oxford University Press.

    1. eLife Assessment

      This study offers valuable insights into the genetic and evolutionary basis of the starvation response by confirming the hypothesis about the mito-nuclear etiology of this trait. The level of evidence is currently incomplete but could be improved if controls were designed better, particularly by including analysis of variation in the initial, common population prior to all treatments. Overall, this work would be interesting to a broad audience, beyond the Drosophila community, if the analysis of the candidate genes against Human ortholog loci were to be conducted with more careful controls.

    2. Reviewer #1 (Public review):

      Summary:

      This manuscript presents a genome-wide investigation of the genetic architecture underlying adaptation to prolonged starvation in Drosophila melanogaster, using an E&R experimental design maintained across 60 generations. Four starvation-selected (SS) and four matched control (C) populations were whole-genome resequenced, and two complementary analytical frameworks, selective sweep inference combined with low-heterozygosity mapping, and a diffusion-based drift-filtering approach, were applied to identify genomic regions under selection. As a result, the authors report (1) 62 high-confidence sweep-low-heterozygosity regions encompassing 255 genes, and (2) 3,578 SNPs with allele-frequency shifts exceeding neutral drift expectations shared across all four SS replicates, mapping to 578 genes. Mitochondrial pathways are identified as prominent targets, with a 13.9-fold enrichment of nuclear-encoded mitochondrial genes among candidates and differentiation at the mitochondrial origin of replication. Finally, the authors demonstrate that human orthologs of starvation-responsive fly genes are enriched for highly differentiated variants in four human populations from the 1000 Genomes Project.

      Strengths:

      (1) The experimental design with four evolution replicates provides proper control for false discovery.

      (2) The phenotypic characterisation is thorough. The approximately 3-fold increase in starvation survival and 1.5-fold increase in TAG content provide a clear physiological basis for interpreting the genomic findings, and the observation of increased adult longevity adds a meaningful life-history dimension to the results.

      (3) The mito-nuclear analysis is one of the more novel contributions of this paper. The implicated picture of coordinated mito-nuclear remodelling under sustained nutrient deprivation is compelling.

      (4) The comparative analysis connecting fly selection candidates to human population differentiation is ambitious and adds evolutionary breadth to the study.

      Weaknesses:

      (1) Ne estimation is derived from controls only, not from selected populations

      The entire drift-filtering framework rests on estimates of effective population size obtained from allele-frequency variance among the four control replicates (Ne = 530 for autosomes, Ne = 461 for the X chromosome). This is justified by assuming that divergence among control populations reflects neutral drift alone, a reasonable assumption for C populations maintained on standard food.

      However, the starvation-selected populations experienced 75-80% mortality per generation as an explicit design feature of the selection regime. This severe, recurrent demographic bottleneck would substantially reduce the effective population size within SS lines relative to controls. The authors do not acknowledge this discrepancy, nor do they attempt to estimate Ne within SS replicates or assess the sensitivity of their drift thresholds to plausible reductions in Ne. If Ne in SS populations is appreciably lower than in controls, the drift thresholds derived from the control-based Ne will underestimate the amount of neutral drift occurring in SS lines. Consequently, some allele-frequency shifts that are driven by the repeated bottleneck could be misclassified as candidate loci, inflating the apparent number of selection targets. This is the most consequential methodological concern in the paper. The authors should either estimate Ne separately for SS populations, implement a sensitivity analysis varying Ne over a biologically plausible range, or, at a minimum, provide a thorough discussion of how downward bias in SS Ne would affect their results and conclusions.

      (2) Lack of consideration about binomial sampling noise due to the pool size in the modeling

      With only 100 individuals pooled per population, binomial sampling from the pool contributes a non-trivial additional source of variance to allele-frequency estimates, on top of genetic drift and sequencing error. This is a well-documented issue in Pool-seq data. Critically, the Kimura diffusion framework used for drift modeling does not appear to explicitly incorporate this binomial sampling noise component, an omission that could affect the calibration of drift thresholds, particularly for low-frequency alleles. The authors should discuss whether and how pool-size-induced sampling variance is accounted for in their drift model.

      (3) No benchmarking against established Pool-seq analysis tools

      The authors use Pool-HMM for sweep detection and a custom diffusion-based drift framework for allele-frequency analysis, with PoPoolation (v1) used only for Tajima's D calculations. However, the study does not benchmark its candidate SNP sets or sweep regions against well-established Pool-seq analysis frameworks such as PoPoolation2, which provides CMH tests and FST estimation specifically designed for replicated Pool-seq E&R data, or R/poolSeq, which implements drift-aware testing purpose-built for this experimental design. The authors should either benchmark their approach against at least one established alternative or provide explicit justification for why their custom framework is preferable and how it compares in sensitivity and specificity.

      (4) Absence of negative controls in the human PBS comparative analysis

      A critical missing element in this comparative analysis is a negative control: the authors do not test whether equivalent enrichment is observed in populations with no particular history of famine or nutritional stress, such as European or East Asian populations from the 1000 Genomes Project. The inclusion of at least one negative-control population triplet is necessary to support the cross-species interpretation as stated.

    3. Reviewer #2 (Public review):

      Summary:

      The authors use an Evolve-and-Resequence approach in Drosophila to study the genomic basis of adaptation to long-term starvation. Replicated selection lines and control populations are sequenced and analyzed to identify signals of selection, which are then related to starvation-related phenotypes. The general experimental design is appropriate, and the combination of genomic and phenotypic data is a clear strength of the study.

      Strengths:

      The strongest aspect of the work is the experimental evolution framework combined with population genomic inference across replicate populations. The observed parallelism across replicates supports the robustness of at least a subset of the detected selection signals. However, several key methodological details are either unclear or insufficiently justified. In particular, both the maintenance of control populations and demographic assumptions are not fully described, and the treatment of structural variation (e.g., segregating inversions) is not sufficiently addressed. The phenotypic analyses are broadly appropriate and replicated but would benefit from access to raw data.

      Weaknesses:

      The human ortholog enrichment analysis is an interesting component of the study, but it should be interpreted more cautiously. As currently presented, it is based on correlational signals of differentiation and is therefore sensitive to potential confounding factors. While the analysis may point to intriguing patterns consistent with conserved genetic architecture, the evidence is not sufficient to support strong claims of conserved starvation/malnutrition-related polygenic adaptation in humans. Framing this component more explicitly as exploratory would strengthen the manuscript. In its current form, this analysis is somewhat less conclusive than the experimental evolution results in flies.

      Overall, the study provides a useful dataset and a reasonably solid analysis of starvation adaptation in experimental Drosophila populations, but several methodological clarifications and a more balanced framing of the cross-species comparisons would strengthen the manuscript.

    4. Reviewer #3 (Public review):

      Summary:

      This study tries to identify the genetic signatures of adaptation to starvation conditions. For this, outbred populations of Drosophila melanogaster were selected for starvation resistance by using the 20% surviving adults after starvation to start the next generation. This was done for 60 generations while parallel populations were kept under control conditions. At the end of the experiment, starvation-selected flies showed increased survival, longevity, and TGA storage. DNA poolseq data from control and starvation populations were compared to identify genomic regions with low heterozygosity and signatures of selective sweeps, and SNPs with differences in allele frequency. The candidate regions point to mitochondrial and metabolic pathways as the targets of selection for starvation resistance.

      The authors replicate the experimental design, selection approach, data collection, and analyses from Hardy et al 2018 (https://doi.org/10.1093/molbev/msx254), which also investigated adaptation to starvation conditions but used a different Drosophila melanogaster population. In this sense, the current study recapitulates most of the findings from Hardy et al. (2018). The analyses of the mitochondrial results, including the overlap with human data, are the novelty of this paper. However, those analyses are not very well justified. The fact that this study is almost identical to Hardy et al is not clearly stated nor discussed in the manuscript.

      Strengths:

      The authors made use of an experimental evolution approach to identify the genetic basis underlying adaptation. This is a powerful approach that has proven very successful in the past. They used a good number of replicates (four per condition), an appropriate depth of sequencing, and quantified higher-order phenotypes to validate the claim that the populations had evolved increased starvation resistance.

      Weaknesses :

      Although the findings of this study seem credible based on the known biology of starvation resistance, there are several aspects of the experimental design that weaken my confidence in the results. The points below should be clarified, and the limitations of the experimental design and analyses need to be included in the discussion.

      (1) Pooled genomic data were collected for the four replicates at the end of 60 generations of selection, and four replicates were kept under control conditions. No data were collected at the beginning of the experiment, which is the current standard in Evolve and Resequence experiments. To infer the genomic regions underlying adaptation to starvation, evolved control and starved cages are compared. Although this will identify regions that are possibly truly caused by adaptation to starvation stress, the available data doesn't allow to determine, for example: a) whether the differences between control and starvation regimes are due to changes in control cages relative to the starting population, combined with no changes in starvation cages relative to the starting population; b) whether the differences across replicates are due to different genomic composition at the start of the experiment that could have been amplified by drift.

      (2) Selection was applied by starving flies until ~80% of the population died. The 20% surviving flies were used to seed the next generation. The control populations, on the other hand, were propagated using the whole population. Given that only the starvation populations were subject to such a strong bottleneck, it is not possible to disentangle whether the genomic signatures at the end of the experiment are due to this, and not necessarily to starvation resistance. For example, the low heterozygosity blocks and the very great changes in allele frequency could be a natural result of such a bottleneck. A proper comparison would have been to select a random 20% of the control individuals to seed every generation.

      (3) The analyses that involve human populations are poorly justified, and the enrichment tests are not clearly explained. There is no evidence of signatures of selection for starvation resistance in human datasets (as mentioned in the text, line 112), and yet the authors claim that their analyses that identify branch-specific alleles for a set of four human populations serve as a dataset for it. I don't think the results of this analysis and further overlap with candidate genes identified in the Drosophila experiment support the conclusion that polygenic adaptation of metabolic pathways is conserved across species (line 303).

      (4) The conclusion that adaptation to starvation conditions is repeatable is not justified by the data. The overlap across replicates is very low in every metric.

      (5) The methods are poorly described. In most of the sections, there is not enough information to be able to replicate the experiments or the analyses. Several of the analyses presented in the results are not described in the methods. Without this information, it is very difficult to assess whether the analyses were correctly done or whether the results are robust.

    1. eLife Assessment

      This valuable study combines a chromosome-level genome assembly with population resequencing and demographic modelling to reassess the status of the Formosan landlocked salmon, a critically endangered salmonid at the southern edge of the genus. The assembly and the evidence for extensive lineage-specific chromosomal fusions are genuinely well done, and the discovery of an overlooked, genetically distinct population in Hehuan Creek is sound and of clear management relevance. Evidence for the headline claims is incomplete: species rank is asserted but nowhere argued, and gene flow is excluded using a coalescent model that contains no migration parameter. Support for the population-viability conclusions is also not complete, as the demographic mechanisms invoked in the discussion are not borne out by the supplementary tables.

    2. Reviewer #1 (Public review):

      Summary:

      This is an interesting paper on an important topic, the taxonomic and conservation status of some unusual salmonid populations in Taiwan.

      Strengths:

      The first part of the manuscript is quite strong: the authors sequence and build a reference genome and conduct a phylogenomic analysis. They examine chromosome structure and rearrangements, test for loss-of-function mutations, and do a proteomic analysis. As a stand-alone, this could serve as its own manuscript, perhaps for a more specialized journal.

      Weaknesses:

      I find this manuscript rather disjointed. The first part of the manuscript is related to phylogenomics of the taxon in question, compared to other nearby species from Japan. The authors go on to describe chromosomal rearrangements, sex-chromosome location, and proteomics. All of these fit within a paper about taxon-level issues. I do find the proteomic analysis perhaps unnecessary. I'm not sure we learn much of substance through this analysis, which is highly speculative.

      PSMC analysis seems highly questionable for taxa with such strong genetic structure. If historical Ne and past changes in structure are confounded, what does this analysis provide? I recommend deletion of the analysis included in Figure 1e.

      The second portion of the manuscript deals with population structure of three O. formosanus populations, based on RADSeq data. This part reads as a separate manuscript, in my opinion. I think the authors are trying to squeeze too much into one manuscript.

      For the second part on population genomics, not enough detail is provided to evaluate the methods, results, and interpretations. For example, not enough detail is provided about each of the three Taiwan populations, the stocking history, and the demographic data collection. The only information available is a brief paragraph in the introduction. Was the Luoyewei (L) population stocked from a brook derived from this population or from Qijiawan (Q)? Why do three L fish have such different levels of MLH? Are these stocked from somewhere else? Are the rest of the fish from one pool, and maybe one family (this would also explain the extremely low contemporary Ne)? Only 17 fish were examined from L, and apparently from one site in the stream; more detail is needed. Are L, Q, and H currently isolated? What is the stocking history? The authors conclude that the Hehuan (H) population has more genetic variation and is likely the result of an unknown native population that bred with stocked fish (which arise from Q). This story does align with the genetic results, but again, more detail is needed. Are there alternative explanations? A more careful treatment would be helpful.

      The demographic modeling is not convincing. Not enough detail is provided, and the lack of individual identification of fish makes it so the modeling is very general. It is hard to place too much stock in these vital rate estimates. The methods were fishing, snorkeling, and some electrofishing. Scales were used for ageing, and catch curve analysis was employed. Overall, this is an underdeveloped portion of the paper that is important, but not convincing as written.

    3. Reviewer #2 (Public review):

      Summary:

      Lee et al. is a comprehensive conservation genomics study that combines a chromosome-level genome assembly (sex-specific, too), population resequencing, coalescent species delimitation, and simulations to reassess the evolutionary status and conservation outlook of the Formosan landlocked salmon, Oncorhynchus formosanus. The authors showed a distinctive genome structure, replete with chromosome fusions and an unusual placement of the sex-determining gene sdY. Across sampling sites, they observed variable levels of genetic diversity, but in a way that was surprising given previous census numbers and conservation history. In particular, the authors report a previously unrecognised native population in Hehuan Creek, and conclude that Hehuan is more resilient to typhoon disturbance than the long-protected Qijiawan population - motivating stream-specific rather than range-wide conservation.

      Overall, this is a well-written paper that combines a number of elements that are timely and relevant. It uses state-of-the-art techniques to reach its conclusions and is generally performed to a high standard. It describes a critically endangered species that poses its unique conservation challenges. There are a number of things to like, as well as some substantial shortcomings in this paper.

      Strengths:

      The genomic resource is excellent. The assembly is well validated (97.3% anchored to 25 scaffolds, 95.7% BUSCO), and the authors generated a separate male assembly specifically to resolve the sex-determining region, allowing XY-shared and Y-specific contigs to be distinguished on coverage rather than inference. This is truly well done, and at a high standard. The synteny evidence for telomere-to-telomere fusions involving at least 14 ancestral chromosomes, against two in O. m. masou, is convincing.

      The Hehuan Creek result is the paper's most valuable contribution. Elevated heterozygosity, short and infrequent runs of homozygosity, and private alleles absent from the Qijiawan broodstock are difficult to reconcile with a purely reintroduced origin. The contrast with Luoyewei is a clean and useful cautionary case for hatchery supplementation.

      Weaknesses:

      (1) The species-rank claim is featured in the abstract, but it is made with any level of rigour in the paper. "New species" appears once, in the abstract (l. 32). The Results conclude only that O. formosanus is a distinct evolutionarily significant unit (ll. 188-191), which itself can be well-justified, but it's far from a taxonomic rank (see author's own ref 10). No species concept is explicitly named anywhere, and the taxon is referred to across the manuscript as a subspecies (l. 68), a "new species" (l. 32), and an ESU (l. 189) in turn.

      (2) Gene flow is asserted, not tested, and two divergence estimates disagree twentyfold. The abstract reports "no detectable gene flow for ~50,000 years." That figure is a divergence time from BPP under the A00 model, which contains no migration parameter; a model that cannot fit gene flow cannot report its absence. Separately, Figure 1b shows a split at 1.15-5.09 Mya (Figure 1b), while the ddRAD coalescent places the same split at ~50 kya (Figure 1d). The explanation offered (ll. 417-421, "differing temporal sensitivity of genomic markers") is not a mechanism.

      (3) The placement of sdY is unresolved, and the paper's own figures conflict with its text. Figure S10 and Table S6 both make O. formosanus chr13 homologous to O. m. masou chr32, whereas reference 28 - on which the authors rely - places the sdY contig on O. m. masou chr7, whose O. formosanus homologue is chr5 (Table S6). These cannot both be correct, and Figure S10's caption compounds the confusion by attributing chr13 to masou and omitting the chr32 track entirely.

      (4) The population-viability model's stated mechanisms are contradicted by the authors' own supplementary tables. The Discussion attributes Qijiawan's vulnerability to "lower juvenile survival, decreased fecundity, and narrower terminal age class representation" (ll. 522-525). Table S10 gives Qijiawan higher age-0 survival (0.202/0.616 vs 0.184/0.615); Table S11 gives it higher fecundity at every reproductive age (7.68/19.27/7.86 vs 6.10/8.95/4.14); Table S9 gives it a broader terminal age class (5.0% vs 0.8% age-3 in November). The only parameter favouring Hehuan is age-1 survival - 0.087 (95% CI 0.000-0.180) versus 0.131 (0.093-0.187) under typhoon, and 0.054 (0.000-0.167) versus 0.087 (0.047-0.149) at baseline. Both Qijiawan intervals include zero and overlap Hehuan's, yet a reported extinction odds ratio of 4.48 rests on this difference.

      (5) The two streams were not measured equivalently, and every asymmetry favours the conclusion.

    1. eLife Assessment

      This important study combines cryo-EM, biochemical, and cell-based assays to examine how Gβγ interacts with and potentiates PLCβ3. The authors present evidence for multiple Gβγ interaction surfaces and argue that Gβγ primarily enhances PLCβ3 activity after membrane recruitment rather than serving mainly as a membrane-recruitment factor. Following additional experimental support, the evidence in support of their conclusions is convincing.

    2. Reviewer #1 (Public review):

      The manuscript by Fisher et al describes the molecular mechanism underlying how G beta gamma subunits engage with the beta 3 isoform of PLC. The paper used a combination of cryo EM, BRET assays, and biochemical assays of PLC beta activity. A key discovery is that G beta gamma is not sufficient to drive membrane binding by itself and instead promotes G alpha activation. The work is important, but suffers slightly from some ambiguity in the actual interface that is present in their cryo EM model, as crosslinkers could stabilise a transient and non-native complex. This is somewhat abrogated by the careful mutational analysis, which shows that mutation of any of these three sites does somewhat block PLC beta G beta gamma activation. However, there could be some improvement in the presentation of this data, as well as possible mutant selection. Overall, this paper is a nice complement to the Falzone et al paper showing the membrane bound complex of PLCB3 on membranes, with this work building on this work, highlighting the importance this will have in our full understanding of PLC beta activation.

      Major concerns

      My most major concern is the potential that this interface is artefactual based on the crosslinking strategy utilised. Here are thoughts on how this could be better validated, presented in a more convincing way.

      (1) The authors main claim is that there is a degree of plasticity of G beta gamma binding to the PLC beta 3 isoform, with three possible binding sites. The main complication of this is of course the possibility that the crosslinking stabilises a non-native complex, driven by a mutated cysteine.

      Because of this any other additional details about this interface are going to be critical for the scientific audience to judge if this is accurate.

      What would greatly help figure 1, is an evolutionarily conservation analysis of the novel Gbg interface in PLC, to see how well this is conserved, and compare this to the conservation of the previously annotated sites. Conservation of these sites on both the G beta gamma and PLC side would help justify this as a native complex.

      This also will help orient the reader to the identity of the mutated residues assayed in figure 3.

      (2) The g beta gamma orientation is also different than what I have observed in previous g beta gamma effector structures. Is there any precedent for this as an effector interface? A supplemental figure comparing this structure to other g beta gamma interfaces from other enzymes, for example recent tesmer structure with PI3K.

      (3) The mutational analysis in Figure 2D-G seems to give some strange results, and I have some question why certain residues were chosen rather than others. Mutation of the Gbg side will be more complicated as of course that can effect any of the three surfaces. My main question is that from the way fig 2A is oriented that the main salt bridge in their novel interface to me looks like R199-D228, with K183 being in the wrong orientation to E226, and D167 being far from any charged residues. Why did the authors not make the corresponding R199 to D or E mutation?

      (4) To help reader interpretation of Figure 2A, I would recommend a supplemental figure showing the density for interfacial residues, as that also would increase confidence in the interface.

      Comment on revised version.

      After revision the authors have addressed all of my concerns.

    3. Reviewer #2 (Public review):

      In this manuscript, the authors dissect how Gβγ potentiates PLCβ3 signaling in cells. Using engineered crosslinking to stabilize a Gβγ-PLCβ3 complex, single particle cryo-EM, and cell-based functional assays, they identify map multiple putative Gβγ interaction surfaces on PLCβ3, including a previously unrecognized binding mode. Structure-guided mutagenesis supports the functional relevance of these interactions and suggests that Gβγ potentiation is not primarily mediated by PLCβ3 membrane recruitment, but instead enhances PLCβ3 activity after the lipase is already at the membrane.

      Previous reconstitution work on membrane surface (Falzone & MacKinnon, 2023) proposed a recruitment/partitioning-centric model in which Gβγ increases PLCβ3 output largely by elevating its membrane surface concentration, whereas Gαq primarily increases catalytic turnover; under those reconstitution conditions, the two inputs can combine approximately multiplicatively. In receptor-driven cellular signaling, however, PLCβ3 is robustly recruited to the plasma membrane upon Gαq activation, which raises the question of whether Gβγ contributes mainly through additional recruitment or through a post-recruitment mechanism once PLCβ3 is already at the membrane.

      This manuscript helps address that gap by using membrane-anchored PLCβ3 and complementary cellular readouts to separate "getting PLCβ3 to the membrane" from "boosting activity once PLCβ3 is already there." Their results argue that, in cells, membrane recruitment is largely dominated by Gαq·GTP, while Gβγ can further potentiate PIP2 hydrolysis after membrane association, consistent with a modulatory role at the membrane rather than primary recruitment.

      Overall, the work provides a structural and mechanistic framework for Gβγ-PLCβ3 cooperation and helps clarify the basis of Gq pathway amplification.

      Comments on revised version.

      The authors have reasonably addressed my comments.

    4. Reviewer #3 (Public review):

      Summary:

      PLCβ3 is activated by both Gαq and Gβγ subunits. This paper follows previous solution and cryoEM studies of the PLCβ3 / Gβγ complex to delineate the molecular details of activation using cellular BRET assays and cryoEM.

      Strengths:

      The authors find evidence for multiple binding sites on PLCβ3 for Gβγ and suggest that Gβγ is not bone fide activator per se but enhances Gαq activation by positioning the catalytic site towards substrate. The authors also find that this activation is not through recruitment of the enzyme to the membrane by Gβγ released upon G protein activation in accord with other PLCβ enzymes.

      Weaknesses:

      (1) The main issue is that the author's mechanism does not fully explain how Gβγ activation occurs for PLCβ2 in reconstituted systems in the absence of Gαq subunits but will be investigating this in future studies.

    5. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      The manuscript by Fisher et al describes the molecular mechanism underlying how G beta gamma subunits engage with the beta 3 isoform of PLC. The paper used a combination of cryo EM, BRET assays, and biochemical assays of PLC beta activity. A key discovery is that G beta gamma is not sufficient to drive membrane binding by itself, and instead promotes G alpha activation. The work is important, but suffers slightly from some ambiguity in the actual interface that is present in their cryo EM model, as crosslinkers could stabilise a transient and non-native complex. This is somewhat abrogated by the careful mutational analysis, which shows that mutation of any of these three sites does somewhat block PLC beta G beta gamma activation. However, there could be some improvement in the presentation of this data, as well as possible mutant selection. Overall, this paper is a nice complement to the Falzone et al paper, showing the membrane-bound complex of PLCB3 on membranes, with this work building on this work, highlighting the importance this will have in our full understanding of PLC beta activation.

      Thank you for the positive feedback.

      Major concerns:

      My biggest concern is the potential that this interface is artefactual based on the crosslinking strategy utilised. Here are thoughts on how this could be better validated, presented in a more convincing way.

      (1) The authors' main claim is that there is a degree of plasticity of G beta gamma binding to the PLC beta 3 isoform, with three possible binding sites. The main complication of this is, of course, the possibility that the crosslinking stabilises a non-native complex, driven by a mutated cysteine.

      Because of this, any other additional details about this interface are going to be critical for the scientific audience to judge if this is accurate.

      What would greatly help Figure 1 is an evolutionary conservation analysis of the novel Gbg interface in PLC, to see how well this is conserved, and compare this to the conservation of the previously annotated sites. Conservation of these sites on both the G beta gamma and PLC side would help justify this as a native complex.

      This will also help orient the reader to the identity of the mutated residues assayed in Figure 3.

      We agree that crosslinking can capture non-physiologically relevant interfaces. However, because we do not observe any crosslinking between Gβγ and a PLCβ3 variant that retains a cysteine in the X–Y linker or between PLCβ3 and any other cysteine in the Gβγ heterodimer, we believe it is site-specific.

      The question about sequence conservation in the Gβγ–PLCb3 interfaces is interesting and we have included this information in Figures S8 and S9.

      (2) The g beta gamma orientation is also different than what I have observed in previous g beta gamma effector structures. Is there any precedent for this as an effector interface? A supplemental figure comparing this structure to other g beta gamma interfaces from other enzymes, for example recent Tesmer structure with PI3K.

      We agree that the orientation of Gβ in the crosslinked structure is different. We include a comparison of this reconstruction to other published Gβγ–effector complexes as Figure S6.

      (3) The mutational analysis in Figure 2D-G seems to give some strange results, and I have some question why certain residues were chosen rather than others. Mutation of the Gbg side will be more complicated, as of course that can affect any of the three surfaces. My main question is that, from the way Figure 2A is oriented, the main salt bridge in their novel interface to me looks like R199-D228, with K183 being in the wrong orientation to E226, and D167 being far from any charged residues. Why did the authors not make the corresponding R199 to D or E mutation?

      Thank you for pointing this out, and we expanded our analysis to include this residue. The R199A and R199E mutations had no defects in basal or Ga<sub>q</sub>-stimulated activities. However, R199A had 3-fold lower activation by Gβγ, while R199E was not activated in this assay. The R199E mutation also had significantly decreased agonist-dependent BRET with Gβγ and decreased PI(4,5)P2 hydrolysis, while retaining robust recruitment to the plasma membrane by Gα<sub>q</sub>. These data is included in the main text and Figures 2-4.

      (4) To help the reader's interpretation of Figure 2A, I would recommend a supplemental figure showing the density for interfacial residues, as that also would increase confidence in the interface.

      Thank for the suggestion. In revised Figure S3, we show the Gβγ–PLCb3 D892-PH<sub>cys</sub> complexes determined in this study at different contour levels.

      Reviewer #2 (Public review):

      In this manuscript, the authors dissect how Gβγ potentiates PLCβ3 signaling in cells. Using engineered crosslinking to stabilize a Gβγ-PLCβ3 complex, single particle cryo-EM, and cell-based functional assays, they identify and map multiple putative Gβγ interaction surfaces on PLCβ3, including a previously unrecognized binding mode. Structure-guided mutagenesis supports the functional relevance of these interactions and suggests that Gβγ potentiation is not primarily mediated by PLCβ3 membrane recruitment, but instead enhances PLCβ3 activity after the lipase is already at the membrane.

      Previous reconstitution work on the membrane surface (Falzone & MacKinnon, 2023) proposed a recruitment/partitioning-centric model in which Gβγ increases PLCβ3 output largely by elevating its membrane surface concentration, whereas Gαq primarily increases catalytic turnover; under those reconstitution conditions, the two inputs can combine approximately multiplicatively. In receptor-driven cellular signaling, however, PLCβ3 is robustly recruited to the plasma membrane upon Gαq activation, which raises the question of whether Gβγ contributes mainly through additional recruitment or through a post-recruitment mechanism once PLCβ3 is already at the membrane.

      This manuscript helps address that gap by using membrane-anchored PLCβ3 and complementary cellular readouts to separate "getting PLCβ3 to the membrane" from "boosting activity once PLCβ3 is already there." Their results argue that, in cells, membrane recruitment is largely dominated by Gαq·GTP, while Gβγ can further potentiate PIP2 hydrolysis after membrane association, consistent with a modulatory role at the membrane rather than primary recruitment.

      Overall, the work provides a structural and mechanistic framework for Gβγ-PLCβ3 cooperation and helps clarify the basis of Gq pathway amplification. The manuscript is generally strong, but some issues need to be addressed.

      Thank you for the positive comments.

      Major comments:

      (1) BMOE/BM(PEG)2 crosslinking may enforce a non-native docking geometry, potentially compromising the physiological relevance and precision of the Gβγ-PLCβ3 interface as described. Although a >50% 1:1 crosslinked complex is formed and remains active, the solution maps show lower local resolution for Gβγ, consistent with a dynamic, potentially heterogeneous, interface. One interface is captured via a single engineered cysteine pair (PLCβ3 E60C-Gβ C271), which could potentially bias the pose. It would be helpful if the authors could provide additional orthogonal support (e.g., alternative crosslinked sites) and bolster the clarification of its uniqueness and relevance.

      We did attempt to isolate other crosslinked complexes. PLCβ3-D892 self-crosslinked under all reaction conditions, while PLCβ3-D892 XY<sub>Cys</sub>, which retains an endogenous cysteine within the X–Y linker (C516), did not result in any crosslinked product when incubated with Gβγ. Only the PLCβ3-D892 E60C crosslinked to Gβγ. With the exception the C68S mutation at the C-terminus of Gg to eliminate its prenylation site, all endogenous cysteines were retained in both Gβ and Gγ. Indeed, Gβ contains two solvent-exposed cysteines in its canonical effector binding surface (C204 and C271), but we did not observe any crosslinker density involving C204. While we cannot exclude the possibility that crosslinking occurred between PLCβ3-D892 E60C and other residues in Gβγ, we were unable to identify any 2D classes corresponding to these alternative conformations. These observations, together with the high efficiency of crosslinking, are consistent with a stable and persistent interaction.

      (2) In the crosslinked structure, the authors report that GβD228 interacts with PLCβ3 R199 and K183. In Figure 2A, R199 appears closer to Gβ D228 than K183, yet only K183 is functionally tested. Testing R199 (e.g., R199E/R199A) would strengthen the structure-guided validation of this interface.

      We agree, and functional analysis of PLCb3 R199E is included in the revised manuscript (see Figures 2-4).

      (3) The mutagenesis strategy appears inconsistent across figures/assays, which makes it difficult to interpret phenotypes and directly link the functional data to the proposed interfaces. For example, in Figure 2E, we see R185L but R215E, while residue L40 is mutated to Gly in the IP accumulation assays but to Glu/Lys (L40E/K) in the BRET assays (Figures 3B/3D/3F). The authors should (i) clearly justify the rationale for each substitution (conservative vs charge-reversal, interface disruption, etc.) and (ii), where possible, test the same mutants across assays (or provide evidence that alternative substitutions yield consistent conclusions).

      Mutagenesis experiments were initially carried out independently in the Lambert and Lyon Labs. As the study progressed, additional mutants were identified and/or designed based on results from both groups. The residues subject to mutagenesis are overall consistent across the different assays, with differences in the identity of the mutation varying in some cases. The L40G mutation is one such example, where given its modest impact on Gβγ-mediated activation in the IP accumulation assay, more impactful changes were made (L40E and L40K) for the BRET and signaling assays. In the revision, we now state that mutations were designed to maximally disrupt the three observed interfaces, such as by changing the size of the side chain and/or introducing charge reversal mutants.

      Reviewer #3 (Public review):

      Summary:

      PLCβ3 is activated by both Gαq and Gβγ subunits. This paper follows previous solutions and cryoEM studies of PLCβ3 / Gβγ, trying to understand the molecular details of activation using cellular BRET assays and cryoEM.

      Strengths:

      The authors find evidence for multiple binding sites on PLCβ3 for Gβγ and suggest that Gβγ is not bone fide activator per se but enhances Gαq activation by positioning the catalytic site towards substrate, although this is not completely convincing. Although these sites may not naturally be operative, the authors might want to develop the potential role of these sites.

      The authors also find that this activation is not through recruitment of the enzyme to the membrane by Gβγ released upon G protein activation, in accord with other PLCβ enzymes, but not for PLCβ3, and again, the authors might want to develop this point further.

      Thank you for the suggestions. We are investigating whether the other PLCb isoforms contain multiple Gβγ binding sites and the relative importance of preactivation by Ga<sub>q</sub> for a manuscript in preparation.

      Weaknesses:

      (1) I'm confused as to why the authors feel that their mechanism is distinct from the two-state enzyme, the synergistic activation proposed by Ross in 2011, using a primarily thermodynamic argument. As written, the authors appear to be very reliant on structural and BRET studies that do not give the details that would disprove this interpretation. The main issue is that the author's mechanism does not fully explain how Gβγ activation occurs for PLCβ2 in reconstituted systems in the absence of Gαq subunits.

      The reconstitution experiments are under extremely artificial conditions, using nM-µM of purified proteins and liposomes that contain up to 30% PI(4,5)P2. Under these conditions, we think the increased activity is due to interfacial activation promoted by Gβγ binding to the lipase once it is associated with the liposome surface. This would be sufficient to account for the dose-dependent increase in both PLCb2 and PLCb3 activity as a function of Gβγ concentration. Given the higher basal activity of PLCβ2 and its decreased sensitivity to activation by Ga<sub>q</sub>, one possible explanation is that this isoform differs in its autoinhibition and/or structure of its proximal CTD that Ha2’ displacement is not a prerequisite for activation. In addition, Gβγ may also be a direct activator of PLCβ2. Further studies, ideally in cell-based systems, are needed to answer these questions.

      (2) In a recent study, McKinnon presents a model showing that Gαq and Gβγ activate PLCβ3 by two distinct pathways and that activation by Gβγ occurs through membrane recruitment. It is not surprising that the authors find that this is not true since the pelleting method used by McKinnon is subject to error. The authors should directly address the limitations of this previous work and the changes in proteoliposomes with sedimentation that alter partition coefficients. Although the inability of Gβγ to drive membrane binding is in accord with the quantitative studies of Scarlata, showing that the affinity of PLCβ3 to Gβγ is fairly weak as compared to the intrinsic membrane partition coefficient.

      We have added some of the limitations of proteoliposome sedimentation experiments to the discussion.

      (3) It was proposed many years ago that in signaling complexes Gαq - Gβγ may not have to fully dissociate when binding PLCβ, but rather shift their relative orientation when binding to PLCβ to allow activation. Is their model consistent with this? Is it possible that PLCβ3 keeps Gβγ from diffusing to enhance the rate of Gq / Gβγ re-association?

      Our crosslinked complex is compatible with simultaneous binding of a Gα<sub>q</sub>-Gβγ heterotrimer to the PLCb3, without disrupting the observed interface. If Gαq were to interact with the Gβγ molecules bound to the PH or EF hands, the interaction would be mediated by the N-terminal helix of Gα<sub>q</sub>. It is possible Gβγ–PLCβ3 interactions may slow heterotrimer reassociation, but this may be complicated by the intrinsic GAP activity of the lipase.

      (4) The authors find that Gβγ binds multiple sites, and it is clear that the PH domain site is the primary one in accord with previous work. Could these weaker sites be an artifact of the elevated concentrations used in cryoEM and BRET assays?

      While more studies have focused on the PH domain as a Gβγ binding site, our data does confirm the EF hands are also functionally relevant. To our knowledge, the role of the EF hands has not been investigated in this capacity until very recently, and so we hesitate to label them primary or secondary. It is possible the EF hands may be a lower-affinity site for Gβγ and the protein concentrations needed in cryo-EM drive complex formation. However, it is also possible the concentration of free Gβγ adjacent to an activated receptor may be high enough to saturate the PH and EF hand binding sites.

      (5) Although their assays infer differences in binding affinities, it would strengthen the paper if the authors could estimate the association energies of these different binding sites. This estimation would also address the concern stated above.

      We appreciate this suggestion and quantifying the affinities of the Gβγ–PLCβ3 interactions is the subject of future studies.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Please correct PIP2 to the correct PIP2 (many examples throughout).

      These have been corrected.

      Reviewer #2 (Recommendations for the authors):

      Minor comments:

      (1) Figure S1B: The lane-condition labels above the third gel appear to be incorrect, as both lanes are marked identically (Gβγ +/+, PLCβ variant +/+, BMOE +/+) despite clearly different banding patterns. Please confirm.

      We have confirmed the markings above the gels are correct.

      (2) In Figure S2, the authors show three fitted models, but the helical density for Gβγ cannot be seen in two of them at the displayed contour level. The authors should provide views of the map at different contour levels (thresholds) to better support the model fitting. Otherwise, it is difficult to assess whether the Gβγ subunit could adopt alternative orientations (i.e., whether it may be rotated) within the density.

      We have included a new figure (Figure S3) that provides images of the maps at different contour levels.

      (3) Page 5: "where PLCβ3 is increased by the overexpression of either Gβγ or Gαq" should be revised to "where PLCβ3 activity is ...".

      This sentence has been corrected.

      (4) Figure 2: Please label residue R215 in Figure 2A/2B (or the relevant structural panel), since R215E is tested in 2E but the position is not shown.

      R215 is now included in Figure 2C.

      (5) Page 19, Figure 2 legend: "Changes ... Figure S3" should be "Changes ... Figure S5".

      We have corrected this figure call.

      Reviewer #3 (Recommendations for the authors):

      The studies seem well carried out, although more details regarding the BRET controls and the significance of the values should be included.

      We have revised the captions to provide more details about the experimental controls and a brief description of significance. Individual p-values are included in the supplemental tables.

    1. eLife Assessment

      This fundamental work uncovers an unexpected lysosomal function for NINJ2 and links it to ferroptosis and cancer biology. The evidence supporting the conclusions appears to be convincing. This work will be of general interest to the community of ferroptosis and cancer biology.

    2. Reviewer #1 (Public review):

      Summary:

      This study reports a novel and potentially impactful role for NINJ2 in maintaining lysosomal integrity and regulating cellular susceptibility to ferroptosis. The authors demonstrate that NINJ2 localizes to lysosomes and interacts with LAMP1, a key lysosomal membrane glycoprotein involved in sensing lysosomal stress. Loss of NINJ2 increases lysosomal membrane permeabilization (LMP), resulting in selective leakage of lysosomal contents, including labile iron, into the cytosol. The authors further show that NINJ2 deficiency reduces the expression of ferritin storage proteins, thereby sensitizing cells to ferroptosis induced by RSL3 and erastin. Collectively, the work proposes a mechanistic link between NINJ2-mediated control of LMP, iron homeostasis, and ferroptotic vulnerability, with potential relevance to cancer biology.

      Strengths:

      This study identifies a novel role for NINJ2 in regulating lysosomal integrity and ferroptosis and establishes a mechanistic link between lysosomal membrane permeabilization, iron homeostasis, and ferroptotic sensitivity, with potential translational relevance in cancer.

      Weaknesses:

      The results overall support the authors' conclusions and provide a plausible mechanistic framework; however, additional quantification of western blot data and further discussion of mechanistic questions would strengthen the study.

      The findings are likely to have broad impact by linking lysosomal integrity to ferroptosis and iron homeostasis, both of which are relevant to cancer biology and therapeutic targeting.

      Comments on revised version.

      The authors have addressed all of my comments and questions. I have no further concerns.

    3. Reviewer #2 (Public review):

      This manuscript, "Nerve Injury-Induced Protein 2 preserves lysosomal membrane integrity to suppress ferroptosis", identifies a previously unrecognized function of NINJ2 as a regulator of lysosomal membrane integrity and iron homeostasis, thereby suppressing ferroptosis. The authors demonstrate that NINJ2 localizes to lysosomes, interacts with LAMP1, limits lysosomal membrane permeabilization (LMP), stabilizes ferritin, and protects cells from ferroptotic cell death. They further extend these mechanistic findings to human cancer datasets, showing co overexpression and positive correlation of NINJ2 with ferritin genes in iron addicted cancers.

      Overall, the study is conceptually interesting, technically solid, and integrates cell biology, iron metabolism, and ferroptosis in a coherent framework. The work expands the functional repertoire of the Ninjurin family beyond plasma membrane rupture and inflammation, which will be of interest to researchers in cell death, lysosome biology, and cancer metabolism.

      Strengths:

      (1) The identification of NINJ2 as a lysosome-associated protein that suppresses ferroptosis represents a meaningful advance beyond its previously described roles in inflammation, pyroptosis, and tumorigenesis.

      (2) The work distinguishes NINJ2 functionally from NINJ1, reinforcing the idea that structurally related Ninjurins have divergent membrane-related roles.

      (3) The study presents a logically connected pathway:<br /> NINJ2 loss → LMP → labile iron increase → ferritin degradation → ferroptosis sensitization, which is well supported by the data.

      (4) The link between LAMP1, ferritin turnover, and ferroptosis is particularly compelling and timely given recent interest in lysosomal contributions to ferroptotic signaling.

      (5) The authors use confocal microscopy, proximity ligation assays, biochemical IPs, iron measurements, protein half-life analyses, ferroptosis assays, and TCGA-based analyses, providing convergent evidence for their model.

      (6) Use of two distinct cell lines (MCF7 and Molt4) strengthens generalizability.

      (7) The integration of cancer expression datasets linking NINJ2 with ferritin expression in hepatocellular and breast carcinomas enhances translational relevance.

      (8) Assigning NINJ2 a lysosomal protective function, distinct from NINJ1-mediated plasma membrane rupture is novel.

      (9) Linking NINJ2 to ferroptosis regulation via lysosomal iron handling, rather than canonical GPX4 or system Xc⁻ pathways is also novel, along with proposing a NINJ2-LAMP1-ferritin axis as a buffering mechanism against iron-driven lipid peroxidation.

      (10) These insights are not incremental; they reframe how NINJ2 may function at the intersection of membrane biology, iron metabolism, and regulated cell death.

    4. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This study reports a novel and potentially impactful role for NINJ2 in maintaining lysosomal integrity and regulating cellular susceptibility to ferroptosis. The authors demonstrate that NINJ2 localizes to lysosomes and interacts with LAMP1, a key lysosomal membrane glycoprotein involved in sensing lysosomal stress. Loss of NINJ2 increases lysosomal membrane permeabilization (LMP), resulting in selective leakage of lysosomal contents, including labile iron, into the cytosol. The authors further show that NINJ2 deficiency reduces the expression of ferritin storage proteins, thereby sensitizing cells to ferroptosis induced by RSL3 and erastin. Collectively, the work proposes a mechanistic link between NINJ2-mediated control of LMP, iron homeostasis, and ferroptotic vulnerability, with potential relevance to cancer biology.

      Strengths:

      This study identifies a novel role for NINJ2 in regulating lysosomal integrity and ferroptosis and establishes a mechanistic link between lysosomal membrane permeabilization, iron homeostasis, and ferroptotic sensitivity, with potential translational relevance in cancer.

      Weaknesses:

      (1) The results overall support the authors' conclusions and provide a plausible mechanistic framework; however, additional quantification of Western blot data and further discussion of mechanistic questions would strengthen the study.

      All the western blot data have been quantified throughout the manuscript in the revised version.

      (2) The findings are likely to have a broad impact by linking lysosomal integrity to ferroptosis and iron homeostasis, both of which are relevant to cancer biology and therapeutic targeting.

      We thank the reviewer’s comment. We have discussed the potential implications of these findings for cancer treatment in the “Discussion” section.

      Reviewer #2 (Public review):

      This manuscript, "Nerve Injury-Induced Protein 2 preserves lysosomal membrane integrity to suppress ferroptosis", identifies a previously unrecognized function of NINJ2 as a regulator of lysosomal membrane integrity and iron homeostasis, thereby suppressing ferroptosis. The authors demonstrate that NINJ2 localizes to lysosomes, interacts with LAMP1, limits lysosomal membrane permeabilization (LMP), stabilizes ferritin, and protects cells from ferroptotic cell death. They further extend these mechanistic findings to human cancer datasets, showing cooverexpression and positive correlation of NINJ2 with ferritin genes in iron-addicted cancers.

      Overall, the study is conceptually interesting, technically solid, and integrates cell biology, iron metabolism, and ferroptosis in a coherent framework. The work expands the functional repertoire of the Ninjurin family beyond plasma membrane rupture and inflammation, which will be of interest to researchers in cell death, lysosome biology, and cancer metabolism.

      Strengths:

      (1) The identification of NINJ2 as a lysosome-associated protein that suppresses ferroptosis represents a meaningful advance beyond its previously described roles in inflammation, pyroptosis, and tumorigenesis.

      (2) The work distinguishes NINJ2 functionally from NINJ1, reinforcing the idea that structurally related Ninjurins have divergent membrane-related roles.

      (3) The study presents a logically connected pathway:

      NINJ2 loss → LMP → labile iron increase → ferritin degradation → ferroptosis sensitization, which is well supported by the data.

      (4) The link between LAMP1, ferritin turnover, and ferroptosis is particularly compelling and timely given recent interest in lysosomal contributions to ferroptotic signaling.

      (5) The authors use confocal microscopy, proximity ligation assays, biochemical IPs, iron measurements, protein half-life analyses, ferroptosis assays, and TCGA-based analyses, providing convergent evidence for their model.

      (6) Use of two distinct cell lines (MCF7 and Molt4) strengthens generalizability.

      (7) The integration of cancer expression datasets linking NINJ2 with ferritin expression in hepatocellular and breast carcinomas enhances translational relevance.

      (8) Assigning NINJ2 a lysosomal protective function, distinct from NINJ1-mediated plasma membrane rupture, is novel.

      (9) Linking NINJ2 to ferroptosis regulation via lysosomal iron handling, rather than canonical GPX4 or system Xc</sup>-</sup> pathways, is also novel, along with proposing a NINJ2-LAMP1-ferritin axis as a buffering mechanism against iron-driven lipid peroxidation.

      (10) These insights are not incremental; they reframe how NINJ2 may function at the intersection of membrane biology, iron metabolism, and regulated cell death.

      Areas for improvement:

      While the study is strong, several issues should be addressed for mechanistic depth and general relevance.

      (1) Although NINJ2 is shown to interact with LAMP1 and LAMP1 knockdown rescues ferritin levels, it remains unclear whether the NINJ2-LAMP1 interaction is required for lysosomal protection. The authors could: a) Map the NINJ2 domain required for LAMP1 interaction and test whether an interaction-deficient mutant fails to protect against LMP and ferroptosis. b) Rescue NINJ2 KO cells with wild-type versus mutant NINJ2 to establish causality.

      We thank the reviewer’s comments. Ongoing work in our laboratory is focused on elucidating the molecular mechanism by which the NINJ2-LAMP1 interaction regulates lysosomal membrane integrity and ferroptosis, and we anticipate reporting these findings in a future publication.

      (2) The conclusion that NINJ2 suppresses ferroptosis relies primarily on RSL3 and Erastin sensitivity. A direct assessment of ferroptosis would hence the study, such as:

      (a) Include ferroptosis rescue experiments using ferrostatin 1 or liproxstatin 1.

      (b) Assess lipid peroxidation directly (e.g., C11 BODIPY staining) to strengthen the ferroptosis claim.

      We thank the reviewer for this thoughtful comment. We agree that ferrostatin-1 or liproxstatin-1 rescue experiments, together with direct analysis of lipid peroxidation, would provide complementary evidence for ferroptosis. We will incorporate these additional experiments in future studies to further strengthen the mechanistic basis of our findings.

      (3) The manuscript discusses lysosomal ferritin degradation but does not directly examine NCOA4, a central mediator of ferritinophagy. It would be good to: a) Test whether NCOA4 knockdown rescues ferritin loss and ferroptosis sensitivity in NINJ2 KO cells. b) This would clarify whether NINJ2 acts upstream of canonical ferritinophagy pathways or via an alternative mechanism.

      We appreciate the reviewer's thoughtful suggestion. Defining the contribution of NCOA4 to NINJ2-mediated ferritin degradation is an important question that could further clarify the underlying mechanism. Addressing this issue will require a comprehensive set of additional experiments, which will be addressed in the future studies.

      (4) The study is entirely cell-based, despite references to inflammatory and tumor phenotypes in Ninj2-deficient mice. While not strictly required, even limited in vivo validation (e.g., ferroptosis markers or iron accumulation in existing Ninj2 KO tissues) would substantially strengthen the manuscript.

      We thank the reviewer for this insightful suggestion. We agree that in vivo validation of ferroptosis markers and iron accumulation in Ninj2-deficient tissues would further strengthen our conclusions. However, these experiments will require substantial additional investigation, which will be pursued in the future studies.

      (5) Finally, most imaging data (e.g., Galectin 3/LAMP1 colocalization, PLA signals) and immunoblot data are presented qualitatively. The authors should provide the qualifications of Western blots and other measurements.

      All the western blot data have been quantified throughout the manuscript in the revised version.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) What mechanisms might underlie the regulation of LAMP1 transcript levels by NINJ2?

      A clear mechanism by which NINJ2 regulates LAMP1 transcripts has not been elucidated and warrants further investigation. Nevertheless, several possibilities can be considered. First, LAMP1 transcription is known to be regulated by TFEB (transcription factor EB), a master regulator of the lysosomal–autophagy pathway. Upon lysosomal membrane permeabilization (LMP), TFEB translocates to the nucleus and activates a broad set of lysosome-related genes, including LAMP1. Notably, phosphorylation of TFEB by mTORC1 at Ser211 inhibits its activity by preventing nuclear translocation. Thus, it would be of interest to determine whether NINJ2 modulates TFEB phosphorylation status and subcellular localization. In addition, the tumour suppressor p53 has been reported to engage in complex crosstalk with TFEB in regulating basal autophagy. Interestingly, p53 expression is increased in NINJ2-KO cells. It is therefore plausible that NINJ2 regulates TFEB activity through p53, or alternatively modulates the p53–TFEB signaling axis more broadly to maintain lysosomal integrity.

      (2) Does Ninjurin1 play a similar role in regulating lysosomal membrane permeabilization (LMP)?

      At this moment, it remains unclear whether NINJ1 plays a role similar to that of NINJ2 in regulating LMP. In fact, our previous studies demonstrated that NINJ2 physically interacts with NINJ1 and may antagonize NINJ1-mediated pyroptosis. Furthermore, NINJ1 has recently been reported to promote ferroptosis by interacting with the xCT cystine/glutamate antiporter (PMID: 38464226), a function that contrasts with the protective role of NINJ2 against ferroptosis identified in the present study. These findings suggest that NINJ1 and NINJ2 may have distinct, or even opposing, functions in regulating cell death pathways. Nevertheless, further studies are required to determine whether NINJ1 also participates in the regulation of LMP and to define its relationship with NINJ2 in maintaining lysosomal membrane integrity.

      (3) What are the potential clinical implications of these findings, particularly in the context of cancer progression or therapeutic targeting?

      Targeting NINJ2 may have important clinical implications in cancer therapy. Given its role in maintaining lysosomal membrane integrity, inhibition or loss of NINJ2 could promote lysosomal membrane permeabilization (LMP), thereby sensitizing cancer cells to ferroptosis through increased intracellular labile iron accumulation and disruption of redox homeostasis, ultimately enhancing tumor cell killing. As such, NINJ2 inhibition may represent a strategy to selectively destabilize lysosomal function in cancer cells and improve responsiveness to ferroptosis-inducing agents or other combination therapies that exploit oxidative stress vulnerabilities. Indeed, we previously developed a peptide that targets NINJ2. Whether this NINJ2-targeting peptide can sensitize cancer cells to ferroptosis therefore warrants further investigation.

      (4) The authors should provide quantification of all Western blot data throughout the manuscript to enhance data robustness and reproducibility.

      All the western blot data have been quantified throughout the manuscript in the revised version.

      Reviewer #2 (Recommendations for the authors):

      (1) Controls for knockdown efficiency of NINJ2 (Figure 2D) should be shown.

      NINJ2-KO MCF7 cells were generated previously and published in the article (PMID: 38325550) along with sequencing confirmation.

      (2) In Figure 2A legends, the concentration of LLOMe is 1mM or 1µM - need to be clarified?

      The concentration for LLOME is 1µM. This typo has been corrected.

      (3) In the figure legends section, "Figure 4" is missing.

      We thank the reviewer’s comment. Figure 4 has been added to the Figure legends.

      (4) Some description of NINJ2 ko cells generation should be included in the materials section.

      In the Materials and Methods section, we briefly described how these cell lines were generated and cited the original publication (PMID: 38325550)

      (5) The manuscript would benefit from a schematic model figure summarizing the proposed NINJ2-LAMP1-iron-ferroptosis axis.

      In the revised manuscript, we provided a model to elucidate the role of NINJ2 in modulating lysosomal membrane integrity and iron homeostasis.

      (6) Some sections of the Introduction are lengthy and could be streamlined to focus more directly on lysosomes and ferroptosis.

      We have streamlined the introduction.

      (7) Statistical methods should clarify whether data meet assumptions for Student's t-test and whether multiple comparisons were corrected where applicable.

      Statistical analysis has been added to the figure legends when applicable.

    1. eLife Assessment

      This valuable study uses genetic, behavioural, neuronal activity and cell-based assays to show that the Drosophila ionotropic receptor IR20a contributes to the detection of arginine and low-sodium salt, and suggests that peripheral receptor combinations may help integrate these nutrient cues. The evidence for a role for IR20a and associated co-receptors in these responses is solid, but the evidence that distinct receptor assemblies mediate receptor-level multimodal integration is incomplete, as this central model is inferred from functional data and leaves some discrepancies between behaviour, physiology and prior work unresolved. The study will be of interest to sensory neuroscientists and chemosensory biologists.

    2. Reviewer #1 (Public review):

      Summary:

      The authors investigated the function of a Drosophila chemosensory receptor, IR20a, using genetics, neuronal histology, calcium imaging (in vivo and in cultured cells), and behavioral approaches. They provide evidence that this receptor functions in the detection of the amino acid arginine and of low salt (NaCl) concentrations, functioning in different combinations with "co-receptor" IRs, IR25a and IR76b.

      Strengths:

      The experiments are generally very well-performed and clearly presented, using established methodology. While, unsurprisingly, some puzzles remain (mentioned below), the work provides one of the clearest lines of evidence for the combinatorial coding of sensory information at the periphery through the combined action of distinct sets of chemosensory IRs.

      As taste neurons have long been recognized to express many different combinations of IRs and Gustatory Receptors (GRs), this study will be of interest to chemosensory biologists in general, particularly those studying invertebrate model systems (though co-expression of different families of taste receptors is a feature of mammalian taste cells).

      The precise molecular mechanisms remain unclear: there is no direct evidence here for protein complex formation (though this is likely), the stoichiometry of such complexes, or how subunits interact to confer or suppress sensory sensitivity. Nevertheless, these receptors, and the authors' success in reconstituting functionality in cultured cells, might make these a powerful model to explore such questions in the future.

      Weaknesses:

      Given the particular interest of the data from the heterologous reconstitution in cultured cells, the authors should be quite explicit about the nature of the quantification of the S2 cell responses. It is unclear whether the cited "n" refers to numbers of cells or something else, and whether all or only a fraction of (transfected) cells gave responses.

      There has been some prior work on the context-specific role of IR76b in amino acid-sensing and salt sensing by Ganguly and colleagues (Cell Reports 2017), who also implicated (weakly) a contribution of IR20a in contributing to the amino acid-sensing role. In that work, the authors focussed principally on the labellum and used electrophysiology rather than calcium imaging. The present manuscript appears rather dismissive of the earlier results (only mentioning them in the Discussion), and the authors could be a bit more generous about what was previously determined, where they have confirmed previous findings, where their results diverge, and why this might be. Similarly, the original functional analysis of IR76b (Zhang Science 2013) argued this was a low-salt sensor by itself, which is at least partially corroborated here; it remains unclear how this role relates to the low-salt detecting function of a potential complex of IR20a/IR25a/IR76b. It would be useful to have a summary model of the possible variety of complexes of IRs in different types of sensory neurons, as supported by the results in this and previous studies.

      The discord between the lack of requirement for IR20a for physiological responses to arginine in tarsi versus the necessity for behavioral responses is puzzling (though might reflect a labellar role for IR20a). There appears to be a trend of a decrease in calcium signal in tarsi to 100 mM arginine, which is the highest concentration tested (Figure 2A, C). Would a statistically significant decrease be observed with lower arginine concentrations? (A more substantial experiment would be to perform calcium imaging in the labellar IR20a neurons, or their axonal projections in the SEZ; this is not necessary, but the authors should at least acknowledge that their imaging of tarsal responses, while convenient, only examines a tiny fraction of the entire IR20a neuron population.

      The authors argue for synergistic responses to arginine and NaCl mediated by IR20a/IR25a. It's not clear to me to what extent there is synergism. In Figure 5A, 10 mM arginine or 10 mM NaCl individually lead to c.30-40% PER, and then when both are presented together in the "Mix" (presumably both compounds at 10 mM?), PER rises to c.60%. Is this really synergism, or rather simple additivity of behavioral responses to two attractive compounds? The authors could discuss this more thoroughly. Similarly, in Figure 5G the authors show that 50 mM arginine does not evoke a significant response in S2 cells expressing IR25a/IR20a, but in Figure 4 it would seem likely that a 50 mM dose would produce a significant response (the response to 25 mM arginine in Figure 4F is already elevated above the control, albeit not statistically significant). Is this just a batch effect of the experiments performed at different times (so they are not directly comparable)?

      The legend title to Figure 6 implies cooperation between tonic and state-modulated pathways, but I don't see specific evidence for "cooperation". Rather, as in the results text, they seem to work in parallel, so this analysis seems slightly peripheral to the main focus of the manuscript. It's ultimately unclear how the IR20a/IR76b/IR25a low salt sensor and the sensor containing IR56b functionally interact at the behavioral level. Here, a graphical summary, as mentioned above, of the different salt sensing neurons, the receptors they use, and the behaviors they control could be useful to establish the current knowledge and highlight open questions for the future.

    3. Reviewer #2 (Public review):

      Summary:

      This study identifies IR20a-expressing gustatory neurons in Drosophila as a multimodal sensory population integrating amino acid (arginine) and low-salt signals through combinatorial IR20a/IR25a/IR76b receptor assemblies. The proposed model of peripheral-level signal integration and synergistic enhancement of feeding preference is potentially significant, as it expands current understanding of gustatory coding beyond single-modality labeled lines.

      Strengths:

      Overall, the findings are conceptually interesting and suggest a novel framework for multimodal taste integration, but some mechanistic interpretations remain incompletely supported by direct evidence.

      Weaknesses:

      (1) Although the authors demonstrate co-expression of IR20a, IR25a, and IR76b in the same GRN population, this evidence is insufficient to support the proposed model of distinct receptors coexisting within individual neurons. Additional molecular or structural data would be required to distinguish whether these subunits assemble into complexes.

      (2) Given that IR76b has already been established as a sodium/salt sensing channel, the novelty of this study relies on the proposed role of IR20a in conferring multimodal integration and synergy. However, it remains unclear whether this represents a fundamentally new sensory mechanism or a re-interpretation of known IR76b-dependent salt responses in a different neuronal context.

      (3) Line 127:<br /> -The statement that there is no overlap between IR20a-GAL4 and GR64f-LexA or GR66a-LexA is not sufficiently supported by the presented imaging data. In particular, the resolution and clarity of the confocal images in Figure 1 appear suboptimal, making it difficult to confidently assess co-localization. The authors are encouraged to provide higher-resolution images or additional quantitative co-localization analysis to substantiate this conclusion.<br /> -In addition, the images shown in Figure 1 F1-F2 suggest possible partial overlap between IR20a and GR66a signals, which appears inconsistent with the authors' statement of no co-expression. This discrepancy should be clarified.

      (4) Lines 138-141:<br /> There appears to be a discrepancy between imaging and behavioral data: IR20a is reported as dispensable for arginine-evoked neural responses, yet IR20a mutants show significantly reduced attraction to arginine in behavioral assays. The authors should clarify how behavioral deficits arise in the absence of detectable changes in calcium imaging,

      (5) The manuscript proposes that IR20a functions in combination with IR25a to mediate multimodal detection of arginine and low NaCl. However, the specific role of IR25a in this context remains unclear.

      (6) The authors report that co-expression of IR20a and IR25a confers synergistic responses to combined arginine and NaCl stimulation, whereas the inclusion of IR76b abolishes this response (Figures 5E-K). This is an intriguing and potentially important finding; however, the mechanistic basis for this suppression is not clearly explained.

      (7) The authors propose that IR56b mediates state-dependent modulation of low-salt preference. However, the current data do not clearly distinguish whether IR56b acts as a real nutrient state sensor or just functions as a downstream modulatory component within a broader feeding circuit. Additional evidence linking IR56b activity changes to upstream metabolic state signals would be necessary to support the interpretation that IR56b functions as a primary state sensor.

      (8) The manuscript suggests that IR20a and IR56b define two parallel and functionally independent pathways mediating nutrient detection and state-dependent preference, respectively. However, this conclusion is not fully supported by the current dataset. While the two receptors are shown to be expressed in distinct neuronal populations, the possibility of indirect interactions or convergence at downstream circuit nodes has not been excluded. Given that both pathways ultimately influence feeding behavior, it remains possible that they converge at higher-order interneurons or shared neuromodulatory circuits.

      (9) In the state-dependent feeding assays (Figure 6), using H2O as a control introduces a severe masking effect. Salt-deprived flies actively suppress pure water intake to avoid osmotic shock, which artificially inflates the Preference Index (P.I.) for salt due to the denominator effect. To cleanly isolate salt preference from the thirst/osmotic drive, the authors will need to utilize an "isosmotic sucrose vs. isosmotic sucrose + salt" paradigm (Jaeger et al., 2018, eLife; Puri et al., 2026, PNAS).

    4. Reviewer #3 (Public review):

      Summary:

      Drosophila, like other animals, use sophisticated taste systems with specialized chemoreceptors to identify gustatory cues in their environment. Multiple gustatory cues associated with a food source are often encountered simultaneously, but our understanding of how this sensory information is detected and integrated remains incompletely understood. This valuable study investigates how salt, amino acids, or their combination are detected by specific combinations of peripheral Ionotropic Receptors, leading to behavioral attraction. The authors show that distinct combinations of IR76b, IR25a, and IR20a confer sensitivity to salt, arginine, or both. They also show striking evidence that cells co-expressing IR25a/IR20a display a synergistic response to a mixture of sub-activating concentrations of these tastants. Together, these experiments lead to the conclusion that combinatorial expression of different subunits and synergistic responses to taste mixtures facilitates integration of taste cues beginning in the periphery. However, in its current form, key methodological details are missing or inadequately described, which complicates interpretation. Additionally, characterization is heavily focused on the population of IR20a+ neurons in the tarsi, while the response properties of the newly-identified, functionally distinct population in the labellum are investigated only through behavioral analysis, limiting the description of potentially additional IR20a complexes. Ultimately, more in-depth biochemical characterization of the IR complexes described will be required to fully support the conclusion that combinatorial assembly of distinct IR20a receptors enables peripheral integration of taste mixtures.

      Strengths:

      The authors characterize the expression pattern of IR20a in the tarsi as well as in the labellum, a tissue for which IR20a expression has been a point of debate. Multiple levels of analysis, including behavioral assays, physiological recordings, as well as ectopic and heterologous expression systems, are used to characterize the response properties of different combinations of IR subunits, demonstrating remarkably consistent behavior of the IR-complexes across cell types. Well-controlled genetic analysis and the use of multiple behavioral assays provide additional support for their results, including the surprising demonstration of synergistic responses to mixtures of tastants that supports the idea of peripheral integration of gustatory inputs. This report also identifies a distinct IR, IR56b, required for starvation-enhanced responses to salt.

      Weaknesses:

      (1) The title states that IR20a integrates L-arginine and salt signals via distinct subunit assemblies, though the paper lacks direct evidence that IR20a serves as a multimodal tuning receptor in distinct functional assemblies. Heterologous expression shows that co-expression of IR20a/IR25a confers sensitivity to Arg, IR76b confers sensitivity to NaCl, and IR20a/IR25a/IR76b co-expression confers sensitivity to both Arg and NaCl. This seems to be interpreted to mean that all three subunits are assembling into a single complex. However, current results do not show any difference in salt response when IR76b is expressed alone compared to alongside IR25a+IR20a. Without more direct evidence for co-assembly of all three subunits, it is equally plausible that the responses observed represent activity of distinct IR25a/IR20a and IR76b receptors for Arg and salt, respectively. In this model, genetic disruption resulting in expression of either IR25a or IR20a alone with IR76b could disrupt its activity or membrane trafficking (as seen here and in previous studies) while co-expression of both IR20a and IR25a relieves this inhibition by sequestering IR20a/IR25a into a distinct complex from IR76b. Direct biochemical characterization, for instance in the form of co-immunoprecipitation or FRET, will be required to differentiate between these possibilities.

      (2) Key methodological details are missing throughout the manuscript. For instance, incomplete genotype and staining information is provided for images in Figure 1, making it difficult to interpret what is being shown. Additionally, for the calcium imaging methods, what is the imaging speed? How are max values calculated (is this the average of several images or just a single maximum)? How are ligands diluted and delivered to cells, and were they applied in a manner that allowed for subsequent washout?

      (3) The composition of the S2 imaging bath buffer requires clarification. As described, the bath buffer appears to lack any Ca2+ or other IR-permeable cations. If this is indeed the case, more detail should be provided about why this bath buffer was selected and what this means for the source and mechanism of calcium responses observed, since it would not reflect direct IR-mediated transduction. It is also notable that addition of water gives such a detectable change in the tarsal preps.

      (4) Visualization of IR20a driver activity in the labellum is interesting. Previous descriptions of labellar expression of IR20a range from no expression to expression in bitter neurons, so the current data linking IR20a to a different population of IR76b+ neurons warrants careful analysis in light of this discrepancy. However, some of the strongest presented evidence for expression is found in Figure 1, where the images are quite small, making it difficult to distinguish the morphology and sensillar innervation pattern of the cells labeled by the IR20a driver. In Figure 1A, several of the arrows do not appear to be associated with any visible fluorescence. It is similarly difficult to assess overlap. Including higher-resolution images and/or validating labellar expression, using antibodies, in situ hybridization, RT-PCR, or transcriptomics would strengthen these claims.

      (5) Similarly, Figure 2 shows that IR20a is not required for Ca2+ responses to AAs or KCl in the legs, but is required for behavioral preferences and PER responses in the labellum. This suggests that IR20a receptors may function differently in different tissues, though direct evidence is lacking. Calcium imaging from a weakly expressed driver may be difficult, but electrophysiological recordings from relevant labellar sensilla or ectopic/heterologous reconstitution of the molecular receptors found there would give important insights into the response properties of these other IR20a receptor type(s) and could provide evidence for additional IR20a-containing complexes. The current paper focuses exclusively on IR25a/IR20a/IR76b, which do seem to reliably reproduce the Arg/NaCl responses observed in the tarsi, but even for the tarsal neurons it is unclear that this represents an exhaustive list of all the relevant IR20a-interacting subunits coexpressed in these cells. For instance, Koh et al., 2014 (PMID: 25123314) found several additional IR driver lines, including IR56b, were active in the 5v/s tarsal sensilla.

    5. Author response:

      We thank the editors and three reviewers for their careful evaluation and constructive feedback. We are pleased that our identification of IR20a as a multimodal tuning receptor required for both low-salt and arginine sensing in a distinct gustatory neuron population was recognized as a valuable contribution to sensory coding. We agree with the major points raised and outline our planned revisions below, organized thematically.

      (1) Receptor assembly and integration model

      Our genetic, heterologous, and calcium imaging data show functional cooperation among IR20a, IR25a, and IR76b but do not demonstrate physical association. In the revision we will replace terms like “distinct subunit assemblies” and “peripheral integration” with more cautious language such as “functional receptor combinations.” We will state explicitly that our data cannot resolve whether the three IRs form a single heteromeric complex, and we will discuss the alternative possibility that IR76b and the IR20a/IR25a pair function as separate receptors within the same neuron. This interpretation better accounts for the response patterns observed upon co-expression. Throughout the text and in a revised summary figure, we will clearly differentiate elements directly supported by data from those that remain inferential.

      (2) Tarsal calcium imaging versus behavior

      The mismatch between tarsal calcium imaging and behavioral arginine responses will be addressed directly. We will explain that tarsal recordings sample only a small subset of IR20a neurons, whereas proboscis extension and feeding assays predominantly engage the more numerous labellar sensilla, whose neurons may carry different receptor compositions. We will generate labellar imaging where feasible; if additional functional data cannot be obtained, we will acknowledge this limitation rather than overinterpreting the tarsal results.

      (3) Synergy versus additivity

      We will discuss behavioral and cellular data separately. Recognizing that true synergy is difficult to demonstrate in feeding and proboscis extension assays, we will adopt conservative terminology when interpreting those experiments. For cellular data, where mechanistic insight is stronger, we will present the evidence for functional synergy and discuss why the outcomes may differ between the cellular and organismal levels.

      (4) Contextualizing prior IR76b and IR20a literatures

      We will expand the Introduction and Discussion to cite more fully the works establishing IR76b as a low-salt sensor and IR20a’s role in amino-acid sensing, including the earlier report that IR20a overexpression can inhibit IR76b-dependent salt responses. We will clarify how our single-cell imaging and loss-of-function data obtained in the native context refine models derived from ectopic expression. In its endogenous setting, IR20a marks neurons narrowly tuned to amino acids such as arginine, and IR20a is strictly required for low-salt detection. These findings contrast with the earlier view that IR20a functions broadly as an amino-acid sensor or as a salt-response blocker.

      (5) State-dependent modulation and IR56b

      Our data do not identify IR56b as the direct molecular sensor of internal state. We will reframe IR56b as a necessary component for state-dependent modulation of low-salt preference and retract any claim that it is the sensor itself. We will also clarify that IR20a and IR56b define genetically separable peripheral pathways that may converge on downstream circuits.

      (6) Methodological and presentation issues

      We will address the following points raised across reviews:

      (1) Co-localization: Higher-resolution confocal images and co-localization analysis will be provided.

      (2) Summary model figure: A new figure will illustrate the distinct functions of IRs in low-salt and amino-acid taste, clearly indicating which aspects are directly supported and which are inferential.

      (3) Feeding-assay control: We will either include an isosmotic sucrose control to avoid the water confound or explicitly discuss this limitation and temper the interpretation of feeding-preference results.

      (4) S2 cell quantification: Complete details on response criteria, responder fractions, and statistical reporting will be added.

      (5) Figure and supplementary corrections: All noted errors in figure legends, scale bars, citations, and supplementary-file mismatches will be fixed.

      We are confident these revisions will bring our mechanistic claims into close alignment with the evidence and substantially improve the manuscript. We again thank the editors and reviewers for their detailed and helpful comments.

    1. eLife Assessment

      This Review delineates postmenopausal age-related lobular involution (ARLI) in the breast that remains only partially resolved. In contrast to the conventional view of persistent lobules as passive residual structures, this work defines them as an actively maintained senescence-immune reserve niche. Inclusion of operational definitions for the reserve state in human breast tissue would strengthen the article. The work will be of interest to scientists working in the fields of breast disorders.

    2. Reviewer #1 (Public review):

      Summary:

      This Perspective proposes a conceptual model in which incomplete age-related lobular involution (ARLI) in the breast reflects an actively maintained senescent-immune "reserve niche," rather than simply passive failure of lobular regression after menopause. The authors aim to integrate breast cancer epidemiology, mammary gland biology, cellular senescence, immune surveillance, and comparative reserve-tissue systems to explain why persistent postmenopausal lobules are associated with increased breast cancer risk. The manuscript is ambitious, creative, and potentially useful in shifting attention from residual epithelial quantity alone toward the microenvironmental state of persistent lobules.

      Strengths:

      A major strength of the manuscript is its forward-looking synthesis. The authors bring together several areas that are often considered separately: ARLI as a tissue-level risk marker, inflammatory features of incompletely involuted breast tissue, senescence biology, macrophage-mediated remodeling, and the menopausal transition as a potential window of biological plasticity. The model is conceptually interesting and, if supported by future evidence, could stimulate new approaches to risk stratification and prevention focused on the perimenopausal period.

      Weaknesses:

      However, the current manuscript often presents the proposed model with more certainty than the available evidence supports. The evidence clearly supports associations among incomplete ARLI, inflammatory or immune features, and breast cancer risk, but it does not yet demonstrate that senescent cells maintain persistent lobules, that immune clearance failure causes incomplete involution, or that a self-sustaining senescent-immune "niche lock" exists in human breast tissue. Much of the mechanistic framework is extrapolated from other tissues, postpartum involution, or general senescence biology. These are reasonable sources for hypothesis generation, but the manuscript would be stronger if it more clearly distinguished established observations from inference and speculation.

      The senescence component of the model requires stronger and more direct support. Several claims about senescent burden in the aging breast appear to rely on general senescence literature or mammary aging studies that do not directly demonstrate senescence in persistent human TDLUs. This distinction is important because the manuscript's central model depends on senescent cells being spatially and functionally linked to incomplete ARLI.

      The epidemiologic evidence also requires a more balanced treatment. Although several studies support incomplete ARLI as a breast cancer risk-associated phenotype, other cohorts and quantitative approaches have reported attenuated or null associations. This mixed evidence is acknowledged, but it is treated largely as a caveat rather than incorporated into the central argument. For readers, this uncertainty is important for interpreting the strength and generalizability of the proposed model.

    3. Reviewer #2 (Public review):

      Summary:

      This review constructs a novel theoretical framework to elucidate incomplete postmenopausal age-related lobular involution (ARLI) in the breast. Differing from the conventional view of persistent lobules as passive residual structures, the work innovatively defines them as an actively maintained senescence-immune reserve niche. It comprehensively integrates multidisciplinary evidence from breast epidemiology, stromal biology, cellular senescence and immune surveillance, as well as cross-tissue research findings, and identifies menopause as a core biological turning point regulating ARLI and relevant breast cancer risk, providing a new theoretical perspective for subsequent breast cancer risk assessment and preventive intervention research.

      Strengths:

      This study presents an original, logically rigorous, and well-organized research hypothesis. It innovatively breaks through the traditional cognitive perspective of ARLI and adopts a multidisciplinary and cross-tissue analytical approach to sort out relevant biological mechanisms systematically. The proposed theoretical framework is insightful, with good theoretical innovation and potential translational value for guiding breast cancer risk evaluation and targeted prevention strategies.

      Weaknesses:

      The manuscript currently serves primarily as a conceptual framework rather than a rigorously evidenced synthesis. Its central argument relies heavily on cross-sectional correlations and theoretical analogies to other organ systems, lacking operational definitions for the reserve state in human breast tissue.

    4. Author response:

      We thank the editors and reviewers for recognizing the originality and potential value of the proposed framework. The reviews rightly ask us to distinguish three things more sharply: what is established directly in human breast tissue, what is inferred from mammary and aging studies, and what remains hypothesis. We agree, and the revision will make that distinction explicit throughout.

      We will define the proposed reserve state operationally and specify the findings that would distinguish active niche maintenance from passive persistence. Throughout, we will treat passive persistence as a legitimate competing hypothesis rather than a settled question. Heterogeneity and immune or inflammatory associations will be presented as consistent with active maintenance and causal directionality, not as establishing them. We will also integrate the mixed epidemiologic evidence more centrally into the argument, rather than treating it as a caveat.

      We will reassess the evidence for senescence in the aging breast and describe it more precisely, correcting or narrowing statements that outrun the data. The figures will be revised so that observed inflammatory and immune-regulatory features are clearly separated from proposed senescence- and SASP-mediated mechanisms. We will also clarify the limits of our analogies: postpartum involution and cross-tissue reserve systems will be presented as sources of candidate mechanisms and testable predictions, not as direct evidence for the proposed mechanism in human ARLI. Finally, we will frame the translational implications as contingent. They depend on first demonstrating that senescent cells are enriched near persistent lobules, identifying the relevant cell types and immune states, and establishing causal relevance.

      We appreciate the reviewers' constructive suggestions. We believe these revisions will preserve the conceptual contribution of the model while making its evidentiary status, the alternative explanations, and its falsifiable predictions substantially clearer.

    1. Author response:

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #1 (Public review):

      I thank the authors for the revised manuscript and for the detailed responses.

      I think the main points raised in the review have now been addressed. In particular, the new experiment with TbPLK inhibition and mass spectrometry is an important addition, as it provides direct evidence that phosphorylation of KIN-G at Thr301 and Ser569 depends on TbPLK activity in cells.

      I also appreciate that the authors have toned down the interpretation of the Golgi phenotype. The revised text now makes clear that the fluorescence data show altered Golgi/ERES organization or duplication, but do not prove a structural defect in Golgi biogenesis.

      The added discussion of the T301A result is also helpful. The finding that only a small fraction of KIN-G is phosphorylated at Thr301 in asynchronous cells makes the lack of a strong T301A phenotype more understandable.

      Overall, I am happy with the revision of the beautiful manuscript.

      Thank you for your critical review and insightful comments.

      Reviewer #2 (Public review):

      Summary:

      The authors identify KIN-G as an in vitro substrate for phosphorylation by TbPLK and show that several of the in vitro P-ated sites, including T310, overlap with P-ation sites seen in live cells. The authors further show that PLK-mediated P-ation inhibits KIN-G binding to microtubules in vitro, as does a KIN-G-T301D mutant, and that expression of a KIN-G-T301D Phospho-mimic in T. brucei phenocopies KIN-G RNAi knockdowns, producing defects in cell division, morphogenesis of the centrin arm, FAZ and other cellular structures, as well as misplaced cytokinesis furrow.

      Understanding cytoskeletal rearrangements that drive cell division in T. brucei is an important and unresolved problem, so the work addresses important questions that are of great interest. PLK and KIN-G have previously been shown to be important for cell division and morphogenesis of cytoskeletal structures that drive cell division in T. brucei. The current work advances our understanding by suggesting a potential mechanism by which PLK and KIN-G might participate, namely through PLK-dependent P-ation to control KIN-G MT binding activity.

      Strengths:

      The authors use a rigorous combination of biochemistry, phosphoproteomics, cell biology, and mutant analysis to support their conclusion that PLK-mediated P-ation of KIN-G negatively regulates KIN-G microtubule binding and this may explain the observation that a KIN-G T301 phosphomimic mutant blocks cell division and perturbs biogenesis of cytoskeletal structures that drive cell division and morphogenesis. Combining rigorous and informative in vitro studies with mutant analysis in live cells is a great strength. The work is solid and important, though a few pieces are needed to fully connect the in vitro findings with the in vivo observations, as detailed below.

      Weaknesses:

      Overall, I find this work to be solid, and to provide an important addition to our understanding of mechanisms controlling cell division in T. brucei. The biochemistry, in particular, is rigorous and convincingly demonstrates PLK can P-ate KIN-G, altering its MT-binding ability. Analysis of phospho-mutants of KIN-G in live T. brucei support the conclusion that P-ation of KIN-G at T301 negatively affects KIN-G function in vivo. I think, however, that the results fall short of supporting the title, because, although the data convincingly show that PLK can phosphorylate KIN-G at T301 in vitro, and that T301 is P-ated in vivo, they do formally demonstrate (nor even test) whether PLK is the kinase responsible for this phosphorylation in vivo (experiments to address this seem quite feasible). I also do not see where the authors try to reconcile the absence of phenotype for KIN-G-T301A with the implied importance of KIN-G phosphorylation by PLK in cell division, which calls into question the need for P-ation of KIN-G-T301 in cell division. Suggestions for addressing these concerns are provided below.

      My two main questions are:

      (1) What is the biological relevance of KIN-G P-ation at T301?

      (a) The authors report no defect for the KIN-G-T301A mutant, so what then is the need for T301 P-ation, if the cell gets along fine without it? One step toward addressing this would be to ask what fraction of KIN-G shows P-ation at T301. Although published studies indicate P-ation at T301, it isn't known what percentage of KIN-G in the cell is P-ated. One might anticipate, for example, that T301-P is a small minority of the population in asynchronous cultures and that T301 P-ation increases at specific cell cycle stages.

      (b) Published work links PLK to cell division, FAZ elongation, etc... The current work suggests that one role of PLK is to P-ate KIN-G at T301. In contrast, however, the current work also indicates that P-ation of KIN-G at T301 is unnecessary for normal cell division, FAZ elongation, etc....

      (c) Some experiments or at least commentary on points a and b above would strengthen the paper.

      - The authors have now addressed this question by assessing what % of KING is phosphorylated at T301 and adding commentary on this point in the revised paper.

      - I would suggest that the model (new figure 8) include a dephosphorylation step, as that is proposed by the authors in the text. Also include in the legend some commentary on the role of phosphorylation, which is the center point of this paper, but not currently mentioned.

      We have modified the model in Figure 8 to include dephosphorylation by an unknown protein phosphatase and a statement about the role of TbPLK phosphorylation on KIN-G function. Thank you.

      (2) Is PLK the kinase that P-ates Kin-G T301 in vivo?

      (a) The authors show PLK P-ates T301 (and other residues) in vitro, and that T-301 is P-ated in vivo. To bring the analysis full circle, it would be informative to examine KIN-G P-ation in a PLK mutant or upon inhibition of PLK with published inhibitors. This seems to be a very doable experiment with the tools available.

      - The authors have addressed this question by demonstrating that T301 phosphorylation is reduced upon treatment with a PLK inhibitor, thus supporting that PLK phosphorylated T301 in vivo. It is noted that one might consider an alternate kinase is also able to phosphorylate T301 in absence of PLK activity, as that could explain the relatively low (~27%) reduction in phosphorylation by PLK inhibitor treatment.

      Thank you.

      Reviewer #3 (Public review):

      Summary:

      Here the authors investigate the role of the Trypanosoma brucei polo-like kinase TbPLK in the function of flagellum-associated cellular structures in trypanosomes. They set out to test the hypothesis that a key substrate of TbPLK is the kinesin protein KIN-G, and that TbPLK phosphorylation of KIN-G regulates its functions in cells.

      Strengths:

      Using in vitro biochemistry with purified proteins, the authors convincingly demonstrate that TbPLK phosphorylates KIN-G at 29 sites. Moreover, they convincingly show that phosphorylation at one site, T301, impairs the binding of purified KIN-G to purified microtubules. They further confirm that inhibition of TbPLK in cells reduces KIN-G phosphorylation at T301 (and S569). Using immunofluorescence-based imaging approaches, they also show that TbPLK colocalizes with KIN-G at centrin arms during early S-phase of the cell cycle. Centin arms are structures that are located near the basal body and flagellum and are important for new flagellum biogenesis, Golgi positioning, and cell division. To evaluate the function of KIN-G phosphorylation in cells, they depleted KIN-G by RNAi, simultaneously expressed phospho-mimetic (T301D) and phospho-ablative mutant proteins, and used immunofluorescene to examine the impact on flagellum-associated cellular structures. They show that expression of the phospho-mimetic mutant KIN-G-T301D causes the following defects: reduced cell proliferation, disruption of centrin arm and Golgi biogenesis, impairment of FAZ elongation and flagellum positioning, and misplacement of the cell division plane. The data convincingly support the conclusion that KIN-G phosphorylation on T301 plays an important role in regulating the cellular functions of this kinesin motor protein.

      Weaknesses:

      The authors have addressed prior weaknesses in the manuscript through additional experimentation and rewording of the conclusions.

      Thank you for your critical review of our manuscript and for the very constructive comments and suggestions to improve the manuscript.

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      (There is some redundancy below with my comments in the public review, but I've included here for clarity and further explanation.)

      The authors have addressed my primary concern, as treatment with PLK inhibitor reduces phosphorylation of T301, while also providing some comment on relative impact of PLK-mediated KIN-G phosphorylation.

      It is notable that phosphorylation of T301 was reduced by only ~27%, while phosphorylation of S569 was reduced by ~100% in the presence of PLK inhibitor. The authors note that this might be explained by slower dephosphorylation of T301. In the absence of a phenotype, and with cell doubling continuing unabated in presence of the inhibitor, it is intriguing that more loss is not observed. An alternative explanation is that an alternate kinase might also be able to phosphorylate T301 in the absence of PLK activity, and the authors should consider that possibility.

      We added a sentence in the main text to suggest an alternative explanation.

      The model shown in figure 8 should include a dephosphorylation step, per the authors comments in the text regarding the small fraction of T301 that is phosphorylated and proposal of a phosphorylation/dephosphorylation cycle. The Fig 8 legend needs to have some commentary on the role of phosphorylation, as phosphorylation is the center point of this paper.

      We have modified the model in Figure 8 to include dephosphorylation by an unknown protein phosphatase and a statement about the role of TbPLK phosphorylation on KIN-G function.

      Minor comments for improving the text are:

      (1) The paper overall is clearly written. However, the Discussion starts with a solid sentence, then becomes a bit diffuse in discussing a wide range of PLK activities that were not addressed in the current work. That detracts attention a bit from the central contributions of this paper.

      (2) At least two places in the text state apparent contradictions.

      (a) p.5 and Fig 2C. The authors say microtubule gliding speed was "...insignificantly reduced..." by the TbPLK-K70R mutant, yet they then state that motility was "interfered with". If the effect is "insignificant", why do they claim there is an effect?

      (b) p6 and Fig 3C. The authors report KIN-G-T301A impact on microtubule gliding activity is insignificant, but then say this mutation reduces motility of KIN-G. These statements are contradictory.

      (3) p. 8, and Fig 7. "ventral side" and "leading edge" are not defined but are used to describe the KIN-G RNAi phenotype.

      (4) Fig 7B. Please explain labeling - the new flagellum daughter is indicated as having the old posterior, while the old flagellum daughter cell is indicated as having the new cell posterior. This is counterintuitive to a reader not intimately familiar with the T. brucei cell division process.

      (5) Fig 4, 5, and 7: "% Cells" is reported. Please indicate what number of cells total were examined.

      These minor comments have already been addressed in the previous revision.

    2. Reviewer #3 (Public review):

      Summary:

      Here the authors investigate the role of the Trypanosoma brucei polo-like kinase TbPLK in the function of flagellum-associated cellular structures in trypanosomes. They set out to test the hypothesis that a key substrate of TbPLK is the kinesin protein KIN-G, and that TbPLK phosphorylation of KIN-G regulates its functions in cells.

      Strengths:

      Using in vitro biochemistry with purified proteins, the authors convincingly demonstrate that TbPLK phosphorylates KIN-G at 29 sites. Moreover, they convincingly show that phosphorylation at one site, T301, impairs the binding of purified KIN-G to purified microtubules. They further confirm that inhibition of TbPLK in cells reduces KIN-G phosphorylation at T301 (and S569). Using immunofluorescence-based imaging approaches, they also show that TbPLK colocalizes with KIN-G at centrin arms during early S-phase of the cell cycle. Centin arms are structures that are located near the basal body and flagellum and are important for new flagellum biogenesis, Golgi positioning, and cell division. To evaluate the function of KIN-G phosphorylation in cells, they depleted KIN-G by RNAi, simultaneously expressed phospho-mimetic (T301D) and phospho-ablative mutant proteins, and used immunofluorescene to examine the impact on flagellum-associated cellular structures. They show that expression of the phospho-mimetic mutant KIN-G-T301D causes the following defects: reduced cell proliferation, disruption of centrin arm and Golgi biogenesis, impairment of FAZ elongation and flagellum positioning, and misplacement of the cell division plane. The data convincingly support the conclusion that KIN-G phosphorylation on T301 plays an important role in regulating the cellular functions of this kinesin motor protein.

    3. Reviewer #2 (Public review):

      Summary:

      The authors identify KIN-G as an in vitro substrate for phosphorylation by TbPLK and show that several of the in vitro P-ated sites, including T310, overlap with P-ation sites seen in live cells. The authors further show that PLK-mediated P-ation inhibits KIN-G binding to microtubules in vitro, as does a KIN-G-T301D mutant, and that expression of a KIN-G-T301D Phospho-mimic in T. brucei phenocopies KIN-G RNAi knockdowns, producing defects in cell division, morphogenesis of the centrin arm, FAZ and other cellular structures, as well as misplaced cytokinesis furrow.

      Understanding cytoskeletal rearrangements that drive cell division in T. brucei is an important and unresolved problem, so the work addresses important questions that are of great interest. PLK and KIN-G have previously been shown to be important for cell division and morphogenesis of cytoskeletal structures that drive cell division in T. brucei. The current work advances our understanding by suggesting a potential mechanism by which PLK and KIN-G might participate, namely through PLK-dependent P-ation to control KIN-G MT binding activity.

      Strengths:

      The authors use a rigorous combination of biochemistry, phosphoproteomics, cell biology, and mutant analysis to support their conclusion that PLK-mediated P-ation of KIN-G negatively regulates KIN-G microtubule binding and this may explain the observation that a KIN-G T301 phosphomimic mutant blocks cell division and perturbs biogenesis of cytoskeletal structures that drive cell division and morphogenesis. Combining rigorous and informative in vitro studies with mutant analysis in live cells is a great strength. The work is solid and important, though a few pieces are needed to fully connect the in vitro findings with the in vivo observations.

    4. Reviewer #1 (Public review):

      [Editors' note: this version has been assessed by the Reviewing Editor without further input from the original reviewers.]

      I think the main points raised in the review have now been addressed. In particular, the new experiment with TbPLK inhibition and mass spectrometry is an important addition, as it provides direct evidence that phosphorylation of KIN-G at Thr301 and Ser569 depends on TbPLK activity in cells.

      I also appreciate that the authors have toned down the interpretation of the Golgi phenotype. The revised text now makes clear that the fluorescence data show altered Golgi/ERES organization or duplication, but do not prove a structural defect in Golgi biogenesis.

      The added discussion of the T301A result is also helpful. The finding that only a small fraction of KIN-G is phosphorylated at Thr301 in asynchronous cells makes the lack of a strong T301A phenotype more understandable.

      Overall, I am happy with the revision of the beautiful manuscript.

    5. eLife Assessment

      This important study provides new insights into the regulation of cell organization and division in Trypanosoma brucei through the phosphorylation-dependent control of a kinesin motor protein by a polo-like kinase. The authors present convincing evidence, combining rigorous biochemical, cell biological, and imaging analyses, demonstrating that phosphorylation modulates kinesin localization and function, thereby influencing cellular organization and cytokinesis. The findings advance our understanding of the molecular mechanisms governing trypanosome cell division and will be of broad interest to researchers studying trypanosomes, cytoskeletal regulation, and eukaryotic cell division.

    1. eLife Assessment

      This study provides an important contribution to retinal regeneration research by using overexpression of pro-neural factors to reprogram fetal human RPE cells into retinal neurons. The authors provide solid evidence of fetal RPE reprogramming into neural and photoreceptor-like states using scRNA-seq and imaging validation; however, there are concerns regarding comparisons between the effectiveness of different combinatorial transcription factor codes.

    2. Reviewer #1 (Public review):

      Summary:

      In this study, the authors identified transcription factor combinations capable of inducing retinal neuronal programs in cultured fetal human retinal pigment epithelial (RPE) cells. Using a pooled screening strategy, single-cell RNA sequencing, lineage barcoding, and immunohistochemical analyses, they identified ASCL1 and NEUROD1 as an effective combination for inducing retinal neuron-associated transcriptional states. This work aims to advance the development of therapeutic approaches for retinal regeneration by exploring the plasticity of RPE cells.

      Strengths:

      A major strength of the study is the comprehensive experimental design. The combination of transcription factor screening, lineage tracing, single-cell transcriptomics, and molecular validation provides a detailed characterization of the cellular responses to reprogramming factor expression.

      Weaknesses:

      All experiments were performed using fetal human RPE cells. Because fetal RPE remains relatively immature and retains proliferative capacity, it remains unclear to what extent the observed responses reflect true reprogramming of differentiated RPE cells versus activation of developmental plasticity already present in fetal tissue. The absence of adult human RPE controls limits assessment of the generality and translational relevance of the findings.

    3. Reviewer #2 (Public review):

      Summary:

      This is an interesting study that explores how human RPE could be used as a source for new retinal neurons. This is a welcome addition to the field of retinal regeneration, which is currently focused almost exclusively on the regenerative capacity of Müller glia cells. The line of inquiry is firmly rooted in findings from amphibian and embryonic chick model systems and advances a fetal human retina RPE-based screening system as a rich resource for insights into human RPE biology, including as a potential stem cell source.

      The authors investigate the potential of fetal human RPE cells to be reprogrammed into retinal neurons using overexpression of pro-neural factors. While this is a critical knowledge gap in the field of retinal regeneration with significant promise for developing regenerative therapies, several methodological concerns impact the interpretation of results. Firstly, while the authors sought to evaluate factors that enhance RPE reprogramming when co-expressed with ASCL1, nearly all co-expression constructs tested failed to achieve appreciable expression of ASCL1, leaving a central hypothesis of this study largely untested (Major concern 1). Second, although the authors were able to detect a cluster of photoreceptor-like cells in their screen, they were unable to identify which reprogramming construct generated this cluster (Major concern 2). Finally, an essential control that definitively demonstrates the value of combinatorial transcription factor reprogramming is missing (Major concern 3).

      In summary, the authors establish a valuable new paradigm for culturing and reprogramming fetal human RPE, and even more importantly, demonstrate successful reprogramming to neural fates. However, the discussion and interpretation of results needs to be modified significantly to make it clear that (i) the outcome of many co-expression paradigms remains effectively unknown/untested due to failed over-expression of ASCL1, and that (ii) the reprogramming construct giving rise to photoreceptor-like cells could not be conclusively identified from their initial screen.

      Strengths:

      (1) Powerful new screening system advanced for exploring the regenerative potential of human fetal RPE cells.

      (2) Co-expression vector system for testing additive effects of proneural transcription factors.

      Weaknesses:

      Major concerns:

      (1) The authors executed a screen for combinations of factors that can enhance ASCL1-mediated reprogramming of RPE into retinal neurons. However, the expression level of ASCL1 was remarkably low in virtually all co-expression paradigms (see Figure 3C). Notably, the reprogramming combination with the highest potency (ASCL1 + NEUROD1) was also the one exhibiting the highest level of ASCL1 expression. The "failed" reprogramming of most of the co-expression constructs (ASCL1+LMO1, ASCL1+EZH2, and ASCL1+RAX2) is potentially a false negative resulting from low transgenic expression of ASCL1.

      (2) The authors' interpretation is that the overexpression of NeuroD1 and Ascl1 generated a new cluster that expressed markers of photoreceptors such as RXRG and RCVRN (see Figure 3D). However, there does not actually appear to be any overlap between the ASCL1+NEUROD1 cluster (orange dots, left panel) and the cells expressing markers of photoreceptors (yellow/green/purple?/black? dots, right panel; yellow being ASCL1-EZH2, green being ASCL1, purple being ASCL1-and black being control - though color coding here is admittedly somewhat confusing). Thus, the photoreceptor-like cluster of interest actually seems to correspond to gray cells that were unmapped/exposed to an unknown programming cocktail. So, it remains completely unknown which reprogramming construct generated this cluster.

      (3) To conclusively establish the additive role of NEUROD1 in reprogramming, it would be prudent to compare ASCL1 + FA directly to ASCL1+NeuroD1+FA. This control was not included but is needed for a more complete interpretation of results.

    1. eLife Assessment

      This study addresses a valuable question with implications for the development of EEG-based neurofeedback interventions for pain. However, the strength of evidence is incomplete because the principal conclusions rely on post hoc responder analyses and methodological ambiguities that weaken the support for the claimed causal relationship between gamma modulation and pain reduction.

    2. Reviewer #1 (Public review):

      Summary:

      The authors investigate whether EEG neurofeedback (NFB) can be used to increase spontaneous parieto-occipital gamma oscillations and thereby reduce experimentally induced pain. Healthy participants were randomly assigned to active or sham neurofeedback and completed three consecutive neurofeedback blocks with concurrent EEG measurements and phasic painful stimulation. The study addresses a relevant question regarding the causal role of spontaneous gamma oscillations in pain perception and the potential of neurofeedback as a non-pharmacological pain intervention. While the reported findings appear consistent with an association between increased gamma power and reduced pain in a subset of participants, the current analyses do not provide sufficient support for the strong causal conclusions drawn by the authors.

      Strengths:

      (1) The study addresses an important and timely research question with potential implications for EEG-based neurofeedback approaches to pain modulation.

      (2) The sample size is relatively large for an experimental EEG neurofeedback study and includes a sham-control condition.

      (3) The manuscript is generally well written and clearly organized.

      (3) The authors address an important methodological concern regarding EMG contamination of gamma-band activity by including additional EMG recordings in a subset of participants.

      Weaknesses:

      (1) The manuscript frequently presents the relationship between spontaneous gamma oscillations and pain perception as established fact. Given the continuing debate regarding the functional significance of EEG gamma oscillations in pain processing, these statements should be moderated.

      (2) The responder analysis is the most serious methodological concern. Participants in the active group were retrospectively classified as "responders" based on increased gamma power after neurofeedback, and only these participants appear to have been included in the primary analyses and matched to sham participants. As only 23 of 44 participants (52%) met this criterion, the responder rate alone does not demonstrate successful neurofeedback-induced gamma modulation. More importantly, selecting participants based on the outcome variable and subsequently testing that same outcome constitutes circular analysis (double dipping), invalidating the statistical inference. Consequently, the reported effects should be interpreted as an association within a post hoc selected subgroup rather than evidence that neurofeedback increased gamma activity and reduced pain.

      (3) The criterion for successful neurofeedback-induced gamma modulation was not prespecified. It is therefore unclear whether successful modulation was defined by the responder classification, the main effect of session, the group × session interaction, or one of the post hoc comparisons.

      (4) Several methodological details reduce the reproducibility and replicability of the study. The spectral analysis does not clearly describe how trial-wise power estimates were aggregated within participants before group-level analyses, and the preprocessing pipeline includes manual ICA-based artifact rejection without specifying the criteria used for component selection. In addition, the analysis pipeline and custom neurofeedback software should be made publicly available to enable independent reproduction and verification of the reported findings.

      (5) The neurofeedback implementation also raises questions. Updating the feedback only once per second using a 2-s sliding window results in discontinuous visual feedback that may reduce feedback quality and could introduce visually evoked activity. In addition, the viewing distance of approximately 30 cm likely required substantial eye movements while following the moving feedback object.

      (6) The muscle-confound analysis is insufficiently documented. EMG recordings were acquired only in the second cohort, but the manuscript does not clearly state how many participants contributed to this analysis or whether responder selection was performed before or after restricting the sample. These details should be explicitly reported.

    3. Reviewer #2 (Public review):

      Summary:

      The authors investigated whether neurofeedback (NFB) training targeting spontaneous gamma oscillations (30-60 Hz) at the parieto-occipital region (Pz electrode) could reduce experimental pain perception. They randomized 88 healthy participants to active or sham NFB groups across two cohorts (44 each). Active NFB consisted of real-time feedback based on participants' own gamma power; sham NFB consisted of the preceding participant's gamma power. Participants completed three ~16-min sessions, and approximately 52% of active NFB participants showed increased gamma power in session 3 and were considered responders. Analyses restricted to these 23 responders (matched with 23 sham controls) showed reduced pain intensity, unpleasantness, and laser-evoked potential (LEP) amplitudes, with a significant negative correlation between gamma power and pain intensity after session 3.

      Strengths:

      (1) The distinction between spontaneous and stimulus-evoked gamma oscillations in pain processing is theoretically important.

      (2) The rationale for targeting spontaneous gamma via NFB is clearly articulated.

      (3) The study was sham-controlled, and the blinding was adequate.

      (4) The authors commendably ran a second cohort (n=44) with simultaneous posterior neck EMG recording to address the critical concern of muscle artifact contamination of gamma, in response to a previous review

      Weaknesses:

      (1) The most critical issue is about the exclusion of non-responders from the analysis. I find this problematic, as the reasoning becomes circular (selecting the participants who managed to increased gamma and then asking whether gamma NFB influenced pain), effect sizes are inflated, and the selection itself may introduce biases. For example, the responders may differ from the non-responders with respect to other characteristics (better attention skills, better self-regulation, etc). It would be more principled to present the results for the entire sample and only present the responder analysis as a secondary analysis. In the preregistration, the responder-only analysis was not mentioned.

      (2) Another critical point is about the causal claims made in the abstract, introduction, and discussion. Given that the current results provide only correlational evidence in a subsample, the language should be revised to avoid overinterpretation. If the authors can demonstrate a significant mediation effect (NFB group -> gamma change -> pain change), they may be able to argue that increases in gamma activity mediate the observed reduction in pain.

      Minor points:

      (1) For the sham procedure, the authors used the preceding participant's gamma data for feedback. This raises two questions: How was this handled for the first participant? Did the authors check the discrepancy between actual gamma and presented gamma in the sham NFB group?

      (2) Was baseline gamma power comparable between groups?

    4. Reviewer #3 (Public review):

      Summary:

      The authors aimed to test whether spontaneous gamma-band oscillations over the parieto-occipital region can be volitionally upregulated using EEG neurofeedback, and whether this upregulation reduces subsequent pain perception and nociceptive-evoked brain responses. Gamma-band activity has been repeatedly associated with pain processing, but most available evidence remains correlational, and previous attempts to modulate pain-related gamma activity using non-invasive stimulation have not produced robust analgesic effects. The present study therefore addresses an important question: whether real-time neurofeedback may provide a more effective way to train endogenous gamma activity and thereby influence pain.

      Strengths:

      A major strength of the study is the use of an active/sham neurofeedback design. The authors also combine subjective pain ratings with laser-evoked potentials, which provides converging behavioural and neurophysiological outcome measures. The manuscript is clearly written overall, and the study addresses a question of broad interest for pain neuroscience and neurofeedback research.

      Weaknesses:

      A number of aspects limit the strength of the conclusions. The first and most important issue concerns the interpretation of scalp gamma-band activity. Gamma-band oscillations recorded with scalp EEG are difficult to measure reliably, are not observable in all participants, and can be strongly affected by muscle activity. The authors acknowledge this issue and include posterior neck EMG, but the control remains limited. A lack of correlation between one posterior neck EMG channel and Pz gamma power is not sufficient to exclude muscle contamination, especially because gamma-band artifacts can arise from multiple muscle groups and may not be well captured by a single EMG channel. This is particularly important because changes in posture, facial tension, breathing, and arousal could all influence high-frequency scalp activity.

      Second, the evidence for a causal relationship between parieto-occipital gamma activity and pain perception should be interpreted cautiously. The authors show that gamma power increased in approximately half of the active neurofeedback participants and that these responders showed reduced pain ratings and laser-evoked potentials. However, because the main analgesic effect is tied to responder classification, it remains difficult to separate the specific effect of gamma upregulation from broader individual differences in task engagement, suggestibility, relaxation ability, attentional state, or neurofeedback learning capacity.

      A third limitation concerns the control condition and blinding. Participants were reportedly blinded to group allocation, and the credibility ratings appear similar between groups, which is reassuring. However, it is not clear whether the experimenters were also blinded during data collection and interaction with participants. This matters because neurofeedback studies are particularly vulnerable to expectancy.

      The choice of the two neurofeedback scenarios requires clearer justification. The manuscript describes a deep ocean scene followed by a seaside scene with relaxation instructions, but it is not clear why these two scenarios were selected, and whether they were matched for attentional engagement and affective content. This is not a minor point, because both groups showed reductions in pain ratings after the entire neurofeedback procedure.

      The comparison with tACS is interesting but currently underdeveloped. The authors suggest that neurofeedback may succeed where gamma-frequency tACS failed because it allows real-time, personalized, self-regulatory modulation of ongoing activity. This is plausible, but the manuscript should discuss this distinction more deeply. Neurofeedback may not simply be a different way of modulating gamma; it may recruit volitional control, attentional engagement, immersion, expectation, etc. These mechanisms could be central to the observed pain reduction and may partly explain why neurofeedback effects differ from those of externally applied stimulation.

      Overall, this is an interesting study that introduces a promising neurofeedback approach for experimental pain modulation. The findings are encouraging, especially the convergence between subjective ratings and laser-evoked potentials in responders. However, the conclusions should be tempered. The current evidence supports the feasibility of training gamma-band activity in a subset of participants and suggests that successful training is associated with reduced experimental pain.

    5. Author response:

      Reviewer #1 (Public review):

      R1-Q1: The manuscript frequently presents the relationship between spontaneous gamma oscillations and pain perception as established fact. Given the continuing debate regarding the functional significance of EEG gamma oscillations in pain processing, these statements should be moderated.

      We thank the reviewer for raising this important and thoughtful point. We agree that the relationship between spontaneous gamma oscillations and pain perception remains a matter of active debate, and we will moderate these statements throughout the manuscript. We will acknowledge the ongoing debate and present the gamma-pain relationship as an active area of investigation rather than settled fact.

      R1-Q2: The responder analysis is the most serious methodological concern. Participants in the active group were retrospectively classified as 'responders' based on increased gamma power after neurofeedback, and only these participants appear to have been included in the primary analyses and matched to sham participants. As only 23 of 44 participants (52%) met this criterion, the responder rate alone does not demonstrate successful neurofeedback-induced gamma modulation. More importantly, selecting participants based on the outcome variable and subsequently testing that same outcome constitutes circular analysis (double dipping), invalidating the statistical inference. Consequently, the reported effects should be interpreted as an association within a post hoc selected subgroup rather than evidence that neurofeedback increased gamma activity and reduced pain.

      We appreciate this careful critique. We wish to clarify the rationale behind our analytical approach and address the concern.

      A well-established finding in the neurofeedback literature is that a substantial proportion of participants are "non-learners" — individuals who, despite receiving real feedback, fail to achieve effective control over the targeted neural activity. This is not a failure of the intervention, but reflects individual differences in neurofeedback learning capacity. Our core research question is therefore: "Among individuals who can successfully learn to upregulate gamma oscillations, does this upregulation reduce pain perception and nociceptive brain responses?"

      To address the circularity concern and improve transparency, we will make the following revisions:

      - We will reframe the wording from "NFB increases gamma and reduces pain" to "Successful gamma upregulation via NFB is associated with reduced pain in those who achieve it." All causal language will be replaced with appropriately cautious, correlation-based terminology.

      - We will report full-sample results for completeness.

      R1-Q3: The criterion for successful neurofeedback-induced gamma modulation was not prespecified. It is therefore unclear whether successful modulation was defined by the responder classification, the main effect of session, the group × session interaction, or one of the post hoc comparisons.

      We thank the reviewer for requesting this clarification. We will specify the exact criterion in the revised manuscript, i.e., a participant was classified as a responder if their post-intervention gamma power minus pre-intervention gamma power was positive (i.e., an increase in gamma power following the neurofeedback intervention).

      R1-Q4: Several methodological details reduce the reproducibility and replicability of the study. The spectral analysis does not clearly describe how trial-wise power estimates were aggregated within participants before group-level analyses, and the preprocessing pipeline includes manual ICA-based artifact rejection without specifying the criteria used for component selection. In addition, the analysis pipeline and custom neurofeedback software should be made publicly available to enable independent reproduction and verification of the reported findings.

      We thank the reviewer for these constructive suggestions. We will supplement and refine the methodological details in the revised manuscript, and we will make the analysis code and the experimental program (including the custom neurofeedback software) publicly available via an open repository.

      R1-Q5: Updating the feedback only once per second using a 2-s sliding window results in discontinuous visual feedback that may reduce feedback quality and could introduce visually evoked activity. In addition, the viewing distance of approximately 30 cm likely required substantial eye movements while following the moving feedback object.

      We will discuss the limitations of the discontinuous visual feedback and the viewing distance in the revised manuscript. We acknowledge these as valid methodological concerns and will address them as limitations in the Discussion.

      R1-Q6: The muscle-confound analysis is insufficiently documented. EMG recordings were acquired only in the second cohort, but the manuscript does not clearly state how many participants contributed to this analysis or whether responder selection was performed before or after restricting the sample. These details should be explicitly reported.

      We thank the reviewer for pointing out that the description of the muscle-confound analysis was insufficiently detailed. In the revised manuscript, we will clarify the EMG analysis procedures and explicitly report: (a) the exact number of participants contributing to the EMG analysis; (b) the cohort from which they were drawn; and (c) whether responder selection was performed before or after restricting the sample for EMG analysis.

      Reviewer #2 (Public review):

      R2-Q1: The most critical issue is about the exclusion of non-responders from the analysis. I find this problematic, as the reasoning becomes circular (selecting the participants who managed to increase gamma and then asking whether gamma NFB influenced pain), effect sizes are inflated, and the selection itself may introduce biases. For example, the responders may differ from the non-responders with respect to other characteristics (better attention skills, better self-regulation, etc). It would be more principled to present the results for the entire sample and only present the responder analysis as a secondary analysis. In the preregistration, the responder-only analysis was not mentioned.

      As detailed in our response to R1-Q2, the responder analysis reflects a conceptually motivated subgroup defined by successful neurofeedback learning — a well-documented challenge in NFB research where many participants are non-learners. Our central question is whether successful gamma upregulation (among those capable of achieving it) is associated with pain reduction. We will make this rationale explicit in the revised manuscript. We will also: (a) transparently report full-sample results; (b) discuss potential biases introduced by subgroup selection (e.g., differences in attention, self-regulation); and (c) acknowledge the lack of preregistration for the responder analysis.

      R2-Q2: Another critical point is about the causal claims made in the abstract, introduction, and discussion. Given that the current results provide only correlational evidence in a subsample, the language should be revised to avoid overinterpretation. If the authors can demonstrate a significant mediation effect (NFB group → gamma change → pain change), they may be able to argue that increases in gamma activity mediate the observed reduction in pain.

      We will substantially revise the language throughout the manuscript to avoid causal claims. We also plan to conduct a formal mediation analysis (NFB group → gamma change → pain change) to test whether changes in gamma activity statistically mediate the observed pain reduction.

      R2-Q3: For the sham procedure, the authors used the preceding participant's gamma data for feedback. This raises two questions: How was this handled for the first participant? Did the authors check the discrepancy between actual gamma and presented gamma in the sham NFB group?

      We thank the reviewer for raising this point, and we will clarify both points in the revised manuscript. (a) Because group assignment was randomized, the first participant could in principle have been assigned to the sham group. To prepare for this possibility, we collected EEG data from one participant in advance (equivalent to pilot data) to serve as the sham feedback signal, had the first participant been assigned to the sham group. In the actual experiment, however, the first participant was randomly assigned to the active group, so this pre-collected dataset was never used. (b) We will also compare the discrepancy between actual gamma power and the sham feedback signal in the sham group, and report this result in the revised manuscript.

      R2-Q4: Was baseline gamma power comparable between groups?

      We thank the reviewer for this suggestion. We will report and compare baseline gamma power between the active and sham groups in the revised manuscript.

      Reviewer #3 (Public review):

      R3-Q1: Gamma-band oscillations recorded with scalp EEG are difficult to measure reliably, are not observable in all participants, and can be strongly affected by muscle activity. The authors acknowledge this issue and include posterior neck EMG, but the control remains limited. A lack of correlation between one posterior neck EMG channel and Pz gamma power is not sufficient to exclude muscle contamination, especially because gamma-band artifacts can arise from multiple muscle groups and may not be well captured by a single EMG channel. This is particularly important because changes in posture, facial tension, breathing, and arousal could all influence high-frequency scalp activity.

      We thank the reviewer for this important suggestion. We will revise the manuscript to discuss more explicitly the inherent difficulty of recording pure gamma-band oscillations with scalp EEG. We will acknowledge that scalp gamma is not reliably observable in all participants, is vulnerable to contamination from multiple muscle sources, and that a single posterior neck EMG channel provides only limited control. We will also discuss the possibility that changes in posture, facial tension, breathing, and arousal may contribute to high-frequency scalp activity.

      R3-Q2: The evidence for a causal relationship between parieto-occipital gamma activity and pain perception should be interpreted cautiously. The authors show that gamma power increased in approximately half of the active neurofeedback participants and that these responders showed reduced pain ratings and laser-evoked potentials. However, because the main analgesic effect is tied to responder classification, it remains difficult to separate the specific effect of gamma upregulation from broader individual differences in task engagement, suggestibility, relaxation ability, attentional state, or neurofeedback learning capacity.

      We appreciate the reviewer's careful consideration of this point. We will temper our conclusions, presenting the current evidence as demonstrating the feasibility of gamma-band neurofeedback training in a subset of participants and an association between successful training and pain reduction, rather than a demonstrated causal relationship. We will discuss individual differences (attention, suggestibility, relaxation ability, neurofeedback learning capacity) as potential confounds that cannot be fully disentangled from gamma-specific effects.

      R3-Q3: It is not clear whether the experimenters were also blinded during data collection and interaction with participants. This matters because neurofeedback studies are particularly vulnerable to expectancy.

      We will clarify that the study employed a single-blind design: participants were unaware of their group assignment. We will state this clearly in the revised manuscript.

      R3-Q4: The choice of the two neurofeedback scenarios requires clearer justification. The manuscript describes a deep ocean scene followed by a seaside scene with relaxation instructions, but it is not clear why these two scenarios were selected, and whether they were matched for attentional engagement and affective content. This is not a minor point, because both groups showed reductions in pain ratings after the entire neurofeedback procedure.

      We will provide a stronger rationale for the selection of the two neurofeedback video scenarios. We will also place greater emphasis on the pain reduction observed in both groups, acknowledging the substantial nonspecific analgesic effects associated with the procedure.

      R3-Q5: The comparison with tACS is interesting but currently underdeveloped. The authors suggest that neurofeedback may succeed where gamma-frequency tACS failed because it allows real-time, personalized, self-regulatory modulation of ongoing activity. This is plausible, but the manuscript should discuss this distinction more deeply. Neurofeedback may not simply be a different way of modulating gamma; it may recruit volitional control, attentional engagement, immersion, expectation, etc. These mechanisms could be central to the observed pain reduction and may partly explain why neurofeedback effects differ from those of externally applied stimulation.

      We thank the reviewer for this insightful comment. We will expand the discussion of why neurofeedback may produce effects beyond those achieved by gamma-frequency tACS. In particular, we will elaborate on the potential contributions of volitional control, attentional engagement, immersion, expectation, and self-regulatory processes, and discuss how these factors may be central to the observed pain reduction rather than merely incidental to the gamma modulation.

      R3-Q6: Overall, this is an interesting study that introduces a promising neurofeedback approach for experimental pain modulation. The findings are encouraging, especially the convergence between subjective ratings and laser-evoked potentials in responders. However, the conclusions should be tempered. The current evidence supports the feasibility of training gamma-band activity in a subset of participants and suggests that successful training is associated with reduced experimental pain.

      We thank the reviewer for this balanced assessment. We fully agree that the conclusions should be tempered, and we will revise the manuscript accordingly to reflect that the current evidence supports feasibility and association rather than established causal efficacy.

      Summary

      In summary, the planned revisions include:

      (i) full-sample results reported;

      (ii) moderating causal language throughout and adding a formal mediation analysis;

      (iii) clearly specifying the responder criterion and adding this to the preregistration;

      (iv) providing complete methodological documentation and publicly releasing all analysis code;

      (v) expanding the Discussion to address limitations regarding scalp gamma measurement, EMG control, visual feedback, viewing distance, single-blind design, NFB scenario rationale, and nonspecific effects;

      (vi) adding analyses on baseline gamma comparability and sham-feedback discrepancy.

    1. eLife Assessment

      This study provides important evidence that crickets use bat-associated odors as cues of predation risk, extending the classic bat-insect arms race beyond its traditional acoustic framework. The authors combine fecal metabarcoding, behavioral assays, electrophysiology, chemical analyses, and field observations to show that Loxoblemmus equestris avoids the odor of the insectivorous bat Scotophilus kuhlii, and that synthetic limonene can elicit antennal responses, avoidance in the laboratory, and reduced calling activity in the field. Together, the multidisciplinary results provide a coherent and convincing chain of evidence for an olfactory component of predator detection in this bat-insect system. The work will be of broad interest to researchers in sensory ecology, chemical ecology, predator-prey interactions, and bat-insect coevolution.

    2. Reviewer #1 (Public review):

      [Editors' note: this version has been assessed by the Reviewing Editor without further input from the original reviewers. The authors have addressed the comments raised in the previous round of review.]

      The manuscript examines whether insects can use bat odor as a cue of predation risk. The authors focus on the insectivorous bat Scotophilus kuhlii and the cricket Loxoblemmus equestris. They first use fecal DNA metabarcoding to show that crickets are part of the bat's diet, and field surveys to show that L. equestris is abundant at local foraging sites. In laboratory Y-tube assays, the authors show that crickets strongly avoid air carrying bat body odor. Gas chromatography coupled with electroantennographic detection showed that cricket antennae respond to components of bat odor. Chemical analyses identified several volatile compounds, with 2,2-dimethylheptane and (−)-limonene associated with antennal responses. Further analyses suggested that snout secretions are likely to contribute to the bat's body odor. The authors then tested individual compounds. Among the commercially available candidates, (−)-limonene elicited a strong antennal response and was sufficient to cause avoidance in the olfactometer. In field plots, spraying (−)-limonene reduced cricket calling activity relative to pre-exposure levels, whereas calling increased in control plots treated with hexane. Overall, the study argues that crickets can detect a vertebrate predator through olfactory cues and that a single bat-associated volatile can trigger antipredator behavior.

      This is an interesting and enjoyable study that addresses an understudied aspect of predator-prey interactions. The manuscript is clearly written, the experiments are presented in a logical sequence, and the figures are crisp and easy to follow. I really appreciated the combination of behavioral assays, electrophysiology, chemical analysis, and field observations.

    3. Reviewer #2 (Public review):

      Many insects possess extremely sensitive olfactory systems that can detect chemical signals from distances of several kilometers. For decades, the arms race between bats and insects has served as a prime example of acoustic co-evolution. The auditory adaptations of insects to echolocation have been well documented. Cricket has a multi-sensory predator recognition system with keen olfactory, tactile, and auditory senses. However, whether crickets can use the scent of bats to avoid them remains unknown at present. The authors hypothesized that cricket prey (Loxoblemmus equestris) might eavesdrop on predator bat (Scotophilus kuhlii) VOCs as an early warning. L. equestris is one of the prey species of S. kuhlii, and the authors demonstrated that the body odor of the insectivorous bat S. kuhlii triggers robust avoidance and electrophysiological responses in the cricket L. equestris, and that a single compound, (-)-limonene, is sufficient to elicit this avoidance in the laboratory and suppress calling in the field. Overall, this paper has a complete chain of evidence and should be a highly praised study.

    4. Author response:

      The following is the authors’ response to the original reviews.

      We sincerely thank you and the reviewers for the thoughtful evaluation of our manuscript and for the constructive comments and important suggestions. We are encouraged by the recognition that the study is valuable and of interest to researchers in sensory ecology, chemical ecology, predator-prey interactions, and bat-insect coevolution. We are also grateful that the reviewers acknowledged the integrative approach of our work, including fecal metabarcoding, behavioral assays, electrophysiological recordings, chemical analyses, and field observations.

      We have carefully considered all comments and have revised the manuscript accordingly. In particular, we have made the following major revisions:

      (1) We clarified the biological origin of limonene and expanded the discussion of possible bat-associated sources, including snout secretions and microbial contributions, while avoiding overinterpretation of limonene as an exclusively endogenous mammalian compound.

      (2) We added more detailed descriptions of contamination controls, including instrument cleaning procedures, materials used for housing and handling, and blank-control results, to address the concern that limonene could have originated from human-associated or environmental contamination.

      (3) We added individual-level odor data and revised the presentation of terpenoid profiles to show more clearly where limonene was detected and how its relative contribution varied among samples.

      (4) We clarified the rationale for the concentration choices used in the electrophysiological and field experiments, emphasizing that these assays were designed to test physiological detectability and functional sufficiency rather than to reproduce exact natural emission concentrations or determine response thresholds.

      (5) We revised our interpretation of limonene more cautiously. We now state that limonene is sufficient to trigger avoidance responses, while we acknowledge that its natural ecological specificity, concentration dynamics, and interactions with other bat odor components require further investigation.

      (6) We corrected the terminology related to limonene enantiomers. Because our GC-MS and GC-EAD analyses did not use an enantioselective column, we replaced “(-)-limonene” with “limonene” throughout the manuscript, figures, legends, and supplementary materials.

      (7) We improved the figures and supporting materials by moving the figure illustrating predator-prey relationship into the main text, adding electrophysiological traces from all tested crickets, revising figure legends, clarifying the rationale for comparing bat body odor with air controls, and providing additional chemical-identification details in the supplementary materials.

      (8) We checked and clarified statistical annotations, including the exact adjusted P-values for relevant comparisons.

      We believe these revisions have substantially improved the clarity, rigor, and balance of the manuscript. We hope that the revised manuscript and the detailed point-by-point responses satisfactorily address all concerns raised.

      We are confident that our study represents an important contribution, as it fundamentally expands the traditional acoustic-centred view of bat–insect interactions by demonstrating that crickets can use olfaction as a complementary sensory modality to detect bat odors and initiate avoidance behavior. The multidisciplinary evidence, integrating behavioral assays, electrophysiology, chemical profiling, and field validation, provides convincing support for this novel olfactory pathway.

      eLife Assessment

      This valuable study raises the intriguing possibility that crickets use bat-associated odors as cues of predation risk, extending the classic bat-insect arms race beyond its usual acoustic framework. The authors combine fecal metabarcoding, behavioral assays, electrophysiology, chemical analyses, and field observations to show that Loxoblemmus equestris avoids the odor of the insectivorous bat Scotophilus kuhlii, and that synthetic (-)-limonene can elicit antennal responses, avoidance in the laboratory, and reduced calling activity in the field. However, the evidence is currently incomplete because the identity, biological source, natural concentration, and ecological specificity of limonene as a bat-derived predator cue require stronger support, including clearer quantification, contamination controls, individual-level odor data, and evidence that crickets can distinguish bat-associated limonene from common environmental sources. The work will be of interest to researchers in sensory ecology, chemical ecology, predator-prey interactions, and bat-insect coevolution.

      We sincerely thank the editors for the positive and constructive assessment of our work. We greatly appreciate the recognition of our study’s value and its potential interest to multiple research communities.

      We have carefully considered all comments and have revised the manuscript accordingly. The revisions include clarifications of limonene’s biological origin and ecological specificity, strengthened contamination controls and individual-level odor data, clearer rationales for experimental concentration choices, and more cautious interpretation throughout the manuscript. All changes are addressed in detail in our point-by-point responses below.

      Importantly, the central conclusion of our study—that crickets can detect and avoid bat odor through olfaction, and that limonene is sufficient to trigger avoidance responses in both laboratory and field settings—remains robustly supported by the multidisciplinary evidence we present. We believe the revised manuscript now provides a clearer, more balanced, and scientifically rigorous account of our findings, and we hope it meets the standards of eLife.

      Public Reviews:

      Reviewer #1 (Public review):

      The manuscript examines whether insects can use bat odor as a cue of predation risk. The authors focus on the insectivorous bat Scotophilus kuhlii and the cricket Loxoblemmus equestris. They first use fecal DNA metabarcoding to show that crickets are part of the bat's diet, and field surveys to show that L. equestris is abundant at local foraging sites. In laboratory Y-tube assays, the authors show that crickets strongly avoid air carrying bat body odor. Gas chromatography coupled with electroantennographic detection showed that cricket antennae respond to components of bat odor. Chemical analyses identified several volatile compounds, with 2,2-dimethylheptane and (−)-limonene associated with antennal responses. Further analyses suggested that snout secretions are likely to contribute to the bat's body odor. The authors then tested individual compounds. Among the commercially available candidates, (−)-limonene elicited a strong antennal response and was sufficient to cause avoidance in the olfactometer. In field plots, spraying (−)-limonene reduced cricket calling activity relative to pre-exposure levels, whereas calling increased in control plots treated with hexane. Overall, the study argues that crickets can detect a vertebrate predator through olfactory cues and that a single bat-associated volatile can trigger antipredator behavior.

      This is an interesting and enjoyable study that addresses an understudied aspect of predator-prey interactions. The manuscript is clearly written, the experiments are presented in a logical sequence, and the figures are crisp and easy to follow. I really appreciated the combination of behavioral assays, electrophysiology, chemical analysis, and field observations.

      We sincerely thank you for your very positive and encouraging evaluation of our work. We are delighted that you found the study interesting and enjoyable. We also appreciate your kind remarks about the clarity of the manuscript, the logical flow of the experiments, and the quality of the figures.

      We are especially grateful that you recognized the value of our integrative approach. Combining behavioral assays, electrophysiology, chemical analysis, and field observations was central to our study design, and we are pleased that this approach resonated with you.

      You provided an accurate and comprehensive summary of our work. You confirmed that our main narrative is clear and logically coherent. Specifically, you followed our progression from establishing the predator–prey relationship, to demonstrating olfactory avoidance, to identifying limonene as an active compound, and finally to validating its behavioral effects in both laboratory and field settings.

      We have carefully considered all your constructive comments and suggestions. We address them in detail in our point-by-point responses below. Your feedback has been extremely helpful, and we believe the revised manuscript is substantially stronger as a result.

      My main issue concerns the identity and biological origin of the proposed bat odor cue, (−)-limonene. Limonene seems like an unusual compound to be emitted endogenously by a mammal, particularly by an insectivorous bat. It would be helpful if the authors could clarify whether mammals are known to synthesize this compound de novo, and, if not, what the likely source of this plant-associated terpene would be in S. kuhlii. Possible sources could include environmental exposure, diet, roosting material, handling, or temporary housing conditions.

      I do not doubt that crickets avoid synthetic (−)-limonene. Indeed, this result is quite plausible given that limonene is widely used in insect repellent or repellent-associated fragrance products. However, this also makes contamination an important issue to address explicitly. How did the authors exclude the possibility that limonene entered the samples from human-associated sources, such as insect repellents, soaps, cleaning products, field equipment, cloth bags, cages, gloves, or other materials used while handling wild-caught bats? It would strengthen the manuscript to report limonene levels for individual bat odor collections, all relevant blanks, and any handling or housing controls.

      More broadly, given the common occurrence of limonene in plants and human-associated products, I am not yet convinced that it would function as a reliable "keystone kairomone" as suggested around line 253. How would crickets distinguish bat-associated limonene from limonene emitted by a mint leaf, citrus peel, pine material, or other non-threatening environmental sources? The authors may wish to soften this interpretation or provide additional evidence that crickets respond to limonene in a bat-specific context, perhaps through concentration, temporal patterning, co-occurring volatiles, or enantiomeric composition.

      We sincerely thank you for your critical and constructive comments. Your questions regarding the identity, biological origin, and ecological specificity of limonene are insightful and have helped us substantially strengthen the manuscript. We address each of your points below.

      On the biological origin of limonene and whether mammals synthesize it de novo

      You raised an important point that limonene seems unusual for a mammal to emit endogenously. We fully agree. We agree that direct evidence for de novo synthesis of limonene in mammals is currently limited.

      However, limonene in bat body odor could still have biological origins. First, it may originate from skin- or gland-associated microbiota. Recent work has shown that skin-associated microorganisms can substantially shape bat volatile odor profiles (Sun et al., 2026, BMC Biology), and some microbes possess enzymes capable of terpene biosynthesis. Second, previous studies have independently reported limonene in the secretions of several bat species (Faulkes et al., 2019, PeerJ; Zhang et al., 2022, Ann. N.Y. Acad. Sci.). This suggests that its presence in bats is not unique to our study. Third, our own analyses detected limonene in hair and snout secretions, but not in faeces or blank controls. This pattern is consistent with a biological source associated with the body surface, rather than diet or environmental deposition.

      We have now expanded our Discussion to cover these possibilities more explicitly. We also emphasize that the exact source, i.e., endogenous, microbial, or otherwise, remains an open question that warrants future investigation. Please see lines 258–273 of the clean version of the revised manuscript, or see the excerpt below:

      “Although limonene reliably induced avoidance behaviour in crickets, two related questions still merit careful consideration. One question is whether the limonene we identified genuinely originates from bats or reflects contamination during sampling. Limonene is common in plants and numerous consumer products (Boncan et al., 2020; Schuman, 2023), making its endogenous production by a mammal seem unusual. Nevertheless, multiple lines of evidence militate against contamination. First, we adhered to rigorous protocols. For example, all instruments were cleaned with ethanol and oven-dried before each use; bats were housed in stainless-steel cages, and cloth bags had been rinsed with purified water. Second, limonene was absent from all blank controls, including empty-chamber air samples and clean swabs, and was not detected in bat faecal samples. In contrast, it was consistently identified in hair and snout-secretion samples from bats. Third, independent studies have similarly identified limonene in the secretions of other bat species (Faulkes et al., 2019; Zhang et al., 2022). Furthermore, emerging evidence indicates skin-associated microbes may contribute to bat volatile profiles, with some taxa possessing enzymes involved in terpene biosynthesis (Sun et al., 2026). Taken together, these observations point towards an endogenous or microbe-mediated source, although the exact biosynthetic pathway remains to be determined.”

      On contamination from human-associated sources

      You asked how we excluded the possibility that limonene entered our samples through handling, equipment, cleaning products, or other human-associated sources. We appreciate this concern and have addressed it in detail.

      We believe contamination is highly unlikely for several reasons. First, we followed strict protocols throughout. All instruments were cleaned with ethanol and oven-dried before and after each use. We used stainless-steel cages and cloth bags made of degreased bleached cotton washed with purified water. These materials are not sources of terpenes. Second, we ran multiple blanks. Limonene was not detected in any empty-chamber air controls or in blank cotton swabs. In contrast, it was consistently found in multiple bat snout-secretion samples. This clear difference between samples and blanks strongly argues against contamination. Third, we now report individual-level odor data (see new Figure 3; Supplementary Table 5.xlsx). These data show that limonene was consistently present across individual bats. It did not appear sporadically, as one would expect from accidental contamination.

      In the revised manuscript, we have added detailed descriptions of our contamination controls in the Methods section (Please see lines 405–407, 443–444, 471–477 of the clean version of the revised manuscript, or see the excerpt below). We have also included the blank-control results (Supplementary Table 5.xlsx), as you suggested.

      Lines 405–407: “Prior to sampling, all glassware was thoroughly rinsed with ethanol and dried in an oven at 120°C, and volatile odor collection was conducted in a dedicated odor-free room to minimize environmental contamination.”

      Lines 443–444: “The empty-chamber controls were used to account for potential background signals from the experimental system and to provide a baseline for comparison with bat odor extracts.”

      Lines 471–477: “Upon capture, bats were placed in clean stainless-steel cages and kept in groups consistent with their natural social associations during the brief interval prior to immediate odor sampling. Hair samples (10 mg per individual) were clipped from dorsal and ventral regions. Snout secretions were collected using sterile cotton swabs (CS15-005, Shenzhen SihuaBo Technology Co., Ltd., China), with two blank swabs as controls. These blank swab controls were included to account for potential volatile contamination from ambient air or the swab material itself (Supplementary Table 5).”

      On how crickets distinguish bat-associated limonene from environmental sources

      You raised a thoughtful question about ecological specificity. Given that limonene is abundant in mint, citrus peel, pine, and other non-threatening plants, how would crickets use it as a reliable indicator of bat presence?

      We agree with you completely. We do not claim that limonene alone serves as an unambiguous bat-specific signal. Instead, our interpretation is more nuanced. We argue that elemental perception represents one effective strategy within a broader olfactory toolkit. It does not exclude the importance of other cues.

      In our revised manuscript, we have softened our interpretation accordingly. We now state explicitly that limonene is sufficient to trigger avoidance under our experimental conditions, but we do not interpret it as a uniquely bat-specific keystone kairomone. We also discuss mechanisms that could help crickets reduce false alarms under natural conditions. These include concentration differences, temporal patterning (bats are active at night), spatial context (specific foraging habitats), co-occurrence with other bat-specific volatiles, and possibly enantiomeric composition. Please see lines 274–292 of the clean version of the revised manuscript, or see the excerpt below.

      Lines 274–292: “The second question is how crickets might distinguish bat-derived limonene from environmental sources of this compound, given its prevalence in mint, citrus peel, pine and other non-threatening plants (Boncan et al., 2020; Schuman, 2023). It seems implausible that crickets could simply rely on limonene per se to differentiate a bat from a leaf. Two non-exclusive mechanisms could help resolve this issue. First, limonene need not be the only olfactory cue mediating risk perception. Our findings establish the sufficiency of limonene as an avoidance trigger, but do not preclude a role for other odor components. The crickets’ antennal responses to other bat volatiles in our GC–EAD analyses suggest more complex peripheral perception. Additional compounds, either alone or in synergistic blends, may modulate the full behavioral response in nature. Therefore, elemental perception via limonene likely represents one effective strategy within a broader olfactory toolkit available to insects. Second, crickets may discriminate bat-derived limonene through context-specific cues (e.g., temporal and spatial patterning, co-occurrence with other bat-specific compounds) to minimize false alarms. Comparative studies on enantiomeric specificity and detection thresholds of cricket olfactory sensory neurons will be essential. Equally critical will be future efforts to quantify natural bat odor composition, limonene release rates, ambient exposure concentrations, and odor-plume dynamics, which together will inform ecologically valid stimulus design in controlled assays. Critically, our field data confirm that limonene exposure in nature robustly triggers an adaptive anti-predator response, irrespective of the precise discrimination mechanism.”

      We acknowledge that fully testing these ideas would require substantial additional work. We have therefore framed this as an important direction for future research, rather than as a resolved issue in the present study.

      We thank you again for these insightful comments. Your feedback has helped us present a more rigorous, balanced, and transparent account of our work.

      Reviewer #2 (Public review):

      Summary:

      Many insects possess extremely sensitive olfactory systems that can detect chemical signals from distances of several kilometers. For decades, the arms race between bats and insects has served as a prime example of acoustic co-evolution. The auditory adaptations of insects to echolocation have been well documented. Cricket has a multi-sensory predator recognition system with keen olfactory, tactile, and auditory senses. However, whether crickets can use the scent of bats to avoid them remains unknown at present. The authors hypothesized that cricket prey (Loxoblemmus equestris) might eavesdrop on predator bat (Scotophilus kuhlii) VOCs as an early warning. L. equestris is one of the prey species of S. kuhlii, and the authors demonstrated that the body odor of the insectivorous bat S. kuhlii triggers robust avoidance and electrophysiological responses in the cricket L. equestris, and that a single compound, (-)-limonene, is sufficient to elicit this avoidance in the laboratory and suppress calling in the field. Overall, this paper has a complete chain of evidence and should be a highly praised study.

      We sincerely thank you for your very positive and encouraging evaluation of our work. We are especially gratified that you recognized our study as having a "complete chain of evidence" and as a "highly praised study." This recognition means a great deal to us, given the multidisciplinary nature of our approach and the effort required to integrate behavioral, electrophysiological, chemical, and field data into a coherent narrative.

      We also appreciate your accurate summary of our work. You correctly highlighted the broader context that while acoustic co-evolution between bats and insects is well documented, whether crickets can use bat scent as an early warning cue has remained unknown. Your summary confirms that our main findings are clear: the body odor of S. kuhlii triggers robust avoidance and electrophysiological responses in L. equestris, and that limonene alone is sufficient to elicit avoidance in the laboratory and suppress calling in the field.

      We are particularly grateful that you acknowledged the multi-sensory nature of cricket predator recognition, with keen olfactory, tactile, and auditory senses. We agree that crickets are an excellent model for studying multimodal predator detection, and we hope our study encourages further exploration of olfaction in this classic predator–prey system.

      We have carefully considered all your specific comments and suggestions. These include questions about the novelty framing of olfactory eavesdropping, the rationale for our concentration choices, and the presentation of electrophysiological comparisons. We address each of these points in detail in our point-by-point responses below. Your thoughtful feedback has been extremely helpful in improving the clarity and rigor of the manuscript.

      Comments:

      (1) Olfactory eavesdropping can transcend the evolutionary divide between vertebrate predators and invertebrate prey, enabling invertebrates to trigger defensive avoidance behaviors in response to predator-derived volatile odors. This phenomenon is empirically well-documented and requires no excessive emphasis.

      Thank you for this comment. You are absolutely right that olfactory eavesdropping across the vertebrate–invertebrate divide is not a new concept in itself, and we acknowledge that this phenomenon has been well documented in previous studies.

      However, we would like to clarify our intended emphasis. In the Introduction and Discussion, we have already stated that empirical examples combining chemical identification, electrophysiological validation, behavioral assays, and field confirmation within a direct predator–prey context remain relatively limited. This is especially true for the bat–insect system, where research has historically focused on acoustic interactions rather than olfaction.

      We did not intend to overstate the novelty of olfactory eavesdropping per se. Instead, our emphasis was on providing a complete chain of evidence in a vertebrate–invertebrate predator–prey system that has traditionally been viewed through an acoustic lens. In that sense, we believe our study adds a complementary olfactory perspective to this classic system, rather than claiming to have discovered olfactory eavesdropping as a novel phenomenon. We hope this clarifies our position, as already stated in the original manuscript (lines 71–78, lines 88–90, lines 235–241 of the clean version of the revised manuscript, or see the excerpt below):

      Lines 71–78: “However, a fundamental gap exists in understanding whether such olfactory eavesdropping can operate across the vast phylogenetic divide separating vertebrate predators and invertebrate prey (Apfelbach et al., 2015; Dicke and Grostal, 2001; Schoeppner and Relyea, 2005). Although olfactory interactions across broad taxonomic boundaries are widespread in nature, such as mosquitoes using host odors to blood-feed, elephants and moths sharing pheromonal components, and aroids chemically mimicking carrion to attract pollinating flies (Kang et al., 2023; Zaremska et al., 2022; Zhao et al., 2022), these interactions are primarily shaped by selective pressures tied to foraging, reproduction, or mutualisms, rather than by predation-related selection.”

      Lines 88–90: “The bat–insect system presents an ideal model to address these questions. Despite the clear importance of olfaction to both taxa and its established role in predator–prey ecology, whether it plays any functional role in the iconic bat–insect arms race remains unexplored.”

      Lines 235–241: “Beyond the specific bat–insect model, our work addresses a central question in sensory ecology: how chemical eavesdropping operates within predator–prey systems between phylogenetically distant taxa with fundamentally divergent olfactory systems (Adams et al., 2020; Emerson and Johnson, 2024; Kaupp, 2010). While intraphyletic kairomone detection is well-established (e.g., rodents avoiding carnivore odors, aphids responding to ladybug chemicals), compelling experimental evidence for such olfaction-mediated recognition across broad phylogenetic divides has been limited (Apfelbach et al., 2005; Ferrari et al., 2007; Tanis et al., 2018).”

      (2) Without quantitative analysis and without knowing the relative content of this key substance limonene, I don't quite understand how to determine the concentration of limonene standard for EAD, as well as the concentration in field experiments. How is the concentration of limonene determined in field spraying, and is this actually the case in the wild environment?

      Thank you for this question. We fully agree that knowing the natural concentrations and relative content of limonene in bat odor would be valuable. However, our experimental aims were not to mimic natural emission levels precisely. Instead, they were designed to answer two distinct questions: First, whether cricket antennae are physiologically capable of detecting limonene at all; and second, whether limonene alone is sufficient to trigger behavioral responses under controlled and semi-natural conditions.

      For the EAG experiments, we selected a concentration gradient (0.001%, 0.01%, 0.1%, 1%, and 10% v/v) following standard practices in insect chemical ecology and referencing a previous study (Tang et al., 2024). The goal was to establish dose-dependent antennal sensitivity, not to match a specific natural concentration. Our data clearly show that cricket antennae respond across a range of concentrations, with stronger responses at higher doses.

      For the field experiment, we used a 10% v/v limonene spray over 25 m<sup>2</sup> plots. We acknowledge that this concentration does not quantitatively reflect natural bat emissions. Natural odor plumes are highly dynamic and depend on airflow, turbulence, temperature, humidity, vegetation structure, and distance from the source. Accurately reconstructing these natural dynamics would require detailed quantitative measurements of bat odor release rates and plume modeling, which were beyond the scope of the present study. Instead, our field experiment was designed for a functional purpose: to test whether limonene could alter cricket calling behavior under semi-natural conditions, using a concentration sufficient to produce a detectable odor stimulus in the field.

      We also note that the field-applied concentration is comparable to what has been used in other chemical ecology studies testing the behavioral effects of single volatile compounds under natural or semi-natural conditions. In that context, our positive result supports the ecological relevance of limonene as an avoidance cue, without requiring that the exact applied concentration matches natural bat emissions.

      We agree that quantitative characterization of natural bat odor composition, limonene release rates, and ambient exposure concentrations is an important direction for future research. According to your comments, we have added this point to the revised Discussion as a clear future direction (please see lines 286–290 of the clean version of the revised manuscript, or see the excerpt below). We have also clarified in the Methods that our assays were designed to test physiological detectability and functional sufficiency, rather than to establish concentration thresholds or mimic natural emissions exactly (please see lines 515–522, 556–558 of the clean version of the revised manuscript, or see the excerpt below).

      Lines 286–290: “Comparative studies on enantiomeric specificity and detection thresholds of cricket olfactory sensory neurons will be essential. Equally critical will be future efforts to quantify natural bat odor composition, limonene release rates, ambient exposure concentrations, and odor-plume dynamics, which together will inform ecologically valid stimulus design in controlled assays.”

      Lines 515–522: “For each antenna, a hexane control was first presented to establish baseline antennal activity. For the initial screening, limonene, undecane, pentadecane, and hexadecane were diluted to 10% (v/v) in hexane and delivered individually in a randomized order. To assess dose-dependent responses, limonene was further tested at five concentrations (0.001%, 0.01%, 0.1%, 1%, and 10%, v/v in n-hexane), following the concentration gradient used in a previous study (Tang et al., 2024). Following the initial hexane control, the five limonene concentrations were tested in a randomized order across trials. Each stimulus lasted 0.5 s, with an inter-stimulus interval of 1 min to allow full recovery of antennal responses.”

      Lines 556–558: “Our assays were designed to test physiological detectability and functional sufficiency, rather than to establish concentration thresholds or mimic natural emissions exactly.”

      (3) Figures 1C and D should compare the GC-EAD response of L. equestris to the odor of bat body and the odor of bat nasal secretions. It should not be compared with the air control group. Figure 1D has the same problem.

      Thank you for this suggestion. We understand your point that comparing GC-EAD responses between bat body odor and snout secretions would be a more direct way to identify the anatomical source of active compounds.

      However, we would like to explain why we did not include this comparison in the current study.

      First, the purpose of Figures 1C and 1D (i.e., Figure 2C and 2D in the revised manuscript) was to answer a more fundamental question: whether bat body odor, as a whole, contains volatile compounds that are detectable by cricket antennae. Comparing with an odor-free air control was therefore the appropriate first step. It established the basic phenomenon of olfactory detection before we moved on to source attribution.

      Second, we did conduct chemical profiling of snout secretions, hair, and faeces using HS-SPME-GC-MS (presented in Figure 3). These analyses showed that limonene was consistently present in hair and snout secretions, but absent from faeces and blanks. This allowed us to identify snout secretions as the most likely source of limonene, without requiring GC-EAD recordings from secretion samples themselves.

      Third, we did not perform GC-EAD on snout secretions for practical reasons. The secretion samples were collected in very small amounts. They were almost entirely consumed during the HS-SPME-GC-MS chemical analyses, leaving insufficient material for additional GC-EAD testing.

      We agree with you that directly comparing GC-EAD responses to snout secretions versus whole-body odor would be an excellent experiment. It would further strengthen the source attribution and provide more direct evidence. We have noted this as a valuable direction for future studies in the revised Discussion.

      We hope this clarifies our rationale. Thank you again for your thoughtful suggestion.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) I would suggest moving Supplementary Figure 1 into the main figures. It contains important information about the predator-prey relationship and the ecological relevance of L. equestris, so it seems too central to be placed only in the supplement.

      We agree with your suggestion. The predator–prey relationship and the ecological relevance of L. equestris are indeed central to the biological context of this study. We have therefore moved the original Supplementary Figure 1 into the main text as Figure 1. We have renumbered the remaining figures accordingly, revised the figure legends, and updated the Results text to better highlight this ecological context (revised manuscript, Figure 1 legend, lines 856–864).

      (2) Lines 134 to 136: Only one representative EAD trace is shown. I suggest showing all five traces, either in the main figure or as a supplementary figure, to better illustrate the reproducibility of the antennal responses across individuals.

      We agree. To better illustrate reproducibility across individuals, we have added EAD traces from all five tested crickets as Supplementary Figure 1. The figure legend now describes the sample sizes for both the bat odor treatment and the odor-free control (revised manuscript, Supplementary Figure 1 legend, lines 900–906).

      (3) Lines 144 to 153: It would be helpful if the authors reported which VOC collections contained limonene and in what amounts. Showing individual-level data for the bat odor samples, rather than only pooled or summarized profiles, would strengthen the conclusion that limonene is consistently associated with S. kuhlii body odor.

      We agree. To better show the consistency of limonene detection across individuals, we have added Figure 3B, which displays an individual-level terpenoid profile. Each stacked bar represents one VOC collection, and the limonene segment indicates its presence and relative contribution. We have also revised the Results section to state explicitly that limonene was detected in hair and snout secretions but absent from feces (lines 144–149 in the revised manuscript).

      Lines 144–149: “To identify the biological sources of bat body odor, we analyzed VOCs from hair, faeces, and snout (pararhinal gland) secretions of nine bats using headspace solid–phase microextraction coupled with gas chromatography–mass spectrometry (HS–SPME–GC–MS). Snout secretions and hair shared similar hydrocarbon-rich VOC profiles, whereas faecal volatiles were distinct (Figure 3A). Individual-level terpenoid profiles further showed that limonene was detected in hair and snout secretion VOC collections but was absent from faeces (Figure 3B and Supplementary Table 2).”

      (4) Figure 2A: I recommend adding representative chromatograms or VOC traces for feces, hair, and snout secretions. This would make the source comparison more transparent and would help readers assess the underlying chemical profiles behind the heatmap.

      We agree that representative chromatograms would make the source comparison more transparent. However, due to the analytical workflow, individual chromatograms were not retained in the dataset we received. As an alternative, we added Figure 3B, which shows individual-level terpenoid profiles. This allows readers to assess which samples contained limonene and how its relative contribution varied among sample types and individuals. In addition, we have provided the NIST retention index, quantitative ion, qualitative ion, and molecular weight for each terpenoid compound in Supplementary Table 2.

      (5) The chemical identification of the key compounds would benefit from more detail. The authors state that compound identities were confirmed by matching retention times and mass spectra to authentic standards, including a mixed standard injection. It would be useful to provide the retention times, match and reverse-match scores, blank traces, and, if available, sample-plus-standard co-injection data showing peak augmentation without the appearance of new peaks. For (−)-limonene specifically, the enantiomeric assignment would require an enantioselective method, such as chiral gas chromatography, unless this was already performed and not described.

      We fully agree that comprehensive identification evidence is important for transparency and reproducibility.

      Regarding the identification data: we have already confirmed compound identities by matching retention times and mass spectra to authentic standards, including a mixed standard injection. In the revised version, we will deposit the raw chromatographic data and identification details (retention times, match scores, and blank traces) in Figshare as supporting information.

      Regarding co-injection: we acknowledge that sample-plus-standard co-injection would provide even stronger confirmation. However, our bat odor samples were difficult to obtain and were almost entirely consumed during the GC–EAD and GC–MS analyses. We therefore could not perform additional co-injection validation. We have noted this limitation in the revised Materials and Methods (revised manuscript, lines 463–464).

      Regarding the enantiomer assignment of limonene: you are correct that determining the specific enantiomer requires a chiral GC column, which was not available in this study. To avoid overinterpretation, we have replaced “(-)-limonene” with “limonene” throughout the manuscript.

      Reviewer #2 (Recommendations for the authors):

      (1) This is a typical study in the field of chemical ecology. As long as the source of the active substances is determined and the biological activity has been detected, this research is complete, so Figure 2 seems to be unnecessary.

      We agree that identifying the source of the active compound and confirming its biological activity are central to this study. However, we believe Figure 2 serves an important purpose. Bat body odor could originate from multiple sources, i.e., hair, faeces, or snout secretions, and comparing VOC profiles across these sources helps us determine which source most likely contributes to the odor cues detected by crickets. This is especially important for limonene, because it is a plant-associated terpenoid and not a typical animal-derived volatile, as the reviewer #2 pointed out. Our analysis showed that hair and snout secretions shared similar VOC profiles, while faecal VOCs were distinct and lacked limonene. This supports snout secretions as the likely primary source. To make this purpose clearer, we have revised the Methods section to state that this analysis was conducted to investigate potential biological sources of bat body odor (revised manuscript, lines 494–495, 567–569).”

      Line 494–495: “To identify the biological source of the characteristic body odor, we compared the VOC profiles from hair, faeces, and snout secretions.”

      Line 567–569: “Principal component analysis (PCA) based on a binary (presence/absence) matrix was performed using the vegan package in R to compare profiles from hair, feces, snout secretions, and bat body odor.”

      (2) Figure 3B seems to be incorrect. The significance of n-hexane and 1% limonene is ***P < 0.001, while for 10% limonene, why is it only two stars, **P < 0.01? Please check it.

      We have rechecked the statistical analysis and confirmed that the original annotation was correct. The significance levels in original Figure 3B (Figure 4B in the revised manuscript) were based on Bonferroni-corrected paired t-tests comparing each limonene concentration with the n-hexane control. The adjusted P value was 0.0006 for 1% limonene (P < 0.001) and 0.00485 for 10% limonene (P < 0.01). The higher adjusted P value for 10% limonene reflects greater among-individual variation at this concentration, which may be due to differential sensitivity of individual antennae at higher doses. We have now added the exact adjusted P values in the revised manuscript, lines 163–167.

      Lines 163–167: “EAG responses to limonene were concentration-dependent (repeated-measures ANOVA, F(5, 25) = 24.95, P < 0.001, η<sup>2</sup>p = 0.83; Figure 4B), with both 1% and 10% limonene solutions eliciting significantly stronger responses than the hexane control (Bonferroni-corrected paired t-tests, 1%: t(5) = 10.77, P < 0.001, Hedges' g = 3.82; 10%: t(5) = 6.92, P = 0.005, Hedges' g = 2.46).”

      Finally, we would like to express our sincere gratitude to the editors and the reviewers for the thoughtful feedback. Your comments have significantly improved the quality and clarity of our manuscript. We hope the revised version satisfactorily addresses all concerns raised.

    1. eLife Assessment

      This important study systematically investigates repeat expansion in the plant Arabidopsis thaliana using a new k-mer-based method, expanding on previous smaller studies, to comprehensively identify cis- and trans-acting loci associated with repeat dynamics. The approach is methodologically sound, and the exploration of different k-mer lengths and use of a more complete reference assembly strengthen confidence in the analysis while clarifying its limitations and the method is broadly applicable to large-scale short-read datasets for assessing copy-number variation and genomic repeat content. The findings are convincing in their scope and novelty for A. thaliana. It will be of interest to learn in future how the approach generalizes to species with substantially greater or lower repeat content.

    2. Reviewer #1 (Public review):

      [Editors' note: this version has been assessed by the Reviewing Editor without further input from the original reviewers.]

      Summary:

      Overall, this study is an excellent and systematic investigation of the expansion of repeat sequences in Arabidopsis thaliana, and the genetic mechanisms underlying these expansions. Many of the key findings here confirm smaller studies of both repeat sequence variation and the individual genes associated with the expansion of various repeat classes. The authors present a highly effective and practical approach that requires datasets that are far more readily available than the multiple reference genomes used to annotate repeat variation in recent works. Therefore, they provide an approach that shows significant promise in non-model systems in which far less is known of repeat variation and its underlying drivers.

      Strengths:

      This is a very methodologically sound study that extends the relatively well-studied Arabidopsis thaliana repeat landscape with more systematic sampling, highlights the loci associated with repeat expansions (many of which were previously identified in a piecemeal manner), and provides some evolutionary inference on these.

      Weaknesses:

      Regarding cis-QTLs: I foresee at least two causes of these associations: non-repetitive cis-acting sequences that promote or permit the expansion of local repeats, and variation in repeat sequences themselves that directly tag the expanding sequence itself. It's arguable whether these are truly two distinct classes, but an attempt to discriminate between them may provide some insight as to the local factors that allow for repeat expansion, beyond the mere presence of a repeat sequence. One way to discriminate these could be to map the ~1300 12-mer frequency profiles on the reference genome, and filter any SNPs with elevated 12-mer frequency from the GWAS (or to categorize them independently).

      I also have a question regarding the choice of k=12 in kmer profile analyses. Did the authors perform any GWAS with other values of K? If so, how did the results change? I would expect that as K is increased, the associations would become more specific to individual repeat families, possibly to the point where only cis-acting loci are detected. The authors show convincing evidence that k=12 is appropriate; however, I would be interested to see if/how GWAS results vary among e.g. k=10, 12, 15, 18.

    3. Reviewer #2 (Public review):

      Summary:

      The authors introduce a K-mer-based method for profiling repeat content within a species, applied here to 1,142 A. thaliana genomes sequenced with short reads. This approach allowed them to bypass the challenges of genome assembly, particularly for repetitive regions, while still quantifying copy number variation. Their analysis identified >50 trans-acting loci regulating repeat abundance, enriched for genes involved in DNA repair, replication, and methylation. They also speculate on the role of selection in shaping genome repeat content, arguing that purifying selection tends to suppress alleles that promote repeat expansion.

      The work presents a scalable way to extract meaningful insights from the large quantities of short-read datasets available. However, I have several concerns regarding the methodology, scope of claims, and interpretation of results.

      Strengths:

      The authors leverage a large dataset, >1100 samples, of A. thaliana. The scale of the study is impressive and clearly bolsters their findings. Additionally, this provides a framework for future, large-scale studies and offers a solid foundation for hypothesis generation. The k-mer-based method is generally practical for large-scale analysis and should be transferable to other datasets. Finally, the authors are commendably upfront about many of the project's limitations.

      Weaknesses:

      The decision to use k=12 is loosely justified. While the authors performed a sweep of k-mer lengths (from 5-20) and noted computational constraints, the choice is highly dataset-specific. Benchmarking across different k values with additional datasets (especially including other species) would strengthen confidence in the robustness of the method.

      All analyses rely exclusively on the TAIR10 reference genome, which is incomplete and known to collapse certain repetitive regions. This dependence raises concerns that some repeats (especially recently expanded or highly variable ones) are systematically undercounted. With improved A. thaliana assemblies now available, testing the method against a more complete reference would alleviate these concerns.

      The manuscript's conclusions are framed in very broad terms (e.g., "shaping genome evolution in plants"). However, the study is restricted to a single species, A. thaliana, which may not represent other plants. While the findings may suggest general principles, the claims in the abstract and conclusion should be moderated to reflect the study system more accurately.

      The identification of >50 trans-acting loci enriched for DNA repair and replication genes is compelling, but the conclusions remain correlational.

    4. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      This important study systematically investigates repeat expansion in the plant Arabidopsis thaliana using a new k-mer-based method, expanding on smaller studies to more comprehensively identify cis- and trans-acting loci associated with repeat dynamics. The approach is methodologically sound and broadly applicable to large-scale short-read datasets for assessing copy number variation and genomic repeat content. While convincing in its scope and novelty, the findings would be further strengthened with exploratory analyses of datasets from other species with more or fewer repeats in their genomes.

      We agree with the assessment and appreciate the Editor’s handling of our manuscript and careful consideration of our work.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      Overall, this study is an excellent and systematic investigation of the expansion of repeat sequences in Arabidopsis thaliana, and the genetic mechanisms underlying these expansions. Many of the key findings here confirm smaller studies of both repeat sequence variation and the individual genes associated with the expansion of various repeat classes. The authors present a highly effective and practical approach that requires datasets that are far more readily available than the multiple reference genomes used to annotate repeat variation in recent works. Therefore, they provide an approach that shows significant promise in non-model systems in which far less is known of repeat variation and its underlying drivers.

      Thank you for your comments and careful consideration of our work.

      Strengths:

      This is a very methodologically sound study that extends the relatively well-studied Arabidopsis thaliana repeat landscape with more systematic sampling, highlights the loci associated with repeat expansions (many of which were previously identified in a piecemeal manner), and provides some evolutionary inference on these.

      Weaknesses:

      Regarding cis-QTLs: I foresee at least two causes of these associations: non-repetitive cis-acting sequences that promote or permit the expansion of local repeats, and variation in repeat sequences themselves that directly tag the expanding sequence itself. It's arguable whether these are truly two distinct classes, but an attempt to discriminate between them may provide some insight as to the local factors that allow for repeat expansion, beyond the mere presence of a repeat sequence. One way to discriminate these could be to map the ~1300 12-mer frequency profiles on the reference genome, and filter any SNPs with elevated 12-mer frequency from the GWAS (or to categorize them independently).

      While it would be interesting to further investigate the mechanisms underlying cis-QTLs, we do not believe this dataset can distinguish between the two proposed models of cis-variation. As argued in the manuscript, the observed cis-association signals are linked to repeat-associated SNPs and therefore are most likely driven by variation in repeat copy number, supporting the latter hypothesis.

      I also have a question regarding the choice of k=12 in kmer profile analyses. Did the authors perform any GWAS with other values of K? If so, how did the results change? I would expect that as K is increased, the associations would become more specific to individual repeat families, possibly to the point where only cis-acting loci are detected. The authors show convincing evidence that k=12 is appropriate; however, I would be interested to see if/how GWAS results vary among e.g. k=10, 12, 15, 18.

      We attempted to regenerate the primary datasets using K = 14, but found that generating a complete 14-mer matrix for the number of samples analyzed was computationally infeasible, even after filtering low-frequency K-mers such as singletons. While this analysis could likely be done by redesigning the pipeline around a database-backed approach, we considered such development beyond the scope of this revision.

      As a compromise, we regenerated the primary dataset and repeated the GWAS analyses using 10-mers. To assess the impact of K-mer length, we compare GWAS results generated from both 10-mers and 12-mers in the new Figures S19-21 and describe these results in expanded discussion on the impact of K-mer length on our results.

      Reviewer #2 (Public review):

      Summary:

      The authors introduce a K-mer-based method for profiling repeat content within a species, applied here to 1,142 A. thaliana genomes sequenced with short reads. This approach allowed them to bypass the challenges of genome assembly, particularly for repetitive regions, while still quantifying copy number variation. Their analysis identified >50 trans-acting loci regulating repeat abundance, enriched for genes involved in DNA repair, replication, and methylation. They also speculate on the role of selection in shaping genome repeat content, arguing that purifying selection tends to suppress alleles that promote repeat expansion.

      The work presents a scalable way to extract meaningful insights from the large quantities of short-read datasets available. However, I have several concerns regarding the methodology, scope of claims, and interpretation of results.

      Thank you for your comments and careful consideration of our work.

      Strengths:

      The authors leverage a large dataset, >1100 samples, of A. thaliana. The scale of the study is impressive and clearly bolsters their findings. Additionally, this provides a framework for future, large-scale studies and offers a solid foundation for hypothesis generation. The k-mer-based method is generally practical for large-scale analysis and should be transferable to other datasets. Finally, the authors are commendably upfront about many of the project's limitations.

      Weaknesses:

      The decision to use k=12 is loosely justified. While the authors performed a sweep of k-mer lengths (from 5-20) and noted computational constraints, the choice is highly dataset-specific. Benchmarking across different k values with additional datasets (especially including other species) would strengthen confidence in the robustness of the method.

      Our decision to use 12-mers to profile genome content in A. thaliana was based on an empirical evaluation of the sensitivity and specificity of different K-mer lengths for this application. Because our goal was to characterize intraspecific variation in genome content, we did not extend this analysis beyond A. thaliana. Nevertheless, we agree that alternative K-mer lengths may capture additional and potentially useful information and that using the method in other species would be interesting.

      Although we found generating a 14-mer dataset to be computationally infeasible, we regenerated the primary dataset and repeated the GWAS analyses using 10-mers. To assess the impact of K-mer length, we compare GWAS results generated from both 10-mers and 12-mers in the new Figures S19-21 and describe these results in expanded discussion on the impact of K-mer length on our results.

      All analyses rely exclusively on the TAIR10 reference genome, which is incomplete and known to collapse certain repetitive regions. This dependence raises concerns that some repeats (especially recently expanded or highly variable ones) are systematically undercounted. With improved A. thaliana assemblies now available, testing the method against a more complete reference would alleviate these concerns.

      To our knowledge there is not a published gapless A. thaliana assembly, although several of the recent assemblies are pretty close. Although we continue to use the 1001 Genomes SNP dataset that relies on TAIR10 for the GWAS analyses, we now present Figure 2A and Figure S9 using the Col-PEK assembly (Hou, Wang, Cheng, Wang & Jiao 2022 Molecular Plant).

      The manuscript's conclusions are framed in very broad terms (e.g., "shaping genome evolution in plants"). However, the study is restricted to a single species, A. thaliana, which may not represent other plants. While the findings may suggest general principles, the claims in the abstract and conclusion should be moderated to reflect the study system more accurately.

      We concede that A. thaliana does not possess a typical plant genome and thus have moderated our claims throughout the paper.

      The identification of >50 trans-acting loci enriched for DNA repair and replication genes is compelling, but the conclusions remain correlational.

      We agree that the conclusions of our work are correlational, as they are based on GWAS associations and have not been validated through functional experiments. We would love to see tests of the hypotheses we presented, but we believe this is beyond the scope of the current study.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Minor comments:

      (1) The Snakemake workflow Kmer-it had a few minor bugs that prevented use with modern Snakemake versions, which I fixed as part of the review (submitted as a pull request). Please be sure to validate that the whole workflow works with the latest Snakemake version.

      We appreciate the Reviewer’s testing of our software and have updated it accordingly.

      (2) L162: Since you already map samples to a reference genome to filter organellar DNA, consider adding some form of filter or correction for duplication rate in paired-end data, which I have found to be one major source of batch effects as observed here between sequencing centers.

      We appreciate this suggestion. In earlier versions of the analysis, we removed duplicated reads, but doing so substantially weakened the relationship shown in Figure 1C. Because it is difficult to distinguish duplicate reads arising from genuine repeat copy number variation from those resulting from technical artifacts, we ultimately chose not to include duplicate removal in the final pipeline. However, we did remove the effect of the sequencing center using a linear model.

      (3) Have you used this approach with raw long-read data? Do you see any difference between Illumina and low-error-rate long read data (e.g., HiFi, R10 Nanopore with SUP calling)? This could be compared in a directly paired manner using some of the recent papers with HiFi/ONT data on 1001G accessions, e.g., Lian et al, 2024, Wlodzimierz et al, 2023, Teasdale et al, 2025. (I do not believe such a comparison is required to prove the utility of this method, but it could be of interest as such sequencing technologies become increasingly cost-competitive).

      We have not evaluated this approach using raw long-read sequencing data, although we agree that it would be an interesting direction for future work. As noted in the manuscript, we detected significant batch effects attributable to sequencing center, even among datasets generated with the same sequencing technology. Given these observations, we are cautious about comparing K-mer frequencies across sequencing platforms that have distinct error profiles, as technical differences could confound biological signals.

      Reviewer #2 (Recommendations for the authors):

      (1) I would strongly suggest modulating some of the claims made in the paper, especially with regard to "genome evolution in plants", given that the paper focuses on A. thaliana. Alternatively, the authors could test the method in a different dataset from a different species. This would alleviate some concerns regarding the choice of k-mer length and demonstrate robustness.

      We agree that A. thaliana is not representative of most plant genomes and have revised the manuscript to better reflect this limitation. Since the primary goal of this study was to characterize intraspecific variation in genome content within A. thaliana, we believe that extending the analysis to an additional species falls beyond the scope of the present work. We do not believe that our choice of K-mer length undermines the robustness of the approach. To evaluate this concern, we repeated the GWAS analyses using 10-mers and found broadly consistent results, which are presented in Supplementary Figure S19-21. These findings suggest that the major conclusions are not strongly dependent on the specific K-mer length selected.

      (2) A k-mer length of 12 is quite small. There are 4^12 possible 12-mers (~17 million), so you would expect that all 12-mers would occur in the A. thaliana genome by chance at least once. While larger k-mers may be more computationally expensive to compute, it would be worth repeating the analysis with a larger k-mer length.

      We attempted to regenerate the primary datasets using K = 14, but found that generating a complete 14-mer matrix for the number of samples analyzed was computationally infeasible, even after filtering low-frequency K-mers such as singletons. While this analysis could likely be done by redesigning the pipeline around a database-backed approach, we considered such development beyond the scope of this revision.

      (3) Line 689 on page 31 should include the figure number.

      Thanks! Fixed.

    1. eLife Assessment

      This important study presents an innovative combination of experimental evolution and CRISPR screening to study the genetic basis of a polygenic trait, resistance against octanoic acid in Drosophila. The evidence supporting the authors claims is solid. The work will be of interest to the Drosophila community and researchers interested in functional testing of polygenic traits.

    2. Reviewer #2 (Public review):

      Summary:

      The authors studied the resistance against octanoic acid, a compound of the noni fruit in D. simulans, using experimental evolution and resistance/susceptibility in D. melanogaster cells. They identified novel candidate genes and performed functional tests.

      Strengths:

      The idea of using experimental evolution of a non-resistant species to develop resistance is interesting and the idea of narrowing down a large list of candidate loci by CRISPR based gene knockout in cell culture is innovative. The reviewer also liked the (easy) follow up experiments to validate the results.

      Comments on revised version.

      Weaknesses:

      - The experiments to validate the effect of candidate genes did not match the experimental evolution conditions.

      This point has been confirmed by the authors.

      - The statistical analysis suffers from some problems and insufficient description of the analyses performed.

      Has not been addressed in their response.

      - Although D. simulans GWAS data are available, the authors did not make an attempt to estimate the effect of selected variants in the candidate genes in the GWAS data set.

      This has now been included in the discussion. I would recommend that they make the distinction between genetic and adaptive architecture, as this matches their verbal description.

      - The reviewer would have liked to see more connection between the experimental evolution and GWAS data. As some D. simulans genotypes have similar resistance as D. sechellia, it would have been interesting to test whether this genotype contributed to the observed resistance.

      The reviewer is happy with the response.

      - At several places the authors discuss the challenge of studying a polygenic trait, but at the same time they claim to have detected and validated candidate genes. It would be helpful if the authors could discuss why they consider that their assays could really detect the contribution of single loci to the polygenic trait. In particular, when GWAS did not detect their candidate genes.

      The reviewer is not satisfied with the arm waving explanation of the authors. The important question is how much of the phenotypic variation is explained by the two candidate genes? The reviewer is inclined that based on the weak selection response, the variation is too little to be detected experimentally. Nevertheless, the overexpression of alkbh7 alone was sufficient to generate resistance levels similar to the ones in d. Melanogaster. Hence, it is not adequate to speak of small effects. This discrepancy requires more discussion.

      - It is not clear to the reviewer why the authors did not pay more attention to the highly significant peaks emerging from the experimental evolution study. Their functional validation would have been biologically more plausible.

      This point remains valid, in particular in the light of the discrepancy between the very limited selection response of alkbh7 and its large phenotypic effect after overexpression.

      Impact:

      - Given the obvious challenges of functional testing of polygenic traits and the clear limitations of the interpretation of the results, the study will be helpful for future studies aiming to characterize polygenic traits. Unfortunately, the results are just another piece of controversial results regarding resistance against octanoic acid-a trait that is rather easy to evaluate.

      The reviewer did not find the reply satisfactory.

    3. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Marconcini et al. report results of an ambitious study on the genetic mechanisms that contribute to resistance of Drosophila flies to the toxin octanoic acid (OA). This study was motivated by two observations: first, Drosophila sechellia, a close relative of D. melanogaster, has evolved specialized feeding on fruits of Morinda citrifolia, which contain high concentrations of OA and second, that artificial selection on Drosophila simulans, a sister species of D. melanogaster, can generate higher resistance to OA. Previous studies had performed genetic mapping studies between D. simulans and D. sechellia that implicated certain genomic regions in resistance to OA and, in particular, implicated several Osiris gene paralogs as contributing to resistance, though the molecular mechanisms of resistance remain unclear. In this study, Marconcini et al. performed two major experiments. First, they performed evolution-and-resequence on Drosophila simulans populations exposed to OA for 50 generations and identified candidate regions with excessive shifts in allele frequencies as candidate regions containing OA resistance genes in D. simulans. Second, they performed a CRISPR knock-out screen in a D. melanogaster cell line to identify genes that contribute to OA resistance and susceptibility.

      Evolve-and-resequence yielded many candidate genomic regions with extreme allele frequency shifts, which may be regions containing OA resistance genes, or linked genes, or regions that happen to show a strong shift in all replicate populations by chance. As the authors note, detecting significant shifts in allele frequencies is a challenging problem, and the authors use two measures of allele frequency shifts (the Cochran-Mantel-Haenszel method and Bait-ER) and perform simulations under neutrality to estimate a reasonable significance threshold. I am not entirely convinced by this method of estimating significance levels, because the simulations involve assumptions that may not be met by the real populations. I would think that a permutation test would provide an assumption-free method of estimating significance levels. I have tried to think whether there is something about the design of these experiments that would preclude the use of permutation tests (which are used widely for genome-wide studies, such as QTL), but I can't think of one. Perhaps the authors are aware of a reason permutation tests would be invalid here, and if so, they should state this reason.

      Significance thresholds have been estimated using a variety of approaches in the literature, including false discovery rate (FDR) control, simulations under neutral drift with predefined cut-offs, arbitrary significance thresholds, haplotype-based analyses, permutation tests, amongst other methods. Our choice was guided by a review (doi:10.1186/s13059-019-1770-8), which compared the performance of several of these approaches and found that the relatively simple assumptions underlying the Cochran–Mantel–Haenszel (CMH) test often performed as well as, or better than, more complex alternatives. As with most significance thresholds used in genome-wide analyses, the CMH threshold applied here is ultimately based on a degree of arbitrariness, although we aimed to be as stringent as possible. As a complementary method, we used Bait-ER; here the threshold used followed the recommendations provided in the original publication describing this method (doi:10.1111/jeb.14134).

      One reason that permutation-based approaches may be less widely adopted than theoretical null models is the extensive linkage disequilibrium among SNPs, which results in large blocks of correlated variants that cannot be considered independent observations (and therefore not shuffled). In principle, permutations could be performed at the haplotype-block level; however, defining haplotype blocks itself requires selecting thresholds or criteria that are often user defined or based on significance cut-offs. Although several tools are available for this purpose, our experience has been that the resulting block definitions remain sensitive to these choices and therefore introduce a comparable degree of arbitrariness. In reality, there is not yet a standard method in the field that has emerged as the “best” one.

      There is overlap between regions detected by the two methods, but the methods disagree for many regions. The authors state that a "majority of prominent peaks were found by both methods," but I am unclear on what "prominent" means here. It would be more helpful to be more quantitative about the extent of overlap.

      To quantify the agreement between the two methods, we calculated the overlap in genomic coverage (base pairs) between CMH and Baiter candidate regions. For G25, the two methods shared 1.90 Mb of candidate regions. The overlap encompassed 65.3% of the genomic span identified by CMH and 64.1% of the span identified by Bait-ER. For G50, the methods shared 4.77 Mb of candidate regions. The overlap encompassed 63.6% of the CMH candidate span and 99.9% of the Bait-ER candidate span. We have now replaced the admittedly qualitative statement the reviewer highlighted ("majority of prominent peaks were found by both methods”) with this information.

      The authors hypothesized that the response would be at similar genomic loci in all populations (line 222). It seems at least possible that epistatic interactions would lead to different combinations of alleles evolving in each population. I wonder if it would be possible to test whether there is heterogeneity in the responses across the replicate populations.

      We agree that epistatic interactions could lead to different allelic combinations being favored in different replicate populations. However, both BaitER and the CMH test are designed to detect parallel evolutionary responses across replicates, which was the focus of our study. One approach for testing population-specific responses is the LRT-2 test (doi:10.1534/genetics.118.301824). We did not pursue this analysis because it would likely generate many additional candidate loci, making interpretation more challenging, while providing limited additional insight into the repeatable genomic responses that were the primary focus of this work.

      The evolve-and-resequence method yielded many possible regions contributing to OA resistance in D. simulans, but perhaps too many regions to test directly or even to build sensible hypotheses about the genes involved. Thus, the authors performed a second experiment to try to narrow down the list of possible candidate genes. They performed a CRISPR knockout screen in a D. melanogaster cell line for genes that contribute to resistance or susceptibility to OA. The authors identify several limitations of this experiment, but they nonetheless identified several genes where knockouts contribute to OA susceptibility or resistance. Intersecting top hits with regions that experienced selection identified two "resistance" genes: kraken and Alkbh7. The selection hit at kraken is quite compelling, whereas the evidence at Alkbh7 is less strong because only two SNPs were marginally significant. Further functional assays, including gene knockouts in D. melanogaster and D. sechellia, provide some support for the claim that both of these genes can contribute to resistance to OA in flies.

      Beyond the few issues raised above, I do not have significant questions about methodology or the results. I do think, however, that the authors should be more conservative about the implications and significance of their results. For example, on line 139, the authors claim that this intersection approach provides a "powerful paradigm to investigate ecotoxicology." I am not sure I agree that the identification of two genes that may contribute to OA resistance, after a seemingly heroic selection experiment and CRISPR screen, suggests that this method is all that powerful. It seems that most of the genes that contribute to the selection response remain unidentified.

      We agree with the reviewer and have modified the sentence accordingly. While we believe that integrating the approaches discussed in this paper can provide valuable insights into the genetic basis of ecotoxicological traits, these approaches are not a panacea for traits with highly complex genetic architectures, such as OA resistance. Nevertheless, the identification of two candidate genes with some evidence of contributing to the trait represents a meaningful advance toward understanding its underlying genetic basis.

      Finally, given that one motivation of this project was to identify genes that contribute to evolved resistance to OA, I am surprised that the authors did not generate CRISPR alleles of kraken and Alkbh7 in D. simulans and then use these together with the existing alleles in D. sechellia to perform reciprocal hemizygosity tests to determine if these two genes actually contribute to evolved resistance in D. sechellia. This test is simpler to perform and may be more sensitive than the allelic replacement that the authors propose (lines 446-449).

      While generating null alleles for the candidate genes in D. simulans is beyond the scope of this revision, we note that we did attempt reciprocal hemizygosity tests using the mutants available in D. melanogaster. However, the resulting hybrids were recovered in low numbers, and these animals were rather weak, making them unsuitable for the severe OA exposure conditions employed in our assays. We agree that, where feasible, future studies should incorporate reciprocal hemizygosity tests, as they represent a powerful approach for validating the contribution of candidate genes to the trait of interest. We have revised the final sentence of the Results section accordingly.

      Reviewer #2 (Public review):

      Summary:

      The authors studied the resistance against octanoic acid, a compound of the noni fruit, in D. simulans, using experimental evolution and resistance/susceptibility in D. melanogaster cells. They identified novel candidate genes and performed functional tests.

      Strengths:

      The idea of using experimental evolution of a non-resistant species to develop resistance is interesting, and the idea of narrowing down a large list of candidate loci by CRISPR-based gene knockout in cell culture is innovative. The reviewer also liked the (easy) follow-up experiments to validate the results.

      Weaknesses:

      The reviewer is not convinced of the conceptual idea behind their approach: the intersection of the two approaches implicitly assumes that null alleles (or at least compromised alleles) should be selected during experimental evolution. The reviewer considers this unlikely, and the authors made no attempt to test this implicit hypothesis in their data.

      We respectfully disagree with the reviewer’s interpretation of the conceptual idea behind our approach. Our strategy did not assume that experimental evolution selects for null alleles, but rather for any type of variant that could contribute to the trait being selected for (i.e., increases in OA resistance), pointing to candidate genes contributing to the trait. Like many evolve-and-resequence experiments, this approach identified hundreds of candidate genes. This is why we took an orthogonal, genome-wide CRISPR screening approach, where loss-of-function mutations could lead to increases or decreases in OA tolerance of cultured cells. However, we stress that the naturally selected alleles – which could be gain or loss of function – may have much subtler phenotypic effects than the null alleles used for functional validation.

      Along the same lines, it is not clear how to reconcile an upregulation of candidate genes in resistant flies with the knockout experiments.

      We respectfully disagree that these findings are difficult to reconcile. The observed upregulation of the candidate genes in the selected, resistant D. simulans is consistent with a role in promoting OA resistance, while the knockout experiments (whether in cultured cells or in whole animals) demonstrate that loss of gene function reduces resistance. These observations are complementary: increased expression is associated with enhanced resistance, whereas complete loss of function impairs it. The knockout experiments of kraken and Alkbh7 were intended as functional validation of gene involvement and do not imply that the alleles selected during experimental evolution are loss-of-function alleles (we rather hypothesize that the selected alleles are gain-of-function through some, as yet undetermined, mechanism).

      The experiments to validate the effect of candidate genes did not match the experimental evolution conditions.

      This is correct. As our results suggest that the phenotype is shaped by multiple genes, such that the effect of any individual gene is likely modest compared to their combined contribution. Consequently, detecting and validating the effect of a single gene requires more stringent OA conditions than those needed to observe the overall phenotypic response over the course of several generations. In addition, the shorter-term plate assay was more practical for higher temporal resolution of the analysis of mortality in the presence of OA.

      The statistical analysis suffers from some problems and an insufficient description of the analyses performed.

      Although D. simulans GWAS data are available, the authors did not make an attempt to estimate the effect of selected variants in the candidate genes in the GWAS data set.

      We agree that comparing the experimental evolution and GWAS results (from our previous work, doi:10.1093/g3journal/jkag032) is of considerable interest. We have now expanded the Discussion to explicitly discuss the relationship between the two datasets. Overall, the overlap between the approaches was limited, although two GWAS candidate genes, bez and CG13003, fall within genomic regions exhibiting significant CMH signals at generation 25 (but not generation 50) of the evolve-and-resequence experiment. (We note that these genes could not have been identified in our CRISPR screen, as they are not expressed in S2R+ cells). More broadly, differences between the GWAS and evolve-and-resequence results likely reflect the distinct evolutionary processes captured by each approach: GWAS maps standing phenotypic variation among isofemale lines, whereas experimental evolution tracks allele frequency changes under sustained selection. Understanding why some signals are shared whereas others are not – whether due to effect size, genetic background, epistasis, pleiotropic costs, or the contribution of initially rare variants – remain important open questions.

      The reviewer would have liked to see more connection between the experimental evolution and the GWAS data. As some D. simulans genotypes have similar resistance to D. sechellia, it would have been interesting to test whether this genotype contributed to the observed resistance.

      While D. simulans genotypes displayed a range of OA resistance levels, none approached D. sechellia levels of resistance (see Figure 2e from our previous work, doi:10.1093/g3journal/jkag032) (The reviewer might have conflated the data from our GWAS of D. melanogaster strain, shown in Figure 2b of that paper, where some lines of that species exhibit comparable resistance to D. sechellia under the conditions of that assay). Regardless, our experimental evolution data do not provide sufficient resolution to identify the specific favorable alleles underlying the response to selection. Instead, we detect genomic regions containing many linked variants whose frequencies change under selection. Consequently, a direct comparison between evolved genotypes and GWAS-associated genotypes is currently difficult. We note, however, that expression of the D. sechellia kraken allele in D. melanogaster did not produce a significant effect on resistance, suggesting that even the most promising candidate alleles might not have strong effects in isolation.

      At several places, the authors discuss the challenge of studying a polygenic trait, but at the same time, they claim to have detected and validated candidate genes. It would be helpful if the authors could discuss why they consider that their assays could really detect the contribution of single loci to the polygenic trait. In particular, when GWAS did not detect their candidate genes.

      Our results do not imply that kraken and Alkbh7 are major-effect loci or that they explain a substantial proportion of the phenotypic variation. Rather, our data indicate that these genes make measurable contributions to OA resistance, consistent with the expectation that complex traits are influenced by many loci of individually modest effect. The absence of these genes among the top GWAS candidates does not preclude their involvement, as GWAS and experimental evolution interrogate different aspects of the genetic architecture and differ in their power to detect loci of varying effect sizes and allele frequencies.

      It is not clear to the reviewer why the authors did not pay more attention to the highly significant peaks emerging from the experimental evolution study. Their functional validation would have been biologically more plausible.

      We agree that the significant peaks identified in the evolve-andre sequence experiment represent promising targets for future investigation. However, these peaks typically span large genomic regions containing tens to hundreds of genes, making it difficult to prioritize individual candidates based on the experimental evolution data alone. In this work, we focused our functional validation on genes independently supported by the CRISPR screen, which provided gene-level resolution. We fully acknowledge that additional causal genes are likely to reside within the selected regions and remain to be functionally characterized.

      Impact:

      Given the obvious challenges of functional testing of polygenic traits and the clear limitations of the interpretation of the results, the study will be helpful for future studies aiming to characterize polygenic traits. Unfortunately, the results are just another piece of controversial results regarding resistance against octanoic acid, a trait that is rather easy to evaluate.

      The reviewer appears to imply that a trait being straightforward to phenotype necessarily implies that its genetic basis should also be straightforward to resolve. Many classic complex traits, such as human height, are simple to measure yet have an extraordinarily complex, highly polygenic genetic architecture. We believe that OA resistance represents a similar challenge: while the phenotype is readily assayed, differences in assay conditions, genetic backgrounds, and the contribution of many loci of individually modest effect make its genetic basis difficult to dissect. We have strived to be cautious in our conclusions, in particular the evolutionary interpretations; nevertheless, to our knowledge, this is the first study to provide functional evidence supporting the contribution of specific genes to OA resistance in D. sechellia, combining both loss-of-function phenotypes and expression data. As emphasized by the title of our manuscript, we view the principal contribution of this work as demonstrating how complementary experimental approaches (both of which are fairly novel for study of toxin susceptibility/resistance genetics) can be integrated to prioritize and functionally evaluate candidate genes underlying complex adaptive traits.

      Recommendations for the authors:

      Reviewing Editor Comments:

      Both reviewers propose to include additional statistical and quantitative measures to strengthen the results. Furthermore, certain parts of the manuscript could be rephrased to make sure that the readers can understand more clearly the implications and significance of the results. To test whether the genes identified in the present manuscript (as contributing to octanoic acid resistance) are also involved in the evolution of the resistance, reciprocal hemizygosity tests using CRISPR alleles of kraken and Alkbh7 in D. simulans would be a plus, but such experiments are not required because the main focus of the paper is on the genetic basis of octanoic acid resistance, and not on evolution.

      We thank the Reviewing Editor for these constructive comments. We have revised the manuscript to clarify the interpretation and significance of our findings and have addressed the reviewers’ comments throughout. Regarding reciprocal hemizygosity tests, we agree that they would provide a valuable means of assessing the evolutionary contribution of candidate genes. However, generating the necessary reagents in D. simulans represents a substantial undertaking beyond the scope of the present study, particularly given the expected modest effects of individual loci underlying this highly polygenic trait. We have nevertheless revised the end of the Results section to mention reciprocal hemizygosity tests as an important direction for future work.

      Reviewer #2 (Recommendations for the authors):

      (1) Provide more details about the selection tests: which sequences were used? Please report p-values. It would also be important to discuss the possibility of false positives caused by a bottleneck in D. sechellia. A genome-wide analysis could help to see if the bottleneck increased the signal of positive selection.

      We are not entirely sure what additional analyses are being suggested. The sequences and methods used for the selection analyses are described in the Methods, and the statistical support for these analyses is reported in the Supplementary Material (MK test p-values and FUBAR posterior probabilities). We are also unclear as to how a genome-wide analysis would address the interpretation of the gene-specific selection analyses presented here. If the reviewer intended a different analysis, we would appreciate further clarification.

      (2) The significance level of the CMH test needs to be determined with the effective population size, not with the census size, as done by the authors. This is important, as it is not clear if the candidate genes remain significant after significance adjustment based on the effective population size.

      We thank the reviewer for this comment. In our analyses, effective population sizes were explicitly incorporated by Bait-ER to model the effects of genetic drift. The CMH significance thresholds, however, were obtained from neutral forward simulations following the recommended workflow for the method, which requires census population sizes rather than effective population sizes as input. To make these simulations as realistic as possible, we therefore used the observed census population sizes at each generation of the experimental evolution. Estimating generation-specific effective population sizes for use in an alternative simulation framework would require substantially more temporal data than are available here (e.g., sequencing many additional time points) and is beyond the scope of the present study.

      (3) Include the allele frequency trajectory across time for the candidate genes.

      We thank the reviewer for this suggestion. However, plotting allele frequency trajectories for the candidate genes is not straightforward because the evolve-and-resequence analysis identified broad linked genomic regions rather than individual causal variants. Each candidate region contains numerous SNPs spanning several kilobases, with each SNP exhibiting its own allele frequency trajectory across the 10 replicate populations. Consequently, there is no single representative trajectory for a given candidate gene, and plotting all SNPs within each region would be difficult to interpret. For kraken, however, we identified seven candidate regulatory SNPs and have plotted their individual allele frequency trajectories, which we include in Author response image 1 to illustrate the diverse patterns observed:

      Author response image 1.

      (4) Estimate the effect of candidate loci in the GWAS data (independent of significance).

      We thank the reviewer for this suggestion. However, we are not entirely sure what analysis is being proposed. In particular, it is unclear which variants the reviewer is referring to, as our candidate genes are associated with multiple linked variants rather than a single causal SNP. Consequently, we are unsure how the effect of a candidate locus should be estimated in the GWAS dataset. Moreover, there is no guarantee that the same variants are represented in both datasets. For example, variants filtered out during the GWAS because of low allele frequency may subsequently have increased in frequency during experimental evolution and therefore contributed to the evolve-and-resequence signals. We would appreciate further clarification of the analysis the reviewer has in mind.

      (5) Discuss the challenge of false positives (see: 10.1016/j.cub.2020.12.023).

      We thank the reviewer for this suggestion. We agree that false positives (as well as false negatives) are an important consideration when studying complex polygenic traits, both through the initial “screening” efforts (e.g., GWAS, experimental evolution) and follow-up functional validation (e.g., RNAi, mutant, overexpression analyses). We believe this issue is already addressed in the manuscript through our discussion of the limitations of the individual approaches, the polygenic nature of OA resistance, and our cautious interpretation of the functional validation results. Our conclusions are limited to identifying candidate genes that contribute to OA resistance, rather than claiming to have identified all of the loci or variants underlying the evolution of this trait. We are therefore not sure what additional discussion the reviewer has in mind and would appreciate further clarification if a specific point from the cited study is intended.

      (6) The figures with the Manhattan plots should be improved to indicate the overlapping genes in the Manhattan plots, rather than in the circle figures below. By using different colors this should be quite easy and clean.

      We thank the reviewer for this suggestion. However, we believe Author response image 2 more appropriately illustrate the overlap between the CMH and BaitER analyses. Although the Manhattan plots display individual SNPs, our candidate loci are defined by broader genomic regions comprising blocks of linked significant SNPs rather than by single variants. Simply highlighting SNPs within overlapping regions would therefore add visual complexity without providing additional biological insight beyond that already captured by the circos plots. As an illustration, we provide here an example of the generation 50 Manhattan plot with overlapping SNPs highlighted in red:

      Author response image 2.

      (7) A more focused discussion of the assumption that functional data from D. melanogaster can explain resistance in D. simulans or D. sechellia. At some places, epistatic interactions are mentioned, but the reviewer feels that the entire screen is based on the idea that similar effects are found across species, hence, this needs to be adequately reflected in the discussion.

      We thank the reviewer for raising this important point. We would like to emphasize that our study was not based on the assumption that functional effects identified in D. melanogaster or D. simulans necessarily explain the evolution of OA resistance in D. sechellia. Rather, our motivation stemmed from the longstanding difficulty of identifying individual genes underlying this highly polygenic trait using mapping approaches alone. We therefore sought to combine orthogonal experimental approaches to identify genes contributing to OA susceptibility (of D. melanogaster and D. simulans) and resistance (D. sechellia). We hoped, but did not assume, that genes supported by multiple independent lines of evidence would be informative for understanding the natural evolution of OA resistance in D. sechellia. Indeed, we acknowledge in the manuscript that the cell-based CRISPR screen, experimental evolution, and the natural evolution of D. sechellia occurred under very different selective contexts and timescales. As reflected in the title of our manuscript, our conclusions are intentionally framed around the identification of novel toxin resistance loci, while remaining cautious about their evolutionary interpretation.

      (8) Discuss that the controls in the RNAi test were quite variable. Could this reflect some problems with the assay?

      We do not believe that the variability among the control lines reflects a problem with the assay. Rather, the different controls (Gal4, UAS-RNAi etc.) represent distinct genetic backgrounds, each of which may exhibit a different baseline level of OA resistance. Indeed, in our recent GWAS of OA resistance (doi:10.1093/g3journal/jkag032), we observed substantial natural variation in OA resistance among D. melanogaster and D. simulans lines. Importantly, each RNAi line was compared with its corresponding genetic background control, so differences among control lines do not affect the interpretation of the individual RNAi experiments.

    1. eLife Assessment

      This study presents valuable findings on the functional consequences of nuclear envelope rupture caused by the depletion of nuclear pore complex subunits on the behaviour and organisation of chromosomes during mitosis in C. elegans embryos. The experiments are generally well-designed and executed; however, the evidence for some of the main conclusions is incomplete. This work is of potential interest to cell biologists working on cell division.

    2. Reviewer #1 (Public review):

      In this manuscript, the authors investigate the functional consequences of nuclear envelope rupture caused by the depletion of the nucleoporin NPP-3.

      They observe that loss of NPP-3 causes condensed chromosomes to localize to the nuclear periphery. This anchoring is independent of the pathway required to anchor heterochromatin and telomeres, but it depends on spindle assembly checkpoint proteins as well as centromere and kinetochore proteins. While the authors propose that relocalization of chromosomes to the nuclear periphery protects genome stability, they do not demonstrate this.

      Overall, some of the observations are interesting, but several points should be addressed. Furthermore, the manuscript could be much clearer if certain sections were shortened, simplified, or removed.

      Major points:

      (1) The title is misleading because the authors provide no experimental evidence that chromosome relocalisation protects genome stability. They are more cautious in the abstract, where they state that it 'may serve a protective role'. If they could provide stronger experimental evidence that chromosome relocalization protects genome stability, this would significantly strengthen the manuscript.

      (2) Here, the authors use acute inactivation of npp-3. Do chromosomes also localize to the periphery upon partial npp-3 inactivation? What are the minimal levels of nuclear envelope rupture that cause chromosomes to localize to the periphery? Given that NPP-3 and NPCs have pleiotropic functions, it would be important to analyze conditions where only a few nuclear envelope ruptures are induced. In such conditions, they might be able to explore the link between chromosome localization and genome stability.

      (3) The authors primarily examined P1 cells. Is the behaviour of the chromosome different between cells of different lineages?

      (4) The authors mentioned that defective chromosomal localisation does not occur upon npp-2 or npp-4 depletion. How do they explain this? Did they attempt to inactivate other NPPs in the Y complexes, and can they be certain that NPP-2 depletion is complete?

      (5) The section on AIR-1 (line 147) is confusing and could be removed. To my knowledge, air-1 depletion does not cause the appearance of multiple centrosomes, except maybe in a very few embryos. air-1 depletion causes major defects, so it is difficult to draw a parallel with npp-3 depletion.

      (6) The authors show that condensed chromosomes tend to localize to the nuclear envelope upon NPP-3 depletion. Do they condense at the nuclear envelope (NE), or do they condense first and then move to the periphery? This is unclear from the data presented in Figure 1D. Also, why do chromosomes condense earlier? This point could be discussed.

      (7) The authors evaluated the consequences of NPP-3 depletion on transcription using RNA sequencing. The relevance of this experiment is questionable, however, as npp-3(RNAi) embryos have significant general defects and not only mislocalised chromosomes.

      (8) In the co-depletion experiment npp-3(RNAi), X(RNAi) presented in Figure 3B, the levels of NPP-3 depletion seem highly variable. All the images shown are not similarly exposed, so it is difficult to evaluate these data.

      (9) Inactivation of mdf-1/2 suppresses the mislocalization of the chromosomes observed upon npp-3 inactivation. Does it also suppress the premature chromosome condensation phenotype?

      (10) Figure 5B: The delay induced by npp-3 depletion is not severe, based on the micrographs presented. The authors should show more representative images. The graph shows the elapsed time between NEBD and NER, and not NER to NEBD, as indicated.

      (11) The authors observed that depleting mdf-1 slightly enhanced the lethality associated with npp-3 inactivation. Based on this observation, they conclude that loss of chromosome anchoring exacerbates genomic instability and severely impairs embryonic survival. However, the genetic interaction is not strong, as npp-3(RNAi) embryos already present more than 95% embryonic lethality and have defects other than just mislocalized chromosomes (e.g., defects in kinetochore and spindle assembly).

    3. Reviewer #2 (Public review):

      Summary:

      The authors aimed to determine the molecular mechanisms by which nuclear pore component NPP-3/NUP205 regulates chromosome localization in C. elegans embryos. Previous studies had shown that depletion of NPP-3 caused premature chromosome condensation and movement of chromosomes to the nuclear periphery. Peripheral location of chromosomes is also observed under respiratory stress conditions, suggesting that peripheral chromosome positioning could act as a protective response to stress conditions. How NPP-3 affects chromosome positioning was unknown. Here, the authors conduct a screen to identify factors that promote chromosome relocation to the periphery in npp-3-depleted embryos, identifying an important role for spindle assembly checkpoint components in this process.

      Strengths:

      Using cytological tools to visualise chromosomes and nuclear envelope markers, the authors show that, in addition to the peripheral location of chromosomes, NPP-3 depletion causes partial rupture of the nuclear envelope and premature chromosome condensation. By systematically co-depleting NPP-3 and factors required for heterochromatin association with nuclear lamina (CEC-4), telomere binding to nuclear envelope (SUN-1 and POT-1), proteins required for the nuclear rupture repair machinery (BAF-1 and LEM-2), kinetochore proteins and components of the spindle assembly checkpoint (SAC) (MDF-1 and MDF-2), the authors convincingly show that SAC components are required for peripheral relocation of chromosomes in absence of NPP-3. The study also provides convincing evidence that peripheral relocation of chromosomes in the absence of NPP-3 has functional implications as it causes transcriptional deregulation and premature relocation of SAC components from the nuclear envelope to chromosomes. Co-depletion of NPP-3 and SAC components accelerates progression through miotic prophase and increases the incidence of defects in chromosome segregation during mitosis. These findings demonstrate that SAC proteins play an important role in regulating chromosome positioning during prophase (at least in the absence of NPP-3) and that they can regulate cell cycle progression at earlier stages than previously thought.

      Weaknesses:

      The authors also propose that NPP-3 depletion causes DNA damage; however, the evidence presented to support this claim is not as strong as that presented for the effects mentioned above. Also, the premature condensation of chromosomes appears as a clear consequence of NPP-3 depletion, but this intriguing phenotype remains unexplored.

    4. Reviewer #3 (Public review):

      Summary:

      This manuscript reports that RNAi depletion of the inner-ring nucleoporin NPP-3/NUP205 in Caenorhabditis elegans embryos causes nuclear envelope rupture, premature chromatin condensation, and relocalization of condensed prophase chromosomes to the nuclear periphery. Through a candidate epistasis screen, the authors argue that this relocalization requires spindle assembly checkpoint (SAC) components (MDF-1, MDF-2, SAN-1), inner kinetochore proteins (HCP-3, HCP-4, and partially KNL-1), and NE rupture-repair factors (BAF-1, LEM-2), but not the CEC-4 heterochromatin- or SUN-1/POT-1 telomere-anchoring pathways. They further show that NPP-3 loss extends prophase and the NEBD-to-anaphase interval in a SAC-dependent manner, redistributes MDF-1/MDF-2, and reduces import of KNL-1/BUB-1/HCP-1. Co-depletion of NPP-3 with MDF-1 abolishes both the arrest and the peripheral localization while increasing lagging chromosomes, HUS-1 foci, micronuclei, and lethality, which the authors interpret as evidence that peripheral positioning is protective.

      Weaknesses:

      (1) The "protective" conclusion is largely correlative. The protective claim rests on the observation that co-depleting MDF-1 (or MDF-2) with NPP-3 removes the peripheral localization and simultaneously increases DNA damage, micronuclei, and lethality. However, depleting a SAC component removes at least three things at once: the peripheral localization, the prophase extension, and the NEBD-to-anaphase arrest. Because loss of the SAC independently causes premature anaphase and genomic instability through well-established mechanisms unrelated to chromosome positioning, the current design cannot separate damage caused by loss of a protective peripheral location from damage caused by checkpoint bypass. As presented, the increased damage is at least as consistent with simple SAC bypass. To support the protective model, the authors should provide a manipulation that disrupts peripheral positioning without abrogating the SAC-dependent arrest (for example, via the BAF-1/LEM-2 or kinetochore depletion) and show that damage still increases. The LEM-2 co-depletion, which partially suppresses positioning, is a natural place to test whether micronuclei and HUS-1 foci also rise.

      (2) Knockdown efficiency of the partner gene in double RNAi is not verified. The double depletions are performed by cloning both gene fragments into a single vector. This risks reducing the effective dose of each dsRNA, so an apparent suppression in an npp-3; gene X (RNAi) condition could reflect weaker NPP-3 knockdown rather than a true epistatic relationship. The authors partially address this by showing that NPP-3::mCherry is still reduced in npp-3;mdf-1 (Figure S4A/B), which is helpful, but they do not demonstrate efficient knockdown of the partner genes in any double condition. For the key epistasis conclusions (MDF-1, MDF-2, HCP-3, HCP-4 suppressions), the knockdown of the second gene should be independently validated with a reporter strain for the second protein.

      (3) Alternative explanations for the transcriptomic and H3K9me3 data are not excluded. NPP-3 depletion blocks nuclear import of molecules smaller than ~70 kDa and arrests development at early gastrulation. Both the RNA-seq changes (30% of genes downregulated) and the increased H3K9me3 signal could therefore be secondary consequences of nucleocytoplasmic transport failure and developmental arrest rather than evidence of position-dependent transcriptional repression. Notably, the authors' own finding that up- and down-regulated genes show no chromosomal positional bias (Figure S2C/D) argues against a model in which peripheral repositioning drives silencing of specific chromatin domains. This section should be reframed more cautiously, with the transport/arrest confound explicitly discussed, and RNA-seq replicate number and differential-expression thresholds reported.

      (4) Evidence for SAC "activation in prophase" is indirect, and the effect is small. The claim of a novel prophase role for the SAC rests on MDF-1/MDF-2 intensity changes that are repeatedly described as "modest," "slight," or "mild," measured with small n and Student's t-tests, together with phenotypic suppression of prophase extension. There is no direct readout of SAC catalytic activity (for example, MCC assembly). The prophase-extension suppression by MDF-1 is the strongest evidence; the intensity data are weak support. I recommend tempering "the SAC is activated in prophase" to a hypothesis, and strengthening it with a more direct assay if feasible.

      (5) The BAF-1 arm of the model is inferred rather than demonstrated. The authors state that baf-1(RNAi) and npp-3;baf-1 produced clustering too severe for epistasis, so BAF-1's requirement for peripheral localization is not actually established genetically; it rests on increased BAF-1 accumulation (correlative) plus the LEM-2 partial suppression. The proposed BAF-1/CENP-C bridge is extrapolated from Drosophila (ref. 71). This is reasonable as a discussion hypothesis but should not be presented in the abstract or summary model as an established dependency.

    5. Author response:

      Public Reviews:

      Reviewer #1 (Public review):

      In this manuscript, the authors investigate the functional consequences of nuclear envelope rupture caused by the depletion of the nucleoporin NPP-3.

      They observe that loss of NPP-3 causes condensed chromosomes to localize to the nuclear periphery. This anchoring is independent of the pathway required to anchor heterochromatin and telomeres, but it depends on spindle assembly checkpoint proteins as well as centromere and kinetochore proteins. While the authors propose that relocalization of chromosomes to the nuclear periphery protects genome stability, they do not demonstrate this.

      Overall, some of the observations are interesting, but several points should be addressed. Furthermore, the manuscript could be much clearer if certain sections were shortened, simplified, or removed.

      Major points:

      (1) The title is misleading because the authors provide no experimental evidence that chromosome relocalisation protects genome stability. They are more cautious in the abstract, where they state that it 'may serve a protective role'. If they could provide stronger experimental evidence that chromosome relocalization protects genome stability, this would significantly strengthen the manuscript.

      We understand the more severe chromosome missegregation and DNA damage phenotypes in the npp-3 mdf-1 or npp-3 mdf-2 double RNAi may suggest that npp-3 RNAi is more sensitized to the loss of spindle assembly checkpoint (SAC) component MDF-1 or MDF-2 for chromosome protection, but may not clearly show that the localization of chromosomes to the nuclear periphery serves a chromosome protective function. Based on our current phenotypic analyses, we will tone down the title to “Prophase Chromosomes Relocalization to Nuclear Periphery in NPP-3/NUP205 Depletion Depends on the Spindle Assembly Checkpoint and Inner Kinetochore Proteins”.

      We have thought about artificially tethering the chromosomes to the nuclear periphery. However, without the same stimulus/defect and activation of the response pathway, the effects could be different. In future, to further clarify the functions of chromosome relocalization, we could analyze the chromosomal missegregation and DNA damage phenotypes in npp-3 lem-2 RNAi, which has a partial chromosomal nuclear periphery phenotype, to see if there is any quantitative relationship between chromosome nuclear periphery location and DNA protection.

      (2) Here, the authors use acute inactivation of npp-3. Do chromosomes also localize to the periphery upon partial npp-3 inactivation? What are the minimal levels of nuclear envelope rupture that cause chromosomes to localize to the periphery? Given that NPP-3 and NPCs have pleiotropic functions, it would be important to analyze conditions where only a few nuclear envelope ruptures are induced. In such conditions, they might be able to explore the link between chromosome localization and genome stability.

      We have not performed partial npp-3 RNAi yet. To explore whether the level of nuclear envelope rupture correlates with chromosome localization to the periphery or if a minimal level of nuclear rupture is required, we have performed the lacI::GFP reporter assay in different NPP RNAi to indicate the nuclear permeability defects (in Fig. S1A and B). npp-2, npp4 or npp-5 RNAi also causes increased permeability, at a comparable level as npp-3, npp-7 and npp-13 RNAi, yet only npp-3, npp-7 and npp-13 RNAi causes chromosome relocalization. This may explain the peripheral chromosome relocalization response may be more closely linked to the specific NPC subcomplex’s (inner ring and nuclear basket) or individual NPP’s function, rather than the general permeability defect or rupture size.

      However, we have also attempted to use other methods to induce targeted nuclear rupture with specific rupture size by laser ablation (Author response image 1, 355 nm UV pulsed laser marked by the white rectangle), and we could occasionally observe all chromosomes relocalizing to all over the nuclear periphery after 660 s upon laser ablation in different strains. However, the chromosome relocalization phenotype is not very consistent (~33%, n =12). Importantly, the laser also introduces DNA damage at the rupture sites, as marked by HUS-1 (Author response image 2), complicating the interpretation, so we did not include this attempt in the manuscript.

      Author response image 1.

      Time-lapse imaging of embryos expressing LEM-2::GFP and mCherry::H2B, a 15*30pixel rectangle was laser-microirradiated to induce nuclear envelope rupture (white rectangle). 0s is the time applying laser ablation. Scale bar 5 μm.

      Author response image 2.

      Time-lapse imaging of embryos expressing HUS-1::GFP and mCherry::H2B, a 15*30pixel rectangle was laser-microirradiated to induce DNA damage (white rectangle). 0 s is the time applying laser ablation. Scale bar 5 μm.

      Another attempt to achieve different nuclear rupture sizes is by observing nucleus in different stages of embryos. By npp-3 feeding RNAi approach, we observed an increase of the proportion of the nuclear circumference with nuclear rupture during embryonic development (Fig. S3D). Yet, the chromosome periphery phenotype is displayed in all embryo stages, suggesting that the localization of chromosomes to the nuclear periphery occurs across a range of rupture sizes.

      Given these observations, it appears that the extent of nuclear rupture or permeability increase may not be the sole determinant of chromosome relocalization. Instead, disruption of certain NUP subcomplexes or NPPs might stimulate this process.

      (3) The authors primarily examined P1 cells. Is the behaviour of the chromosome different between cells of different lineages?

      We observed that chromosomes localize to the nuclear periphery in all cells during embryonic development until the embryo dies, as demonstrated by snapshots and time-lapse imaging (Fig. S1C and Fig. S1E).

      We focused on the P1 blastomere for our analyses because its cell cycle timing and division orientation is well characterized, which facilitates precise examination of chromosome localization and cell cycle dynamics under different perturbation conditions. We will add a sentence at the beginning of that result section to clarify this choice: "The chromosome nuclear periphery localization in npp-3 RNAi is consistent in all cells at different embryonic stages (Fig. S1C and Fig. S1E). For analysis in distinct perturbation conditions, we specifically examined the P1 cells, whose cell cycle timing and division orientation is well characterized and suitable for such studies.”

      (4) The authors mentioned that defective chromosomal localisation does not occur upon npp-2 or npp-4 depletion. How do they explain this? Did they attempt to inactivate other NPPs in the Y complexes, and can they be certain that NPP-2 depletion is complete?

      (see response to point 2) We observed that npp-2 or npp-4 RNAi can cause nuclear rupture with certain permeability defect compared to wildtype (Fig. S1A and S1B), but not chromosomal nuclear periphery localization, suggesting that such chromosome localization to the nuclear periphery does not just depend on the nuclear rupture but could be a more specific phenotype related to individual NUP subcomplex’s (inner ring and nuclear basket) or NPP’s function. 

      For Y complex, we also now tested NPP-5 depletion, which shows permeability defect but did not show nuclear rupture marked by NPP-1:GFP or chromosomal periphery phenotype (Fig. S1A and S1B), which is consistent to the paper cited [1]. 

      To confirm the efficiency of NPP-2 depletion, we have observed the presence of smaller nuclei (as described in phenobank) by NPP-1::GFP marker (Fig. S1A) and a marked reduction in NPP2::GFP signal in npp-2 RNAi-treated embryos (unpublished data). These evidences support that NPP-2 depletion was effective. 

      (5) The section on AIR-1 (line 147) is confusing and could be removed. To my knowledge, air-1 depletion does not cause the appearance of multiple centrosomes, except maybe in a very few embryos. air-1 depletion causes major defects, so it is difficult to draw a parallel with npp-3 depletion.

      We agree with this suggestion and have decided to remove the section on AIR-1 in the manuscript text to avoid confusion. To clarify, air-1(RNAi) causes multiple centrosomes in only a small proportion of embryos (approximately 13%) [2] , primarily by affecting centrosome positioning [3]. and does not routinely lead to significant centrosome amplification.

      Our interest in AIR-1 was inspired by previous findings by Hachet et al., showing that AIR1/Aurora A localizes at sites without NPP-3 during mitotic entry—these sites are believed to correspond to centrosome locations [4]. This relationship prompted us to explore how AIR-1 depletion might influence NPP-3 localization at the nuclear envelope and its potential effects on chromosome positioning. Interestingly, we discovered that following air-1(RNAi) treatment, nuclei exhibit discontinuous NPP-3 localization on the nuclear envelope. Thereby, we could investigate the interplay between NPP-3 and chromosomal dynamics. Chromosomes localize to nuclear periphery specifically without NPP-3 (Author response image 3). This suggests a potential negative correlation of NPP-3 localization with the chromosomes. We did not imply any regulation by AIR-1.

      Author response image 3.

      Condensed chromosomes tend to localize at nuclear envelope sites without NPP-3 or NPP-7. (A) (C) Selected representative confocal images of all chromosomes localizing at the nuclear periphery in the control and air-1(RNAi) embryos expressing H2B::GFP and mCherry::NPP-3 (A) and GFP::NPP-7 and mCherry::H2B (C). The upper right image is the zoom-in view of the nucleus. Scale bar 10 μm. (B) (D) The intensity of GFP and mCherry is normalized to the average intensity along the nuclear envelope and plotted in the control and air- 1(RNAi) nucleus from (A) and (C).

      (6) The authors show that condensed chromosomes tend to localize to the nuclear envelope upon NPP-3 depletion. Do they condense at the nuclear envelope (NE), or do they condense first and then move to the periphery? This is unclear from the data presented in Figure 1D. Also, why do chromosomes condense earlier? This point could be discussed.

      Based on our time-lapse data in Figure 1D, most chromosomes appear to condense at or near the nuclear periphery, but there are still some chromosomes inside the nuclear space at approximately -600 seconds before NEBD, and then subsequently moving to the periphery during condensation (with full condensation at 100 s past NEBD). This suggests that initial condensation may occur both within the nucleus and at the periphery. Over time, chromosomes condense and cluster at the nuclear envelope. However, we did not separately measure the condensation level of individual chromosomes at the nuclear periphery versus in the middle of the nucleus, which could be challenging in live cells. So far, we cannot separate the chromosome nuclear periphery phenotype and the condensation.

      To confirm whether chromosome condensation precedes or follows relocation to the nuclear periphery, we could perform depletion of condensin II component, e.g. hcp-6, and see if lack of condensation affects chromosome relocalization.

      (7) The authors evaluated the consequences of NPP-3 depletion on transcription using RNA sequencing. The relevance of this experiment is questionable, however, as npp-3(RNAi) embryos have significant general defects and not only mislocalised chromosomes.

      We recognize that npp-3 RNAi embryos exhibit broad developmental defects, and therefore global transcriptional changes. Our RNA-seq analysis revealed that differentially expressed genes did not display positional bias within the genome. NPP-3 depletion downregulates many pathways, including pathways related to RNA polymerase II activity and cell cycle regulation, consistent with a global transcriptional downregulation. Thus, while the transcriptomic data are broad, it is consistent with the chromosome condensation phenotype.

      (8) In the co-depletion experiment npp-3(RNAi), X(RNAi) presented in Figure 3B, the levels of NPP-3 depletion seem highly variable. All the images shown are not similarly exposed, so it is difficult to evaluate these data.

      In our co-depletion experiments, the images for npp-3(RNAi) and X(RNAi) (Fig. 4) are displayed side-by-side under the same exposure conditions and scaled the same way with the same intensity thresholds to facilitate comparison. Despite that, some differences in the background intensity can be observed across samples. Thus, the normalised mCherry::NPP-3 signal intensity (subtracting the background) the single and double RNAi samples will be quantified and added to supplementary figure S4.

      To ensure consistency of double RNAi, we used ligation-based RNAi constructs designed to simultaneously target both NPP-3 and X, aiming to achieve comparable knockdown efficiencies. Nonetheless, variability in RNAi efficiency is a recognized limitation, and we interpret our data within this context. To further validate the knockdown, we will also assess the efficiency of the other gene X using the corresponding fluorescent reporter (see response to Reviewer 3, point 2).

      (9) Inactivation of mdf-1/2 suppresses the mislocalization of the chromosomes observed upon npp-3 inactivation. Does it also suppress the premature chromosome condensation phenotype?

      Inactivation of MDF-1 or MDF-2 suppresses both the chromosome nuclear periphery localization and extended duration of interphase to prophase and prometaphase in npp-3 inactivation. We have performed the analysis to assess the timing and extent of chromosome condensation in the single and double depletion conditions (Author response image 4). The chromosome condensation dynamics is similar in the mdf-1 RNAi and npp-3 mdf-1 double RNAi, as well as the control group, indicating that MDF-1 is also involved the premature chromosome condensation phenotype caused by NPP-3 depletion.

      Author response image 4.

      The dynamic changes of chromosome condensation parameter, in which 30% of pixels in the ROI analyzed is below the threshold scaled intensity (<77), in the different groups. The sample size is 5. Error bars show mean ± SEM.

      (10) Figure 5B: The delay induced by npp-3 depletion is not severe, based on the micrographs presented. The authors should show more representative images. The graph shows the elapsed time between NEBD and NER, and not NER to NEBD, as indicated.

      We have aligned the nuclear envelope reassembly (NER) time and highlighted the time point in the images (Fig. 5A). Our data show that in control embryos, this duration is approximately 780 seconds, while in npp-3(RNAi) embryos, it extends to about 930 seconds. This difference is statistically significant, as determined by one-way ANOVA (or appropriate nonparametric/mixed tests). The representative image is consistent with the quantification presented in Fig. 5B. We could add the corresponding videos to the supplementary information.

      (11) The authors observed that depleting mdf-1 slightly enhanced the lethality associated with npp-3 inactivation. Based on this observation, they conclude that loss of chromosome anchoring exacerbates genomic instability and severely impairs embryonic survival. However, the genetic interaction is not strong, as npp-3(RNAi) embryos already present more than 95% embryonic lethality and have defects other than just mislocalized chromosomes (e.g., defects in kinetochore and spindle assembly).

      It is correct that the average embryonic lethality observed in npp-3(RNAi) embryos reachs 95% (Fig. S6). Given the broad developmental defects and such high baseline lethality, the genetic interaction with mdf-1 is modest.

      Nevertheless, our findings highlight that in the npp-3 mdf-1 double RNAi condition, we observe significantly increased rates of lagging chromosomes (Fig. 5D), micronuclei formation (Fig. 7A), and elevated DNA damage (Fig. 7B). These effects support the idea that MDF-1-, MDF-2dependent chromosome anchoring to the nuclear periphery (and condensation) plays a positive role in NPP-3 depleted cells.

      Reviewer #2 (Public review):

      Summary:

      The authors aimed to determine the molecular mechanisms by which nuclear pore component NPP- 3/NUP205 regulates chromosome localization in C. elegans embryos. Previous studies had shown that depletion of NPP-3 caused premature chromosome condensation and movement of chromosomes to the nuclear periphery. Peripheral location of chromosomes is also observed under respiratory stress conditions, suggesting that peripheral chromosome positioning could act as a protective response to stress conditions. How NPP-3 affects chromosome positioning was unknown. Here, the authors conduct a screen to identify factors that promote chromosome relocation to the periphery in npp-3-depleted embryos, identifying an important role for spindle assembly checkpoint components in this process.

      Strengths:

      Using cytological tools to visualise chromosomes and nuclear envelope markers, the authors show that, in addition to the peripheral location of chromosomes, NPP-3 depletion causes partial rupture of the nuclear envelope and premature chromosome condensation. By systematically codepleting NPP-3 and factors required for heterochromatin association with nuclear lamina (CEC4), telomere binding to nuclear envelope (SUN-1 and POT-1), proteins required for the nuclear rupture repair machinery (BAF-1 and LEM-2), kinetochore proteins and components of the spindle assembly checkpoint (SAC) (MDF-1 and MDF-2), the authors convincingly show that SAC components are required for peripheral relocation of chromosomes in absence of NPP-3. The study also provides convincing evidence that peripheral relocation of chromosomes in the absence of NPP-3 has functional implications as it causes transcriptional deregulation and premature relocation of SAC components from the nuclear envelope to chromosomes. Codepletion of NPP-3 and SAC components accelerates progression through miotic prophase and increases the incidence of defects in chromosome segregation during mitosis. These findings demonstrate that SAC proteins play an important role in regulating chromosome positioning during prophase (at least in the absence of NPP-3) and that they can regulate cell cycle progression at earlier stages than previously thought.

      Weaknesses:

      The authors also propose that NPP-3 depletion causes DNA damage; however, the evidence presented to support this claim is not as strong as that presented for the effects mentioned above. Also, the premature condensation of chromosomes appears as a clear consequence of NPP-3 depletion, but this intriguing phenotype remains unexplored.

      The DNA damage evidence is based on lagging chromosomes in 2-cell stages, the HUS-1 reporter, and micronuclei in embryos at the 20-30 cell stage. It is noted that npp-3 RNAi is pleiotropic and also causes DNA damage. There is additional DNA damage caused by loss of MDF-1 and MDF-2 in npp-3 RNAi, but we agree that it is difficult to say whether the effect is additive or not, complicating the interpretation. Thus, we will tone down in our title to describe the dependency of the chromosomal nuclear periphery phenotype (see response to Reviewer 1 point 1).

      As for the premature chromosome condensation phenotype in NPP-3 depletion, we hypothesize it may result from accumulation of factors such as BAF-1 at the nuclear periphery, which could facilitate chromatin condensation. To confirm whether chromosome condensation precedes or follows relocation to the nuclear periphery, we could perform depletion of condensin II component, e.g. hcp-6, and see if lack of condensation affects chromosome relocalization (also see response to Reviewer 1 point 6).

      Reviewer #3 (Public review):

      Summary:

      This manuscript reports that RNAi depletion of the inner-ring nucleoporin NPP-3/NUP205 in Caenorhabditis elegans embryos causes nuclear envelope rupture, premature chromatin condensation, and relocalization of condensed prophase chromosomes to the nuclear periphery. Through a candidate epistasis screen, the authors argue that this relocalization requires spindle assembly checkpoint (SAC) components (MDF-1, MDF-2, SAN-1), inner kinetochore proteins (HCP- 3, HCP-4, and partially KNL-1), and NE rupture-repair factors (BAF-1, LEM-2), but not the CEC-4 heterochromatin- or SUN-1/POT-1 telomere-anchoring pathways. They further show that NPP-3 loss extends prophase and the NEBD-to-anaphase interval in a SAC-dependent manner, redistributes MDF-1/MDF-2, and reduces import of KNL-1/BUB-1/HCP-1. Codepletion of NPP-3 with MDF-1 abolishes both the arrest and the peripheral localization while increasing lagging chromosomes, HUS-1 foci, micronuclei, and lethality, which the authors interpret as evidence that peripheral positioning is protective.

      Weaknesses:

      (1) The "protective" conclusion is largely correlative. The protective claim rests on the observation that co-depleting MDF-1 (or MDF-2) with NPP-3 removes the peripheral localization and simultaneously increases DNA damage, micronuclei, and lethality. However, depleting a SAC component removes at least three things at once: the peripheral localization, the prophase extension, and the NEBD-to-anaphase arrest. , they unrelated to chromosome positioning, the current design cannot separate damage caused by loss of a protective peripheral location from damage caused by checkpoint bypass. As presented, the increased damage is at least as consistent with simple SAC bypass. To support the protective model, the authors should provide a manipulation that disrupts peripheral positioning without abrogating the SACdependent arrest (for example, via the BAF-1/LEM-2 or kinetochore depletion) and show that damage still increases. The LEM-2 co-depletion, which partially suppresses positioning, is a natural place to test whether micronuclei and HUS-1 foci also rise.

      We agree that the current data are largely correlative. In the early C. elegans embryos, single depletion of SAC components MDF-1 or MDF-2 does not affect the mitosis duration or chromosome segregation. When spindles are defective, the functional SAC delays progression through mitosis [5]. Depleting SAC components such as MDF-1 in npp-3 RNAi indeed impacts multiple processes, including prophase and prometaphase duration, chromosome repositioning and condensation, making it challenging to disentangle effects specifically to related chromosome positioning.

      To address this, future experiments involving NPP-3 LEM-2 co-depletion, which has been shown to partially impair peripheral chromosome positioning, will be utilized to assess whether disruption of partial peripheral localization results in increased DNA damage, micronuclei formation, or HUS-1 foci accumulation, independent of SAC function (see response to Reviewer 1 point 1). 

      (2) Knockdown efficiency of the partner gene in double RNAi is not verified. The double depletions are performed by cloning both gene fragments into a single vector. This risks reducing the effective dose of each dsRNA, so an apparent suppression in an npp-3; gene X (RNAi) condition could reflect weaker NPP-3 knockdown rather than a true epistatic relationship. The authors partially address this by showing that NPP-3::mCherry is still reduced in npp-3;mdf-1 (Figure S4A/B), which is helpful, but they do not demonstrate efficient knockdown of the partner genes in any double condition. For the key epistasis conclusions (MDF-1, MDF-2, HCP-3, HCP4 suppressions), the knockdown of the second gene should be independently validated with a reporter strain for the second protein.

      We acknowledge that in our double RNAi experiments, the knockdown efficiency of the npp-3 is checked by imaging (see response to Reviewer 1 point 8), whereas that of the second gene was not validated in each condition. To address this, we confirmed the effectiveness of certain partner gene depletions, e.g. HCP-3, by examining the levels of the respective proteins using available GFP-marked strains or immunofluorescence (Author response image 5). Additionally, for genes like HCP-3, KNL-1, and BUB-1, we also assessed the functional consequences on chromosome segregation, where severe defects observed (Fig. S4C) can support effective depletion. However, for some strains, we do not have GFP makers and will need to check the RNA levels. 

      Author response image 5.

      Representative confocal image of GFP::HCP-3 and mCherry::H2B at the NEBD time point of P1 cell in the control, single and double RNAi. Scale bar, 5 μm.

      (3) Alternative explanations for the transcriptomic and H3K9me3 data are not excluded. NPP-3 depletion blocks nuclear import of molecules smaller than ~70 kDa and arrests development at early gastrulation. Both the RNA-seq changes (30% of genes downregulated) and the increased H3K9me3 signal could therefore be secondary consequences of nucleocytoplasmic transport failure and developmental arrest rather than evidence of position-dependent transcriptional repression. Notably, the authors' own finding that up- and down-regulated genes show no chromosomal positional bias (Figure S2C/D) argues against a model in which peripheral repositioning drives silencing of specific chromatin domains. This section should be reframed more cautiously, with the transport/arrest confound explicitly discussed, and RNA-seq replicate number and differential-expression thresholds reported.

      We agree that these chromatin modifications and transcriptional alterations could be related to the nuclear transport failure and developmental delay in npp-3 disruption. We will discuss this possibility in results and discussion. 

      (4) Evidence for SAC "activation in prophase" is indirect, and the effect is small. The claim of a novel prophase role for the SAC rests on MDF-1/MDF-2 intensity changes that are repeatedly described as "modest," "slight," or "mild," measured with small n and Student's t-tests, together with phenotypic suppression of prophase extension. There is no direct readout of SAC catalytic activity (for example, MCC assembly). The prophase-extension suppression by MDF-1 is the strongest evidence; the intensity data are weak support. I recommend tempering "the SAC is activated in prophase" to a hypothesis, and strengthening it with a more direct assay if feasible.

      We agree that the evidence for SAC activation during prophase is indirect. The prophase extension (100 s) in npp-3 RNAi is a functional assay to support SAC activation, and the suppression in npp-3 mdf-1 double RNAi suggests dependency. The changes in MDF1/MDF-2 intensities are modest. Biochemical analyses of MCC assembly in C. elegans mixed cell cycle stage embryos is challenging to demonstrate SAC activity in prophase. 

      (5) The BAF-1 arm of the model is inferred rather than demonstrated. The authors state that baf- 1(RNAi) and npp-3;baf-1 produced clustering too severe for epistasis, so BAF-1's requirement for peripheral localization is not actually established genetically; it rests on increased BAF-1 accumulation (correlative) plus the LEM-2 partial suppression. The proposed BAF- 1/CENP-C bridge is extrapolated from Drosophila (ref. 71). This is reasonable as a discussion hypothesis but should not be presented in the abstract or summary model as an established dependency.

      We did not include the BAF-1/CENP-C bridge hypothesis in the abstract or the model, and we will discuss this as a speculative mechanism rather than an established dependency.

      References:

      (1) Rodenas, E., Gonzalez-Aguilera, C., Ayuso, C. & Askjaer, P. Dissection of the NUP107 nuclear pore subcomplex reveals a novel interaction with spindle assembly checkpoint protein MAD1 in Caenorhabditis elegans. Mol Biol Cell 23, 930-944 (2012).

      (2) Schumacher, J.M., Ashcroft, N., Donovan, P.J. & Golden, A. A highly conserved centrosomal kinase, AIR-1, is required for accurate cell cycle progression and segregation of developmental factors in Caenorhabditis elegans embryos. Development 125, 4391-4402 (1998).

      (3) Kotak, S., Afshar, K., Busso, C. & Gonczy, P. Aurora A kinase regulates proper spindle positioning in C. elegans and in human cells. J Cell Sci 129, 3015-3025 (2016).

      (4) Hachet, V. et al. The nucleoporin Nup205/NPP-3 is lost near centrosomes at mitotic onset and can modulate the timing of this process in Caenorhabditis elegans embryos. Molecular Biology of the Cell 23, 3111-3121 (2012).

      (5) Encalada, S.E., Willis, J., Lyczak, R. & Bowerman, B. A spindle checkpoint functions during mitosis in the early Caenorhabditis elegans embryo. Molecular Biology of the Cell 16, 1056-1070 (2005).

    1. eLife Assessment

      This useful study employs an innovative chemoproteomic and multi-modal experimental approach to support VDAC2 as a key functional mediator of STX effects on mitochondrial bioenergetics. However, concerns remain regarding the physiological and disease relevance, including the reliance on cell lines, lack of loss-of-function validation, unresolved binding specificity, and limited justification for focusing on POMC neurons in the context of Alzheimer's disease. As a result, the evidence is incomplete for a link between VDAC2 and Alzheimer's pathogenesis, and aspects of the binding data raise alternative interpretations, including a potential role for other VDAC isoforms.

    2. Reviewer #1 (Public review):

      Summary:

      In this study, Qiu et al. examine the effects of the estrogen mimic STX on mitochondrial function and its interaction with VDAC2 in PMOC neurons.

      Strengths:

      The authors employ a broad range of molecular, cellular, and chemoproteomic approaches with generally sound methodology.

      Weaknesses:

      The work suffers from major conceptual and experimental issues that substantially limit its scientific impact.

      Major Concerns

      (1) Lack of Rationale.<br /> The study provides no justification for investigating sex specific aspects of Alzheimer's disease by focusing on VDAC-mediated mitochondrial dysfunction in PMOC neurons. These hypothalamic neurons are not recognized as early or primary sites of AD vulnerability, making the biological premise unclear.

      (2) Weak Link to AD Pathogenesis.<br /> Although mitochondrial dysfunction is well established in AD, the authors do not convincingly demonstrate a mechanistic or pathological connection between VDAC2 and AD. VDACs are not established contributors to AD etiology, and the manuscript does not strengthen this association.

      (3) Unclear Relevance to AD Contexts.<br /> While the data support an interaction between STX and VDAC2 affecting mitochondrial parameters (ATP production, membrane potential, glycolysis, respiration) in PMOC neurons, the study does not show whether this mechanism is relevant to mitochondrial dysfunction in AD. No validation is provided in AD-related models or in contexts related to sex specific AD phenotypes.

      (4) Interpretation of Competitive Binding Data.<br /> The competitive binding results in Figure S4B are not adequately interpreted. The dose-dependent competition observed for VDAC3 suggests it may be a stronger candidate than VDAC2, yet this possibility is not addressed.

    3. Reviewer #2 (Public review):

      Summary:

      STX is a non-steroidal, CNS-selective estrogenic compound with neuroprotective effects in stroke and Alzheimer's disease models, but its molecular target has remained unknown for nearly 20 years. In this study, the authors identify VDAC proteins as the direct mitochondrial targets of STX using chemoproteomics, single-cell qPCR, electrophysiology, and metabolic flux analyses. They further show that VDAC2 is the primary functional target in female POMC neurons, linking STX-mediated VDAC modulation to enhanced mitochondrial bioenergetics and neuroprotection.

      Strengths:

      This study is strengthened by its innovative chemoproteomic approach, in which the authors developed a novel bifunctional STX probe (BF-STX) containing a photo-crosslinkable diazirine group and an alkyne handle to capture transient STX-protein interactions in living cells. The experimental design is further reinforced by rigorous controls, including no-UV negative controls and competition assays with excess unlabeled STX, which provide convincing evidence that VDAC1, VDAC2, and VDAC3 are genuine STX-binding targets rather than nonspecific artifacts. Finally, the authors validate the STX-VDAC interaction using multiple complementary approaches, including chemoproteomics, single-cell qPCR, planar lipid membrane electrophysiology, and Seahorse metabolic flux analyses, providing strong mechanistic support for their conclusions.

      Weaknesses:

      While the study provides convincing evidence that STX directly modulates VDAC function, several limitations remain. Most experiments were performed in immortalized cell lines rather than primary neurons or in vivo models, limiting their physiological relevance. In addition, the exact structural binding site of STX on VDAC remains unresolved, and no loss-of-function experiments (e.g., VDAC2 knockdown) were performed to establish a direct causal link between VDAC2 and STX's bioenergetic and neuroprotective effects. The non-linear dose-response at higher STX concentrations also requires further investigation.

    4. Author response:

      Reviewer #1:

      We thank Reviewer #1 for the thoughtful critique. We have revised the manuscript to clarify the physiological rationale for the experimental system, the potential relevance of mitochondrial STX–VDAC signaling to neurodegeneration, and the appropriate scope of our conclusions.

      (1) Lack of Rationale

      The study provides no justification for investigating sex-specific aspects of Alzheimer's disease by focusing on VDAC-mediated mitochondrial dysfunction in POMC neurons. These hypothalamic neurons are not recognized as early or primary sites of AD vulnerability, making the biological premise unclear.

      We thank the reviewer for raising this important point. We agree that the original manuscript did not sufficiently explain the rationale for using POMC neurons.

      Our rationale is primarily physiological rather than disease-specific. POMC neurons are a well-characterized, metabolically sensitive, and estrogen-responsive neuronal population in which membrane-initiated estrogen signaling and the actions of STX have been extensively characterized. They therefore provide a physiologically relevant neuronal system for identifying the molecular mechanisms through which STX regulates mitochondrial function.

      The potential relevance to neurodegeneration is supported by evidence that hypothalamic and POMC neuronal function can be disrupted in neurodegenerative disease models. Do and colleagues (2018) reported hypothalamic neurodegeneration, increased inflammatory and apoptotic markers, and reduced POMC neuronal populations in 3xTg-AD mice and further showed that exercise attenuated hypothalamic apoptosis and restored POMC neuronal populations (Do, Laing et al. 2018). In addition, Shen and colleagues (2016) demonstrated disruption of POMC/MC4R signaling in APP/PS1 mice and showed that restoration of this pathway improved synaptic function (Shen, Tian et al. 2016).

      These studies do not establish POMC neurons as a primary site of AD pathology, but they demonstrate that POMC-related neuronal systems can be vulnerable to neurodegenerative processes. This provides a broader biological context for investigating mitochondrial mechanisms in this neuronal population.

      There is also a strong physiological rationale for examining estrogen-sensitive mechanisms in POMC neurons. These neurons are established targets of 17β-estradiol and are highly responsive to metabolic and mitochondrial state. STX is a non-steroidal estrogenic compound that activates membrane-initiated estrogen signaling, and our previous studies demonstrated neuroprotective and mitochondrial effects of STX in a neurodegenerative disease model.

      Thus, POMC neurons were used because they provide a well-defined estrogen-responsive and metabolically sensitive neuronal population in which STX signaling can be mechanistically investigated—not because we consider them an initiating site of AD pathology.

      We have revised the Introduction and Discussion accordingly. The revised manuscript now emphasizes the physiological significance of STX–VDAC signaling for neuronal mitochondrial function and presents its potential relevance to neurodegeneration as an important area for future investigation.

      (2) Weak Link to AD Pathogenesis

      Although mitochondrial dysfunction is well established in AD, the authors do not convincingly demonstrate a mechanistic or pathological connection between VDAC2 and AD. VDACs are not established contributors to AD etiology, and the manuscript does not strengthen this association.

      We agree that the present experiments do not establish VDAC2 as an etiological or pathological driver of AD. This was not the objective of the present study.

      The primary goal was to identify the molecular target(s) through which STX influences mitochondrial function. Using BF-STX chemoproteomic capture and competition experiments together with single-cell gene-expression analysis, planar lipid membrane electrophysiology, and mitochondrial metabolic analyses, we identify VDAC proteins as mitochondrial targets of STX and demonstrate functional effects of STX on VDAC channel properties and mitochondrial bioenergetics.

      Our previous studies demonstrated neuroprotective and mitochondrial effects of STX in the 5xFAD model (Lee, Bostick et al. 2025), providing a neurodegenerative context that motivated the present mechanistic investigation. However, the current findings should not be interpreted as establishing VDAC2 as an AD pathogenic mechanism.

      We have revised the manuscript accordingly. The STX–VDAC interaction is now presented principally as a mitochondrial mechanism with potential relevance to neuronal physiology and neurodegeneration. Whether this pathway contributes to the neuroprotective actions of STX in disease models will require direct experimental testing.

      (3) Unclear Relevance to AD Contexts

      While the data support an interaction between STX and VDAC2 affecting mitochondrial parameters (ATP production, membrane potential, glycolysis, respiration) in POMC neurons, the study does not show whether this mechanism is relevant to mitochondrial dysfunction in AD. No validation is provided in AD-related models or in contexts related to sex-specific AD phenotypes.

      We agree that the present experiments do not directly establish the relevance of the STX–VDAC interaction to mitochondrial dysfunction in AD.

      The current study was designed to identify and characterize the molecular mechanism underlying the mitochondrial actions of STX, rather than to test this pathway in a specific neurodegenerative disease model. Our findings demonstrate that STX interacts with VDAC proteins, modifies VDAC channel properties, and alters mitochondrial bioenergetics.

      The study builds on our previous findings in the 5xFAD model, in which STX reduced amyloid-β-associated pathology and affected mitochondrial function (Lee, Bostick et al. 2025). The identification of VDAC proteins as STX targets therefore provides a mechanistic foundation for future studies examining whether this pathway contributes to neuronal protection under neurodegenerative conditions.

      We have revised the Discussion to acknowledge the absence of direct disease-model validation. Future studies using conditional or neuron-specific manipulation of VDAC isoforms will be important for determining the physiological and neuroprotective significance of STX–VDAC signaling in vivo.

      Accordingly, our conclusions now emphasize what is directly supported by the present experiments: STX interacts with VDAC proteins and modulates VDAC channel function and mitochondrial bioenergetics. The relevance of this mechanism to neurodegenerative disease remains to be established.

      (4) Interpretation of Competitive Binding Data

      The competitive binding results in Figure S4B are not adequately interpreted. The dose-dependent competition observed for VDAC3 suggests it may be a stronger candidate than VDAC2, yet this possibility is not addressed.

      We thank the reviewer for highlighting this important point. We agree that the original manuscript placed too much emphasis on VDAC2 based on the chemoproteomic data.

      Our chemoproteomic experiments identified VDAC1, VDAC2, and VDAC3 as STX-interacting proteins. Competition with unlabeled STX produced a particularly clear reduction in VDAC3 labeling, including complete loss of the VDAC3 signal at three molar equivalents of unlabeled STX. We therefore agree that VDAC3 represents an important candidate STX target.

      We have revised the Results and Discussion so that the competition experiment is no longer interpreted as demonstrating preferential or exclusive binding to VDAC2.

      However, the persistence of VDAC1 and VDAC2 labeling may also be influenced by properties of the BF-STX photoprobe as alkyl diazirine probes can preferentially photolabel membrane proteins (Kleiner, Heydenreuter et al. 2017). We now present this only as a possible technical consideration rather than an explanation established by our data.

      Our subsequent emphasis on VDAC2 was based on the integration of several observations. Single-cell qPCR demonstrated that Vdac2 is the predominant transcript in the native POMC neurons examined (revised Figure 3B), with an approximate expression hierarchy of Vdac2 > Vdac3 >> Vdac1. In addition on a technical note, recombinant VDAC3 is more difficult to reconstitute reliably into artificial membranes because of its lower stability in detergent.

      The revised manuscript therefore recognizes all three VDAC isoforms as candidate STX targets but does not claim that VDAC2 is the exclusive or preferential target. Direct comparisons of STX binding and functional modulation among the three isoforms will be required to establish isoform selectivity.

      Reviewer #2:

      We thank Reviewer #2 for the positive assessment of our chemoproteomic strategy and multidisciplinary characterization of the STX–VDAC interaction. We also appreciate the reviewer’s identification of important limitations and directions for future investigation.

      (1) Physiological Relevance of Cell-Line Experiments

      Most experiments were performed in immortalized cell lines rather than primary neurons or in vivo models, limiting their physiological relevance.

      We agree that the use of immortalized neuronal cell models represents an important limitation.

      However, these models provided the cellular material, reproducibility, and experimental control required for chemoproteomic target capture and Seahorse metabolic measurements. Importantly, we complemented these studies with single-cell analysis of native POMC neurons and electrophysiological characterization of recombinant VDAC channels.

      Nevertheless, these approaches do not substitute for direct demonstration of STX–VDAC signaling in primary POMC neurons or in vivo. We have therefore revised the Discussion to emphasize that establishing the physiological significance of this mitochondrial pathway will require validation in primary neuronal preparations and whole-animal models.

      (2) Structural Binding Site of STX on VDAC

      The exact structural binding site of STX on VDAC remains unresolved.

      We agree. BF-STX chemoproteomics identifies proteins interacting with STX in a cellular environment but does not resolve the amino acid residues or structural pocket responsible for binding.

      We have clarified this limitation in the Discussion. Determining the STX-binding site will require complementary approaches such as targeted mutagenesis, direct binding measurements with purified VDAC proteins, and structural studies in membrane-like environments, including lipid nanodiscs. Such studies should also determine whether STX recognizes a conserved feature among VDAC isoforms or exhibits isoform selectivity.

      (3) Lack of VDAC2 Loss-of-Function Experiments

      No loss-of-function experiments (e.g., VDAC2 knockdown) were performed to establish a direct causal link between VDAC2 and STX's bioenergetic and neuroprotective effects.

      We agree that VDAC loss-of-function experiments would provide an important additional test of causality.

      Our present evidence is convergent: chemoproteomics identifies VDAC proteins as STX-interacting targets; single-cell analyses demonstrate VDAC expression in POMC neurons; electrophysiological studies show that STX modifies VDAC channel properties; and metabolic analyses demonstrate STX-dependent changes in mitochondrial bioenergetics. Together, these findings support a STX–VDAC mechanism but do not establish that VDAC2 alone is necessary for the mitochondrial or neuroprotective actions of STX.

      We have revised the manuscript accordingly and explicitly identify the absence of loss-of-function experiments as a limitation.

      Because whole-body VDAC2 loss-of-function is associated with severe developmental consequences, future studies will require conditional or neuron-specific approaches. Parallel evaluation of VDAC1 and VDAC3 will also be important because all three isoforms were identified by chemoproteomics and potential functional redundancy may complicate single-isoform manipulations.

      (4) Non-linear Dose-Response at Higher STX Concentrations

      The non-linear dose-response at higher STX concentrations also requires further investigation.

      We agree that the non-linear concentration-response warrants further investigation and have revised the Discussion to avoid interpreting the STX response as a simple monotonic concentration-response relationship.

      One possibility is that STX engages multiple mitochondrial targets with different apparent affinities and opposing effects on respiration. At lower concentrations, STX may preferentially engage a higher-affinity target, potentially VDAC, whereas higher concentrations may recruit lower-affinity targets that constrain this response. Consistent with this possibility, our BF-STX dataset (Supplemental Tables) identified several mitochondrial proteins involved in oxidative phosphorylation and metabolite transport, including ATP5F1C, NNT, SLC25A4, SLC25A5 and SLC25A3. ATP5F1C is required for efficient mitochondrial ATP production (Fiorillo, Scatena et al. 2021), and estrogenic regulation of ATP synthase has been reported (Massart, Paolini et al. 2002, Moreno, Moreira et al. 2013). NNT, SLC25A4, SLC25A5 and SLC25A3 also regulate mitochondrial respiration, redox balance, and ATP production (Mayr, Merkel et al. 2007, Lopert and Patel 2014). However, BF-STX enrichment does not establish direct STX binding, relative affinity, or functional modulation of these proteins. We therefore present the multi-target explanation only as a hypothesis. Direct binding and concentration-dependent target-engagement studies will be required to determine whether the non-linear response reflects recruitment of a lower-affinity mitochondrial target, concentration-dependent effects on VDAC itself, or downstream mitochondrial feedback.

      We thank the editors and reviewers again for their constructive comments. The revised manuscript more clearly defines the physiological rationale for the experimental system, establishes the scope of the STX–VDAC mitochondrial mechanism supported by our data, and places its potential relevance to neurodegeneration in an appropriately forward-looking context.

      References cited in the response

      Do, K., B. T. Laing, T. Landry, W. Bunner, N. Mersaud, T. Matsubara, P. Li, Y. Yuan, Q. Lu and H. Huang (2018). "The effects of exercise on hypothalamic neurodegeneration of Alzheimer's disease mouse model." PLoS One 13(1): e0190205.

      Fiorillo, M., C. Scatena, A. G. Naccarato, F. Sotgia and M. P. Lisanti (2021). "Bedaquiline, an FDA-approved drug, inhibits mitochondrial ATP production and metastasis in vivo, by targeting the gamma subunit (ATP5F1C) of the ATP synthase." Cell Death Differ 28(9): 2797-2817.

      Kleiner, P., W. Heydenreuter, M. Stahl, V. S. Korotkov and S. A. Sieber (2017). "A Whole Proteome Inventory of Background Photocrosslinker Binding." Angew Chem Int Ed Engl 56(5): 1396-1401.

      Lee, H.-J., Z. Bostick, J. Doherty, T. L. Swanson, M. J. Kelly, J. F. Quinn, N. E. Gray and P. F. Copenhaver (2025). "Neuroprotection against beta-amyloid toxicity by the novel estrogen receptor modulator STX requires convergent signaling pathways." Frontiers in Molecular Neuroscience Volume 18 - 2025.

      Lopert, P. and M. Patel (2014). "Nicotinamide nucleotide transhydrogenase (Nnt) links the substrate requirement in brain mitochondria for hydrogen peroxide removal to the thioredoxin/peroxiredoxin (Trx/Prx) system." J Biol Chem 289(22): 15611-15620.

      Massart, F., S. Paolini, E. Piscitelli, M. L. Brandi and G. Solaini (2002). "Dose-dependent inhibition of mitochondrial ATP synthase by 17 beta-estradiol." Gynecol Endocrinol 16(5): 373-377.

      Mayr, J. A., O. Merkel, S. D. Kohlwein, B. R. Gebhardt, H. Böhles, U. Fötschl, J. Koch, M. Jaksch, H. Lochmüller, R. Horváth, P. Freisinger and W. Sperl (2007). "Mitochondrial phosphate-carrier deficiency: a novel disorder of oxidative phosphorylation." Am J Hum Genet 80(3): 478-484.

      Moreno, A. J., P. I. Moreira, J. B. Custódio and M. S. Santos (2013). "Mechanism of inhibition of mitochondrial ATP synthase by 17β-estradiol." J Bioenerg Biomembr 45(3): 261-270.

      Shen, Y., M. Tian, Y. Zheng, F. Gong, A. K. Y. Fu and N. Y. Ip (2016). "Stimulation of the Hippocampal POMC/MC4R Circuit Alleviates Synaptic Plasticity Impairment in an Alzheimer's Disease Model." Cell Rep 17(7): 1819-1831.

    1. eLife Assessment

      This revised study offers valuable insights into how nuclear export influences protein condensates and TDP-43 phase behaviour. The findings are solid and highlight several noteworthy observations for the field, such as RNA-dependent stabilisation of TDP-43 condensates and the inhibition of nuclear export in an ALS organoid model. The work suggests a possible mechanistic link between nuclear export and TDP-43 aggregation in ALS/FTD; however, the exact mechanistic connection remains unclear, as much of the evidence is based on indirect observations within sensitised model systems. Although the authors carefully acknowledge these limitations and moderate their conclusions, further validation in more physiologically relevant models will be necessary to demonstrate a direct causal role of nuclear export in regulating pathological TDP-43 aggregation.

    2. Reviewer #1 (Public review):

      [Editors' note: this version has been assessed by the Reviewing Editor without further input from the original reviewers. The authors have addressed the comments raised in the previous round of review.]

      This revised manuscript represents a partial response to the concerns raised in the first round of review. The authors have made one genuine mechanistic addition in the form of the semi-permeabilized cell reconstitution assay, removed the most overreaching conclusions regarding the contribution of cytoplasmic TDP-43 aggregation to disease, and made several minor presentational improvements. However, the central weaknesses of the original submission remain substantially unaddressed. The exclusive reliance on non-physiological TDP-43 variants, the incompletely resolved mechanism linking XPO1 to TDP-43 phase behavior, and the limited organoid validation continue to limit confidence in the major claims. The authors have, in several instances, responded by removing contested data rather than by providing the additional evidence that was requested.

      (1) The justification for the 2KQ acetylation-mimetic system remains inadequate.<br /> The authors respond to the concern about the non-physiological nature of the 2KQ mutant by citing published evidence that TDP-43 acetylation occurs in ALS patient spinal cord and is upregulated under oxidative and proteotoxic stress conditions. While these references are real and support the relevance of acetylation as a pathological post-translational modification, they do not resolve the central concern: there is no quantification of how much endogenous TDP-43 is acetylated at the specific lysine residues mimicked by 2KQ in degenerating human neurons, and no evidence that the degree of RNA-binding disruption imposed by the double glutamine substitution is ever achieved by endogenous acetylation in vivo. The 2KQ mutant eliminates RNA binding essentially completely, whereas physiological acetylation events are graded, reversible, and likely partial. The response conflates the existence of TDP-43 acetylation as a phenomenon with validation that 2KQ is a physiologically accurate model of that phenomenon. None of the new experiments address the request to test whether wild-type TDP-43 expressed at near-physiological levels, or a bona fide heterozygous ALS-linked TARDBP mutant in iPSC-derived neurons, responds to XPO1 modulation in a qualitatively similar fashion. Until this is shown, the mechanistic conclusions of this paper remain constrained to a highly artificial overexpression system and cannot be extrapolated to physiological or pathological TDP-43 biology with confidence.

      (2) The homozygous K181E organoid model is still not adequately justified, and no heterozygous comparison has been provided.<br /> The authors acknowledge that the homozygous background is "more sensitive for detecting phospho-TDP-43" and argue that homozygous conditions are commonly used in experimental TDP-43 research. However, the critical issue is not whether homozygous models are used in general, but whether the homozygous background specifically alters the relative contribution of cytoplasmic aggregation versus nuclear RNA-processing dysfunction in this study. In a homozygous K181E model, both alleles produce an RNA-binding-defective TDP-43, meaning that every molecule of endogenous TDP-43 in the cell is dysfunctional. This is categorically different from the patient situation in which one wild-type allele is present, and it may substantially exaggerate nuclear loss-of-function relative to cytoplasmic gain-of-function phenotypes. The authors have not performed the requested comparison with heterozygous K181E/+ organoids, nor have they acknowledged that the organoid genotype itself could bias the interpretation of what KPT-276 treatment rescues. Given that the organoid section is now the sole in-disease-model validation of the XPO1 mechanism, this limitation is more consequential than it was in the original submission.

      (3) The new semi-permeabilized cell data is a genuine contribution, but the mechanistic interpretation remains insufficiently constrained.<br /> The development of the streptolysin O semi-permeabilized cell reconstitution system is the most substantive new addition to this revision. The finding that LMB-stabilized anisosomes resist cytosol washout but dissolve upon RNase T1 treatment is interesting and provides a plausible indirect mechanism: XPO1 inhibition retains nuclear RNA, and this elevated nuclear RNA availability contributes to maintaining the liquid LLPS state of the TDP-43 2KQ condensate. This is a meaningful mechanistic advance and deserves credit. However, several important limitations of this new data are not adequately discussed. First, RNase T1 degrades single-stranded RNA globally during permeabilization, so the experiment does not identify which specific RNA species stabilize the anisosome, nor whether these are pre-mRNA splicing intermediates, mature mRNA, non-coding RNA, or another class. Second, the same nuclear export blockade that retains RNA will also retain the nuclear concentrations of many RNA-binding proteins, splicing factors, and other XPO1-dependent cargos. The RNase T1 experiment does not exclude the possibility that the relevant effect is mediated by an RNA-binding protein whose nuclear concentration increases upon LMB treatment and which, upon RNase digestion, can no longer engage TDP-43 or the anisosome shell. Third, the permeabilized cell system is by definition not intact and has lost cytosolic factors; whether the RNA-dependent stabilization of anisosomes operates in the same way in intact cells during physiological or pathological nuclear export perturbation is an assumption, not a demonstrated fact. The authors should more carefully frame these data as hypothesis-generating and explicitly note these alternative interpretations in the Discussion.

      (4) The conceptual asymmetry between XPO1 inhibition and XPO1 overexpression phenotypes is not resolved by the new mechanism.<br /> The paper continues to present two XPO1 perturbation phenotypes that are difficult to reconcile within a single mechanistic model. XPO1 inhibition enlarges anisosomes, maintains their liquid character by FRAP, and retains them in the nucleus. XPO1 overexpression also enlarges TDP-43 puncta, but these are FRAP-impaired, gel-like, and appear in the cytoplasm. The RNA-retention model proposed by the new semi-permeabilized data explains why XPO1 inhibition stabilizes the liquid state, but it does not explain why XPO1 overexpression drives the opposite outcome: gel-like hardening and cytoplasmic redistribution. If increased nuclear RNA availability is the key variable downstream of XPO1 inhibition, then XPO1 overexpression would be expected to decrease nuclear RNA and thereby destabilize anisosomes toward dissolution or hardening. The paper does not test whether nuclear RNA levels are indeed altered by XPO1 overexpression, nor whether the cytoplasmic gel-like puncta seen in XPO1-overexpressing cells are RNA-poor relative to control anisosomes. The revised Discussion does not engage with this asymmetry in a satisfying way, and the figure model remains qualitative. A quantitative or at least semi-quantitative model that accounts for both arms of the XPO1 perturbation is needed.

      (5) The removal of RNA-seq data weakens rather than strengthens the organoid section.<br /> The authors have removed the bulk RNA-seq analysis from the revised manuscript in response to concerns that the modest transcriptional rescue was being over-interpreted. While the decision to remove over-interpretation is appropriate, the result is that the organoid section now rests entirely on pTDP-43 immunostaining as its sole readout. The revised paper thus uses reduction in immunofluorescent pTDP-43 puncta in homozygous K181E organoids as the only evidence that nuclear export inhibition mitigates TDP-43 proteinopathy in a disease-relevant context. This is a weaker evidentiary base than before the revision, not an improvement. The originally requested more sensitive orthogonal readouts, including biochemical fractionation for SDS-insoluble TDP-43, filter-trap assays, or RNA aptamer-based detection of TDP-43 aggregates, remain absent. Without at least one additional independent measure confirming that cytoplasmic TDP-43 aggregation is genuinely reduced rather than simply rendered antigenically undetectable, the organoid conclusion is not adequately supported. At minimum, the authors should provide total and cytoplasmic TDP-43 fractionation data from organoid lysates to corroborate the immunostaining result.

      (6) No functional neuronal readout has been provided for the organoid model.<br /> The organoid section now makes the claim that "nuclear export is required for the formation of p-TDP-43-containing aggregates in a disease-relevant organoid model," but no measure of neuronal health, integrity, or function is reported in association with this. Even a simple assessment of neuron survival by TUJ1 or MAP2 quantification, neurite complexity, or cleaved caspase-3 staining before and after KPT-276 treatment would substantially strengthen the biological significance of the pTDP-43 reduction. The current data establish a pharmacological effect on a pathological marker but do not demonstrate that this has any consequence for neuronal biology in the organoid, which is what the disease-relevance framing implies.

      (7) The abstract and title continue to overstate the mechanistic conclusions.<br /> Despite the stated intent to reframe the study as a screening study and to temper the conclusions, the revised abstract retains the language: "These findings establish nuclear export as a key regulator of TDP-43 phase transitions and define a mechanistic framework that links altered nuclear transport and phase dynamics to TDP-43 aggregation potential." Similarly, the Discussion still states: "a particularly compelling aspect of our study is the discovery that the nuclear export receptor XPO1 governs TDP-43 liquid-to-solid transitions and subcellular localization." The word "governs" and the phrase "establish nuclear export as a key regulator" are not warranted by data that derive entirely from an overexpressed acetylation-mimetic mutant in a colon cancer cell line and a homozygous K181E organoid model. A more accurate framing would describe these findings as identifying nuclear export as one of several cellular processes that modulate TDP-43 phase behavior in a sensitized model system, with an indirect RNA-mediated mechanism that remains to be defined at the molecular level. The title change from "governs" to "modulates" is appreciated but does not extend into the abstract and Discussion, where the strong causal language persists.

      (8) Individual siRNA knockdown validation for XPO1 has not been provided.<br /> The authors argue that validation with 6 independent siRNAs across two rounds of screening, combined with convergent pharmacological data, is sufficient to establish XPO1 as a genuine hit. While the convergence of chemical and genetic evidence is reassuring, the specific request was for protein-level confirmation of XPO1 knockdown efficiency in the DLD1 TDP-43 2KQ cells used for mechanistic follow-up, together with demonstration that the anisosome phenotype is specifically caused by loss of XPO1 and not by off-target effects. This is a straightforward experiment, and its absence is particularly notable given that the entire mechanistic XPO1 narrative hinges on this specificity. At minimum, an immunoblot confirming XPO1 protein depletion in cells treated with the siRNA pool identified in the screen, in the same cell background and induction conditions as the follow-up experiments, should be provided.

      (9) The identity of XPO1-dependent cargos that regulate anisosome dynamics remains entirely unknown.<br /> The authors acknowledge that XPO1 does not directly bind TDP-43 and that the mechanism is likely indirect. The new RNA data provides one plausible indirect pathway. However, the possibility that one or more specific RNA-binding proteins or splicing factors, whose nuclear levels rise upon XPO1 inhibition, are the proximate drivers of anisosome stabilization has not been addressed. This matters because if the relevant mechanism operates through a specific cargo rather than bulk RNA retention, the model for how nuclear export connects to TDP-43 aggregation in disease would be fundamentally different. The authors decline to pursue adaptor identification on grounds of scope, which is a defensible position for future work. However, the framing should explicitly state that the current data cannot distinguish between bulk RNA retention and cargo-specific effects, and that the conclusion that nuclear export modulates TDP-43 phase behavior via RNA accumulation is a working hypothesis supported by but not proven by the RNase T1 experiment.

      Minor remaining issues.

      The number of independent iPSC clones and organoid batches used for the KPT-276 treatment experiment is now stated as two batches per condition, which is minimal for a 3D organoid study and does not fully address the concern about clone-level variability. Ideally, organoids from at least two independently derived isogenic clones per genotype would be used. The mCherry overexpression control added in Supplemental Figure 4 is a useful addition and is acknowledged. The immunoblotting confirmation that drug treatments do not alter total TDP-43 levels addresses a prior concern adequately. The addition of the sentence noting that anisosomes have not been validated in human patient samples is appreciated and appropriate. Statistical detail has been improved in figure legends. These minor improvements are noted positively but do not compensate for the major unresolved concerns above.

    3. Reviewer #2 (Public review):

      This manuscript addresses an important and timely question in TDP-43 biology by systematically identifying regulators of TDP-43 anisosome formation, with a particular focus on nuclear export via XPO1. Using a combination of unbiased chemical screening, genetic perturbation, and advanced imaging approaches, the authors propose that inhibition of nuclear export modulates the abundance and biophysical properties of TDP-43 anisosomes. They further strengthen their findings by introducing an additional model system, a semi-permeabilized in vitro assay, which provides mechanistic evidence that XPO1 activity prevents anisosome dissolution by retaining nuclear RNAs. The study is conceptually innovative and has potential relevance for neurodegenerative diseases characterized by TDP-43 pathology. Some minor concerns remain, mostly about experimental design of the newly added data.

      Strengths:

      (1) The study employs an unbiased, hypothesis-free compound screen to identify regulators of TDP-43 anisosome formation, which is a major strength and reduces confirmation bias.

      (2) The authors combine chemical and genetic screening approaches, providing orthogonal validation of key pathways and increasing confidence in the biological relevance of top hits.

      (3) The focus on biophysical properties of TDP-43 assemblies, assessed through imaging and FRAP, moves beyond simple presence/absence of aggregates and provides mechanistic insight into the biophysical states of TDP-43.

      (4) The use of multiple experimental modalities, including live-cell imaging, FRAP, pharmacological perturbation, and transcriptomic analysis, reflects a technically sophisticated and ambitious study design.

      (5) The authors attempt to extend findings beyond immortalized cancer cell lines by incorporating organoid models, demonstrating awareness of disease relevance and translational importance.

      (6) The authors extend their study by incorporating a semi-permeabilized in vitro system, which provides compelling evidence that inhibition of nuclear export promotes the retention of nuclear anisosomes, an effect driven by the accumulation of nuclear RNAs.

      Overall, the manuscript is clearly written and logically structured, making complex experimental workflows accessible and the central hypotheses easy to follow.

      Weaknesses:

      (1) The manuscript has significantly improved with the revisions. Some experimental procedures and method details, as well has statements remain incompletely described:<br /> a) What is the smear in Figure S1 after VLX treatment?<br /> b) The authors state that "The reduction in TDP-43 signal was not due to protein elimination.", however no data is provided to support that statement.<br /> c) The authors state that "TDP-43 shifts from phase-separated state to a soluble state ...", however no data is provided to support that statement.<br /> d) Why did the authors choose cow lover cytosol for this study?<br /> e) The experimental setup for supplementing with cytosol/ATP/GTP is unclear. A more detailed schematic would be helpful to understand at what stage in the experiment these factors were added. Which step of the protocol was performed at 37 {degree sign}C, which is indicated in the figure schematic but not described in the methods.<br /> f) In the organoid model, the authors mention that they observe similar levels of total TDP-43, however they do not provide quantification. Instead, they provide a graph that shows highly significant changes in nuclear TDP-43, which was not addressed in the text.

      Additionally, some questions remain unclear:

      (1) The anisosomes induced by ATP/GTP or cytosol are insufficiently characterized. It remains unclear whether these structures correspond to canonical ring-shaped anisosomes, and whether they exhibit dynamic (liquid-like) or more static (gel-like) properties.

      (2) The contribution of the cytosol and ATP/GTP supplementation experiments to the overall narrative is unclear. While the findings are intriguing, their interpretation within the context of the study is not well articulated. In particular, the rationale for including cytosol is not sufficiently justified, given that ATP/GTP alone induces a pronounced effect, whereas cytosol alone does not.

      (3) The authors should address why endogenous XPO1 does not co-localize with anisosomes, whereas overexpressed XPO1 does. This raises the possibility that the observed co-localization may be an artifact of non-physiological protein levels, which should be discussed.

      (4) The iPSC-based model remains insufficiently characterized. While the authors propose that this system recapitulates the accumulation of liquid and solid aggregates resembling anisosomes, it is unclear whether this phenotype is robustly observed and whether KPT treatment effectively modulates it.

      (5) The rationale for the selected treatment durations is unclear, and the timing appears inconsistent across experiments (ranging from 3 to 16 hours), including within experiments involving the same compound. This variability should be justified or standardized.

      (6) Several figure legends require clarification:<br /> a) In the section stating "Collectively, our results suggest that the stability and dynamics of anisosomes are modulated by XPO1-mediated nuclear export ...", the cited figure appears to be incorrect. This should refer to Figure 5L rather than Figure 5J.<br /> b) Figure 1B: Please specify the number of replicates per concentration, the number of cells analyzed, and the model used for regression analysis. Additionally, the legend indicates a treatment duration of 15 hours, whereas Figure 1A states 24 hours.<br /> c) Figure 2G: The authors state "7 anisosomes per condition," but the graph displays only 4-6 data points. Please clarify what each data point represents.<br /> d) Figures 3B and 3G: Please clarify whether a defined threshold was used to determine a "reduction in anisosome number."<br /> e) Figure 4B: These do not represent biological replicates, as all samples derive from a single cell line; rather, they constitute independent experimental replicates.<br /> f) Figures 5B and 5H: The legend states "n = 3 biological repeats," but the number of data points shown appears higher. Please clarify.<br /> g) Figures 5K, 6C, and 6E: "Mean Fluorescence Intensity (MPI)" should be corrected to "MFI."<br /> h) Figure 6C: Please include the number of cells analyzed and provide relevant statistical measures (e.g., R², p-value).<br /> i) Figure 6D: The experimental timeline is unclear. Please specify the duration of incubation and the timing of each step.<br /> j) Figure 7B: Improved labeling is needed (e.g., clarification of "mean spot volume") to better align with the figure legend.

    4. Reviewer #3 (Public review):

      Summary:

      TDP-43 proteinopathy is broadly found in neurodegenerative diseases. This manuscript investigates how nuclear export influences the biophysical properties of TDP-43. The authors use a combination of chemical screening and genome-wide siRNA screening to identify pathways that modulate TDP-43 liquid-to-solid transitions. Overall, the study employs a broad array of approaches and addresses an important question in TDP-43 pathobiology. The identification of nuclear export as a central regulator is compelling and conceptually aligns with the emerging view that TDP-43 nucleocytoplasmic trafficking is a major defect in neurodegeneration.

      Strengths:

      This work integrates chemical and genetic screening to identify novel modifiers. The candidates were validated in both reporter cell lines and iPS-differentiated organoids. The findings support the nucleocytoplasmic transport is important for the biophysical properties of TDP-43.

      Comments on revised version.

      The manuscript has been improved with more data and clarification. The RNase T1 treatment experiment suggests that RNA is required for anisosome integrity. However, this does not directly demonstrate LMB increases nuclear RNA availability as changes in protein composition or other RNA-dependent mechanisms may also contribute. The conclusion and discussion need to be edited to consider these alternative scenarios. Overall, as most of the evidence remains indirect, the manuscript should avoid overinterpretation regarding the mechanisms underlying TDP-43 phase transition and aggregation.

    5. Author response:

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #1 (Public review):

      This revised manuscript represents a partial response to the concerns raised in the first round of review. The authors have made one genuine mechanistic addition in the form of the semi-permeabilized cell reconstitution assay, removed the most overreaching conclusions regarding the contribution of cytoplasmic TDP-43 aggregation to disease, and made several minor presentational improvements. However, the central weaknesses of the original submission remain substantially unaddressed. The exclusive reliance on non-physiological TDP-43 variants, the incompletely resolved mechanism linking XPO1 to TDP-43 phase behavior, and the limited organoid validation continue to limit confidence in the major claims. The authors have, in several instances, responded by removing contested data rather than by providing the additional evidence that was requested.

      (1) The justification for the 2KQ acetylation-mimetic system remains inadequate.

      The authors respond to the concern about the non-physiological nature of the 2KQ mutant by citing published evidence that TDP-43 acetylation occurs in ALS patient spinal cord and is upregulated under oxidative and proteotoxic stress conditions. While these references are real and support the relevance of acetylation as a pathological post-translational modification, they do not resolve the central concern: there is no quantification of how much endogenous TDP-43 is acetylated at the specific lysine residues mimicked by 2KQ in degenerating human neurons, and no evidence that the degree of RNA-binding disruption imposed by the double glutamine substitution is ever achieved by endogenous acetylation in vivo. The 2KQ mutant eliminates RNA binding essentially completely, whereas physiological acetylation events are graded, reversible, and likely partial. The response conflates the existence of TDP-43 acetylation as a phenomenon with validation that 2KQ is a physiologically accurate model of that phenomenon. None of the new experiments address the request to test whether wild-type TDP-43 expressed at near-physiological levels, or a bona fide heterozygous ALS-linked TARDBP mutant in iPSC-derived neurons, responds to XPO1 modulation in a qualitatively similar fashion. Until this is shown, the mechanistic conclusions of this paper remain constrained to a highly artificial overexpression system and cannot be extrapolated to physiological or pathological TDP-43 biology with confidence.

      We agree with the reviewer that the TDP-43 2KQ mutant is a non-physiological variant. To address this concern, we have removed all statements that extrapolate our findings to disease pathogenesis. As reflected in the revised title, we now present this study as an investigation of factors that modulate TDP-43 phase transition and aggregation using a sensitized model system, rather than as a direct disease model. The Abstract and Discussion have also been revised accordingly.

      The choice of the 2KQ mutant was dictated by the requirements of the screening strategy. To identify modulators of TDP-43 phase transition, it was necessary to use a TDP-43 variant that reliably undergoes phase separation in cells within an experimentally practical timeframe. In this context, we believe the 2KQ mutant provides a suitable and justified experimental tool. We have revised the manuscript to clearly distinguish observations made with this engineered construct from conclusions regarding physiological or disease-associated TDP-43.

      (2) The homozygous K181E organoid model is still not adequately justified, and no heterozygous comparison has been provided.

      The authors acknowledge that the homozygous background is "more sensitive for detecting phospho-TDP-43" and argue that homozygous conditions are commonly used in experimental TDP-43 research. However, the critical issue is not whether homozygous models are used in general, but whether the homozygous background specifically alters the relative contribution of cytoplasmic aggregation versus nuclear RNA-processing dysfunction in this study. In a homozygous K181E model, both alleles produce an RNA-binding-defective TDP-43, meaning that every molecule of endogenous TDP-43 in the cell is dysfunctional. This is categorically different from the patient situation in which one wild-type allele is present, and it may substantially exaggerate nuclear loss-of-function relative to cytoplasmic gain-of-function phenotypes. The authors have not performed the requested comparison with heterozygous K181E/+ organoids, nor have they acknowledged that the organoid genotype itself could bias the interpretation of what KPT-276 treatment rescues. Given that the organoid section is now the sole in-disease-model validation of the XPO1 mechanism, this limitation is more consequential than it was in the original submission.

      We agree with the reviewer and have removed all speculative statements regarding the relative contributions of cytoplasmic aggregation and RNA splicing defects to disease pathogenesis. The organoid section has also been revised to focus solely on the experimental findings. Specifically, we only present evidence that inhibition of nuclear export in a sensitized organoid model promotes the accumulation of cytoplasmic phosphorylated TDP-43 without making broader claims regarding its role in disease pathogenesis.

      (3) The new semi-permeabilized cell data is a genuine contribution, but the mechanistic interpretation remains insufficiently constrained.

      The development of the streptolysin O semi-permeabilized cell reconstitution system is the most substantive new addition to this revision. The finding that LMB-stabilized anisosomes resist cytosol washout but dissolve upon RNase T1 treatment is interesting and provides a plausible indirect mechanism: XPO1 inhibition retains nuclear RNA, and this elevated nuclear RNA availability contributes to maintaining the liquid LLPS state of the TDP-43 2KQ condensate. This is a meaningful mechanistic advance and deserves credit. However, several important limitations of this new data are not adequately discussed. First, RNase T1 degrades single-stranded RNA globally during permeabilization, so the experiment does not identify which specific RNA species stabilize the anisosome, nor whether these are pre-mRNA splicing intermediates, mature mRNA, non-coding RNA, or another class. Second, the same nuclear export blockade that retains RNA will also retain the nuclear concentrations of many RNA-binding proteins, splicing factors, and other XPO1-dependent cargos. The RNase T1 experiment does not exclude the possibility that the relevant effect is mediated by an RNA-binding protein whose nuclear concentration increases upon LMB treatment and which, upon RNase digestion, can no longer engage TDP-43 or the anisosome shell. Third, the permeabilized cell system is by definition not intact and has lost cytosolic factors; whether the RNA-dependent stabilization of anisosomes operates in the same way in intact cells during physiological or pathological nuclear export perturbation is an assumption, not a demonstrated fact. The authors should more carefully frame these data as hypothesis-generating and explicitly note these alternative interpretations in the Discussion.

      We have now added some sentences on page 11 to acknowledge the limitation of our experiments. It reads as “However, our data does not exclude other RNA species or RNA-binding proteins as anisosome stabilizer. Whether RNA-dependent stabilization of anisosomes operates in the same way in intact cells also requires further validation.”

      (4) The conceptual asymmetry between XPO1 inhibition and XPO1 overexpression phenotypes is not resolved by the new mechanism.

      The paper continues to present two XPO1 perturbation phenotypes that are difficult to reconcile within a single mechanistic model. XPO1 inhibition enlarges anisosomes, maintains their liquid character by FRAP, and retains them in the nucleus. XPO1 overexpression also enlarges TDP-43 puncta, but these are FRAP-impaired, gel-like, and appear in the cytoplasm. The RNA-retention model proposed by the new semi-permeabilized data explains why XPO1 inhibition stabilizes the liquid state, but it does not explain why XPO1 overexpression drives the opposite outcome: gel-like hardening and cytoplasmic redistribution. If increased nuclear RNA availability is the key variable downstream of XPO1 inhibition, then XPO1 overexpression would be expected to decrease nuclear RNA and thereby destabilize anisosomes toward dissolution or hardening. The paper does not test whether nuclear RNA levels are indeed altered by XPO1 overexpression, nor whether the cytoplasmic gel-like puncta seen in XPO1-overexpressing cells are RNA-poor relative to control anisosomes. The revised Discussion does not engage with this asymmetry in a satisfying way, and the figure model remains qualitative. A quantitative or at least semi-quantitative model that accounts for both arms of the XPO1 perturbation is needed.

      We thank the reviewer for this point. We have now explicitly mentioned in the discussion that the effect of XPO-1 on anisosome dynamics is likely mediated by an indirect mechanism. We also acknowledge that we do not fully understand why over-expression of XPO-1 causes TDP-43 to accumulate in gel-like structures in the cytoplasm. Although we did not check whether overexpression of XPO1 increases cargo export, we cited studies showing that over-expressed XPO1 disrupts the normal distribution of cargos between nucleus and cytoplasm on page 7. To avoid confusion, we also revised the result part on page 7, emphasizing on the difference rather than the similar increase in puncta size by opposing manipulations.

      (5) The removal of RNA-seq data weakens rather than strengthens the organoid section.

      The authors have removed the bulk RNA-seq analysis from the revised manuscript in response to concerns that the modest transcriptional rescue was being over-interpreted. While the decision to remove over-interpretation is appropriate, the result is that the organoid section now rests entirely on pTDP-43 immunostaining as its sole readout. The revised paper thus uses reduction in immunofluorescent pTDP-43 puncta in homozygous K181E organoids as the only evidence that nuclear export inhibition mitigates TDP-43 proteinopathy in a disease-relevant context. This is a weaker evidentiary base than before the revision, not an improvement. The originally requested more sensitive orthogonal readouts, including biochemical fractionation for SDS-insoluble TDP-43, filter-trap assays, or RNA aptamer-based detection of TDP-43 aggregates, remain absent. Without at least one additional independent measure confirming that cytoplasmic TDP-43 aggregation is genuinely reduced rather than simply rendered antigenically undetectable, the organoid conclusion is not adequately supported. At minimum, the authors should provide total and cytoplasmic TDP-43 fractionation data from organoid lysates to corroborate the immunostaining result.

      We removed the RNA-seq analysis because both the reviewers and editors agreed that the modest transcriptional rescue should not be overinterpreted. Upon reconsideration, we believe that reinstating these data would not address the reviewer's principal concern, namely whether nuclear export inhibition reduces TDP-43 aggregation in organoids. We have therefore chosen to limit our conclusions to the direct observation supported by the current data, namely a reduction in cytoplasmic phosphorylated TDP-43 immunoreactivity. We also add a sentence to acknowledge that “Whether the reduction in pTDP-43 immunoreactivity reflects a decrease in insoluble TDP-43 aggregates remains to be determined” on page 10.

      (6) No functional neuronal readout has been provided for the organoid model.

      The organoid section now makes the claim that "nuclear export is required for the formation of p-TDP-43-containing aggregates in a disease-relevant organoid model," but no measure of neuronal health, integrity, or function is reported in association with this. Even a simple assessment of neuron survival by TUJ1 or MAP2 quantification, neurite complexity, or cleaved caspase-3 staining before and after KPT-276 treatment would substantially strengthen the biological significance of the pTDP-43 reduction. The current data establish a pharmacological effect on a pathological marker but do not demonstrate that this has any consequence for neuronal biology in the organoid, which is what the disease-relevance framing implies.

      We thank the reviewer for this helpful suggestion. We agree that assessments of neuronal survival or function would be important if the manuscript were claiming that nuclear export inhibition improves neuronal health or rescues disease phenotypes in the organoid model. However, in response to the reviewers' comments regarding the physiological relevance of the homozygous K181E organoids, we have substantially revised both the framing and interpretation of this section.

      Specifically, we have revised the statement to read, "These results imply that maintaining TDP-43 in the nuclear demixed liquid state might diminish p-TDP-43 accumulation but whether the reduction of pTDP-43 immunoreactivity reflects a decrease in insoluble TDP-43 aggregates remains to be determined," thereby limiting our conclusion to the direct experimental observation. We no longer make claims regarding disease modification or functional rescue in the organoid model. Given this revised scope, we believe that additional measurements of neuronal survival or function, while certainly of interest, are not essential to support the conclusions presented in this study. We have also revised the conclusion to explicitly acknowledge that the functional consequences of reducing p-TDP-43-positive puncta (whether this can be translated to reduced aggregation) remain to be determined by future studies (page 10).

      (7) The abstract and title continue to overstate the mechanistic conclusions.

      Despite the stated intent to reframe the study as a screening study and to temper the conclusions, the revised abstract retains the language: "These findings establish nuclear export as a key regulator of TDP-43 phase transitions and define a mechanistic framework that links altered nuclear transport and phase dynamics to TDP-43 aggregation potential." Similarly, the Discussion still states: "a particularly compelling aspect of our study is the discovery that the nuclear export receptor XPO1 governs TDP-43 liquid-to-solid transitions and subcellular localization." The word "governs" and the phrase "establish nuclear export as a key regulator" are not warranted by data that derive entirely from an overexpressed acetylation-mimetic mutant in a colon cancer cell line and a homozygous K181E organoid model. A more accurate framing would describe these findings as identifying nuclear export as one of several cellular processes that modulate TDP-43 phase behavior in a sensitized model system, with an indirect RNA-mediated mechanism that remains to be defined at the molecular level. The title change from "governs" to "modulates" is appreciated but does not extend into the abstract and Discussion, where the strong causal language persists.

      We have revised the title of the paper, reframing it as a screen that reveals modulators of TDP-43 phase separation. The last sentence of the abstract is also revised accordingly. It now reads as “These findings identify multiple modulators of TDP-43 phase transitions in a sensitized model system and establish a framework for further dissecting the link between nuclear transport and TDP-43 phase dynamics.” We also tone down our conclusions and discussions.

      (8) Individual siRNA knockdown validation for XPO1 has not been provided.

      The authors argue that validation with 6 independent siRNAs across two rounds of screening, combined with convergent pharmacological data, is sufficient to establish XPO1 as a genuine hit. While the convergence of chemical and genetic evidence is reassuring, the specific request was for protein-level confirmation of XPO1 knockdown efficiency in the DLD1 TDP-43 2KQ cells used for mechanistic follow-up, together with demonstration that the anisosome phenotype is specifically caused by loss of XPO1 and not by off-target effects. This is a straightforward experiment, and its absence is particularly notable given that the entire mechanistic XPO1 narrative hinges on this specificity. At minimum, an immunoblot confirming XPO1 protein depletion in cells treated with the siRNA pool identified in the screen, in the same cell background and induction conditions as the follow-up experiments, should be provided.

      While we agree with the reviewer that studies relying on siRNA should provide sufficient information regarding knockdown efficiency and specificity, we respectfully disagree that this should be a major concern in the present study. As explained in the manuscript, we deliberately chose not to pursue mechanistic studies using chronic XPO1 knockdown because prolonged depletion of this essential nuclear export factor is likely to produce secondary effects that could complicate data interpretation. Instead, we employed multiple chemically distinct XPO1 inhibitors to achieve acute inhibition, thereby minimizing indirect consequences while providing a more appropriate approach for mechanistic analysis.

      We agree that assessing knockdown efficiency is technically straightforward. However, because our mechanistic conclusions are based primarily on acute pharmacological inhibition rather than siRNA-mediated depletion, we prioritized experiments that directly addressed the central mechanistic questions raised by the reviewers, particularly the semi-permeabilized cell assay. Moreover, the XPO1 inhibitors used in this study are well-characterized, highly specific compounds that have been extensively validated and widely used in the literature. We therefore believe that our experimental strategy provides a reliable basis for the conclusions presented.

      (9) The identity of XPO1-dependent cargos that regulate anisosome dynamics remains entirely unknown.

      The authors acknowledge that XPO1 does not directly bind TDP-43 and that the mechanism is likely indirect. The new RNA data provides one plausible indirect pathway. However, the possibility that one or more specific RNA-binding proteins or splicing factors, whose nuclear levels rise upon XPO1 inhibition, are the proximate drivers of anisosome stabilization has not been addressed. This matters because if the relevant mechanism operates through a specific cargo rather than bulk RNA retention, the model for how nuclear export connects to TDP-43 aggregation in disease would be fundamentally different. The authors decline to pursue adaptor identification on grounds of scope, which is a defensible position for future work. However, the framing should explicitly state that the current data cannot distinguish between bulk RNA retention and cargo-specific effects, and that the conclusion that nuclear export modulates TDP-43 phase behavior via RNA accumulation is a working hypothesis supported by but not proven by the RNase T1 experiment.

      We thank the reviewer for this helpful suggestion. We have now added a sentence on page 11, which state that “our data does not exclude other RNA species or RNA-binding proteins as anisosome stabilizer. Whether RNA-dependent stabilization of anisosomes operates in the same way in intact cells also requires validation.”

      Minor remaining issues.

      The number of independent iPSC clones and organoid batches used for the KPT-276 treatment experiment is now stated as two batches per condition, which is minimal for a 3D organoid study and does not fully address the concern about clone-level variability. Ideally, organoids from at least two independently derived isogenic clones per genotype would be used. The mCherry overexpression control added in Supplemental Figure 4 is a useful addition and is acknowledged. The immunoblotting confirmation that drug treatments do not alter total TDP-43 levels addresses a prior concern adequately. The addition of the sentence noting that anisosomes have not been validated in human patient samples is appreciated and appropriate. Statistical detail has been improved in figure legends. These minor improvements are noted positively but do not compensate for the major unresolved concerns above.

      We thank the reviewer for his/her appreciation of our previous revision. We hope that the new changes now satisfactorily address the remaining concerns.

      Reviewer #2 (Public review):

      This manuscript addresses an important and timely question in TDP-43 biology by systematically identifying regulators of TDP-43 anisosome formation, with a particular focus on nuclear export via XPO1. Using a combination of unbiased chemical screening, genetic perturbation, and advanced imaging approaches, the authors propose that inhibition of nuclear export modulates the abundance and biophysical properties of TDP-43 anisosomes. They further strengthen their findings by introducing an additional model system, a semi-permeabilized in vitro assay, which provides mechanistic evidence that XPO1 activity prevents anisosome dissolution by retaining nuclear RNAs. The study is conceptually innovative and has potential relevance for neurodegenerative diseases characterized by TDP-43 pathology. Some minor concerns remain, mostly about experimental design of the newly added data.

      Strengths:

      (1) The study employs an unbiased, hypothesis-free compound screen to identify regulators of TDP-43 anisosome formation, which is a major strength and reduces confirmation bias.

      (2) The authors combine chemical and genetic screening approaches, providing orthogonal validation of key pathways and increasing confidence in the biological relevance of top hits.

      (3) The focus on biophysical properties of TDP-43 assemblies, assessed through imaging and FRAP, moves beyond simple presence/absence of aggregates and provides mechanistic insight into the biophysical states of TDP-43.

      (4) The use of multiple experimental modalities, including live-cell imaging, FRAP, pharmacological perturbation, and transcriptomic analysis, reflects a technically sophisticated and ambitious study design.

      (5) The authors attempt to extend findings beyond immortalized cancer cell lines by incorporating organoid models, demonstrating awareness of disease relevance and translational importance.

      (6) The authors extend their study by incorporating a semi-permeabilized in vitro system, which provides compelling evidence that inhibition of nuclear export promotes the retention of nuclear anisosomes, an effect driven by the accumulation of nuclear RNAs.

      Overall, the manuscript is clearly written and logically structured, making complex experimental workflows accessible and the central hypotheses easy to follow.

      We thank the reviewer for acknowledging the strength and the potential significance of our study.

      Weaknesses:

      (1) The manuscript has significantly improved with the revisions. Some experimental procedures and method details, as well has statements remain incompletely described:

      (a) What is the smear in Figure S1 after VLX treatment?

      We thank the reviewers for the positive assessment. We do not know why VLX treatment causes a fraction of TDP-43 to migrate slowly. We suspect that it may form detergent-insoluble aggregates. However, we cannot be sure whether this occurred during drug treatment or sample preparation. We now add a sentence in the figure legend to clarify this point.

      (b) The authors state that "The reduction in TDP-43 signal was not due to protein elimination.", however no data is provided to support that statement.

      We reasoned that the reduction in TDP-43 was probably not caused by protein elimination because the puncta could be reformed when permeabilized cells were incubated with exogenously added cytosol and ATP/GTP. We have revised the text to avoid this confusion. The revision on page 8 reads as “The reduction in TDP-43 signal probably resulted from a shift of TDP-43 from a phase-separated high fluorescent state into a soluble state with reduced fluorescence intensity (Zhang et al., 2026). We attributed this phenotype to the depletion of cytosolic factors and ATP during cell permeabilization because it is known that anisosome formation and maintenance require HSP70, a cytosolic ATPase (Yu et al., 2021).”.

      (c) The authors state that "TDP-43 shifts from phase-separated state to a soluble state ...", however no data is provided to support that statement.

      Since TDP-43 protein was apparently still in the nucleus after cell permeabilization (see above) but became invisible, the best interpretation is that the protein is shifted into a soluble state, which reduces the fluorescence intensity substantially. We have revised the text to clarify this point. We also cited a recent study showing that EGFP-alpha-synuclein oligomerization/aggregation enhances its fluorescence intensity in cells.

      (d) Why did the authors choose cow lover cytosol for this study?

      The main reason is because we have access to a large amount of cow liver cytosol that is known to have activities in in vitro reconstitution assays. We now cite several papers from us that reported the use of the same cytosol in other in vitro assays in the method (page 14). 

      (e) The experimental setup for supplementing with cytosol/ATP/GTP is unclear. A more detailed schematic would be helpful to understand at what stage in the experiment these factors were added. Which step of the protocol was performed at 37 {degree sign}C, which is indicated in the figure schematic but not described in the methods.

      We now revise the schematic in Figure 6A and include more details in the method and figure legend.

      (f) In the organoid model, the authors mention that they observe similar levels of total TDP-43, however they do not provide quantification. Instead, they provide a graph that shows highly significant changes in nuclear TDP-43, which was not addressed in the text.

      The total TDP-43 level was shown by immunostaining in green in Figure 7. This was used as a control to show that the increase in p-TDP-43 was not simply caused by an overall increase in its protein level. We have added the quantification to Figure 7C. We also discuss the reduced nuclear TDP-43 in organoids bearing the disease mutation in the main text.  

      Additionally, some questions remain unclear:

      (1) The anisosomes induced by ATP/GTP or cytosol are insufficiently characterized. It remains unclear whether these structures correspond to canonical ring-shaped anisosomes, and whether they exhibit dynamic (liquid-like) or more static (gel-like) properties.

      We agree that the structures reformed after incubating permeabilized cells with cytosol and ATP/GTP are not fully characterized. Due to their small size, we could not see the typical ring-shaped anisosome morphology. FRAP experiment is also tricky. Due to these issues, we have revised the text to acknowledge that we do not know the exact identity of these structures. We speculate that they are anisosome-related because like anisosome formation, it depends on cytosolic factor and energy (page 9). It is worth noting that whether these structures are anisosomes is not the main conclusion of this experiment. We conclude from this experiment that TDP-43 was still in the nucleus after cell permeabilization (not degraded). The fact that we could not see the protein likely because the protein was in a low-fluorescence soluble state.

      (2) The contribution of the cytosol and ATP/GTP supplementation experiments to the overall narrative is unclear. While the findings are intriguing, their interpretation within the context of the study is not well articulated. In particular, the rationale for including cytosol is not sufficiently justified, given that ATP/GTP alone induces a pronounced effect, whereas cytosol alone does not.

      Since the formation of anisosome requires HSP70, a cytosolic chaperone that likely needs to be imported into the nucleus, we included cytosol and ARS/GTP in our in vitro reaction. We revise the description in the result part to improve clarity (page 8-9).

      (3) The authors should address why endogenous XPO1 does not co-localize with anisosomes, whereas overexpressed XPO1 does. This raises the possibility that the observed co-localization may be an artifact of non-physiological protein levels, which should be discussed.

      As discussed in Yu H et al., Science 2021, proteins in anisosomes cannot be stained by antibodies due to an antibody accessibility issue. It was mentioned in our paper as “since antibody staining could not conclusively demonstrate the sequestration of endogenous XPO-1 in anisosomes due to an antibody penetration barrier {Yu, 2021 #937}.” We now revise this section completely to better clarify this point. We could see overexpressed XPO1 in anisosome because it has a mCherry tag.

      (4) The iPSC-based model remains insufficiently characterized. While the authors propose that this system recapitulates the accumulation of liquid and solid aggregates resembling anisosomes, it is unclear whether this phenotype is robustly observed and whether KPT treatment effectively modulates it.

      The full characterization of the iPSC-derived organoids is presented in a second paper that is posted in BioRxiv (https://www.biorxiv.org/content/10.1101/2025.11.09.687455v2), which is cited. This manuscript reports not only the accumulation of p-TDP43, but also other ALS-related phenotypes including cell death, gene transcriptional changes, cryptic exon inclusion etc. in mutant organoids.

      (5) The rationale for the selected treatment durations is unclear, and the timing appears inconsistent across experiments (ranging from 3 to 16 hours), including within experiments involving the same compound. This variability should be justified or standardized.

      The longer treatment (24 h) was used in the chemical genetic screen in which different drugs may act with different efficiency. To maximize our chance of detecting more drug effect, we used a longer treatment scheme. For later follow-up experiments involving Spuatin-1, Bortezomib, TRP, because the phenotype appears quickly. To avoid secondary effects from long treatment, we shortened the treatment to 3-5 hours. For LMB treatment, we used long treatment to reveal the steady-state phenotype (anisosome enlargement in size and reduction in number has reached maximum). This time point was determined in Figure 5A-C. In contrast, shorter treatment (5 h) was to reveal early changes that might be causal to the end-point phenotypes (e.g. the anisosome fusion phenotype could be detected as early as 5 h post-treatment). We have added some explanations in the result section to make this point clear.

      (6) Several figure legends require clarification:

      We thank the reviewer for pointing out the inconsistencies. We have corrected the outstanding issues, as explained below.

      (a) In the section stating “Collectively, our results suggest that the stability and dynamics of anisosomes are modulated by XPO1-mediated nuclear export ...", the cited figure appears to be incorrect. This should refer to Figure 5L rather than Figure 5J.

      Thanks for pointing out this error. This is now corrected.

      (b) Figure 1B: Please specify the number of replicates per concentration, the number of cells analyzed, and the model used for regression analysis. Additionally, the legend indicates a treatment duration of 15 hours, whereas Figure 1A states 24 hours.

      Due to the large sample size, each concentration was analyzed once. We have added other information to the figure legend. We also remove the redundant inaccurate information from the figure legend. The treatment time shown in the figure is correct as it was also indicated in the method.

      (c) Figure 2G: The authors state "7 anisosomes per condition," but the graph displays only 4-6 data points. Please clarify what each data point represents.

      We thank the reviewer for noticing the discrepancy and apologize for the error. We have corrected the figure legend to indicate that 4-6 anisosomes were analyzed for each condition. In Figure 2G, each data point represents the initial fluorescence loss rate averaged from the first 10 sec after reverse photobleaching.

      (d) Figures 3B and 3G: Please clarify whether a defined threshold was used to determine a "reduction in anisosome number."

      In Figure 3B, we used Z score >2 as the threshold. This is now mentioned in the legend and defined in the method. There is no Figure 3G.

      (e) Figure 4B: These do not represent biological replicates, as all samples derive from a single cell line; rather, they constitute independent experimental replicates.

      We have changed the figure legend throughout the paper accordingly.

      (f) Figures 5B and 5H: The legend states "n = 3 biological repeats," but the number of data points shown appears higher. Please clarify.

      In Figure 5B, the graph reflects data collected from 3 independent replicates. To ensure reliable baseline measurement, for each experiment, two independent control samples were included, which is why it has 6 data points. In Figure 5H, each dot represents a randomly selected imaging field. We now mention the total number of fields analyzed.

      (g) Figures 5K, 6C, and 6E: "Mean Fluorescence Intensity (MPI)" should be corrected to "MFI."

      These are all fixed. Thank you for pointing this out.

      (h) Figure 6C: Please include the number of cells analyzed and provide relevant statistical measures (e.g., R<sup>2</sup>, p-value).

      We now include the cell number in the legend and R<sup>2</sup> and p-value in the figure.

      (i) Figure 6D: The experimental timeline is unclear. Please specify the duration of incubation and the timing of each step.

      We now revise the experimental scheme in Figure 6A to better explain the experiment and the sequence of different events. For Figure 6D, permeabilized cells were incubated with cytosol with or without ARS/GTP for 40 min. This information is added to the figure legend.

      (j) Figure 7B: Improved labeling is needed (e.g., clarification of "mean spot volume") to better align with the figure legend.

      To improve clarity, we change mean spot volume to p-TDP-43 puncta mean volume. This refers to the average volume of segmented phosphorylated TDP-43-positive puncta.  

      Reviewer #3 (Public review):

      Summary:

      TDP-43 proteinopathy is broadly found in neurodegenerative diseases. This manuscript investigates how nuclear export influences the biophysical properties of TDP-43. The authors use a combination of chemical screening and genome-wide siRNA screening to identify pathways that modulate TDP-43 liquid-to-solid transitions. Overall, the study employs a broad array of approaches and addresses an important question in TDP-43 pathobiology. The identification of nuclear export as a central regulator is compelling and conceptually aligns with the emerging view that TDP-43 nucleocytoplasmic trafficking is a major defect in neurodegeneration.

      Strengths:

      This work integrates chemical and genetic screening to identify novel modifiers. The candidates were validated in both reporter cell lines and iPS-differentiated organoids. The findings support the nucleocytoplasmic transport is important for the biophysical properties of TDP-43.

      Comments on revised version.

      The manuscript has been improved with more data and clarification. The RNase T1 treatment experiment suggests that RNA is required for anisosome integrity. However, this does not directly demonstrate LMB increases nuclear RNA availability as changes in protein composition or other RNA-dependent mechanisms may also contribute. The conclusion and discussion need to be edited to consider these alternative scenarios. Overall, as most of the evidence remains indirect, the manuscript should avoid overinterpretation regarding the mechanisms underlying TDP-43 phase transition and aggregation.

      We thank the reviewer for this helpful suggestion. We have added a few sentences in the discussion (page 10) to acknowledge the limitation of the semi-permeabilized cell assay. Specifically, we mentioned that “However, our data does not exclude other RNA species or RNA-binding proteins as anisosome stabilizer. Whether RNA-dependent stabilization of anisosomes operates in the same way in intact cells also requires further validation.” We also revise our manuscript throughout to avoid over-interpretation.  

      Recommendations for the authors:

      Editor's notes:

      The value of the work is clear.

      We also recognise that it may not be possible/practical to get around the 'incomplete' appellation attached to this body of work, by further experiments. However, there may be scope here to retreat from the less well supported mechanistic claims-by editing the title, abstract and discussion and thus earn a 'solid' descriptor on a revised paper that remains a useful addition.

      The authors are best placed to consider creatively how to achieve this.

      We thank the editors for this helpful suggestion. We have revised the manuscript extensively to address every single concerns of the reviewers.

      Reviewer #1 (Recommendations for the authors):

      This revision addresses some minor concerns and adds one mechanistic experiment of genuine value. However, the major deficiencies of the original submission persist: the exclusive reliance on non-physiological TDP-43 model systems without validation in more disease-relevant contexts, the unresolved asymmetry between the two XPO1 perturbation phenotypes, the thin organoid section that now has fewer readouts than before the revision, and overstatement of the mechanistic conclusions in the abstract and Discussion. The manuscript in its current form still does not provide methods, data, and analyses that sufficiently support the primary claim that nuclear export is an established key regulator of TDP-43 phase transitions with mechanistic and disease relevance.

      As mentioned before, we have clarified the interpretation of the data, tempered conclusions where appropriate, and revised the text to explicitly acknowledge the limitations of the current study. We hope that these changes satisfactorily addressed the reviewer’s concern.

      Reviewer #2 (Recommendations for the authors):

      I would suggest adjusting the title to match the data, which shows so much more than just an effect of nuclear export.

      We thank the reviewer for this suggestion. We have changed the title to “Cellular modifiers of TDP-43 phase transition and cytoplasmic aggregation”

      When introducing the semi-permeabilized cell-based in vitro assay, it would be helpful to add a short statement describing what this model resembles, and what advantage it can bring to use this system in the context of the study.

      We have revised this section extensively and hope that improves the clarity. See marked text in page 8-9.

    1. eLife Assessment

      The study presents valuable findings of a new E. coli cell-free protein synthesis (eCFPS) system that has been simplified by reducing the number of core components from 35 to 7; furthermore, the findings communicate a simplified 'fast lysate' preparation that eliminates the need for traditional runoff and dialysis steps. The system's robustness is exhibited by its applicability to nanoluc, an ubiquitous protein and to more challenging proteins like vimentin and the active restriction endonuclease Bsal. The evidence is convincing and supports the main claims on efficiency backed up by investigations on the mechanisms. The paper will be of interest to scientists in cell and molecular biology, microbiology, biotechnology and protein synthesis.

    2. Reviewer #1 (Public review):

      [Editors' note: this version has been assessed by the Reviewing Editor without further input from the original reviewers. The authors have addressed the comments raised in the previous round of review.]

      Summary:

      The authors presented a simplified E. coli cell-free protein synthesis (eCFPS) system reduces core reaction components from 35 to 7, improving protein expression levels. They also presented a "fast lysate" protocol that simplifies extract preparation, enhancing accessibility and robustness for diverse applications.

      Strengths:

      The authors present a valuable new protocol for eCFPS, which simplifies its application.

    3. Reviewer #2 (Public review):

      Summary:

      The authors have made a convincing argument that the current system of in vitro translation using E. coli extracts can be significantly optimized to work with much lesser components, while maintaining activity. They have showcased their improved activity using not only physical but also functional readouts.

      Strengths:

      The experiments are designed in a very logical and easy to understand manner, which makes it easier not only to follow the paper, but also reproduce the results. Functional assays with the synthesized proteins are a good way to demonstrate functionality and applicability of the system. They also benchmark their system against a commercial kit to show superior performance of their system.

      Weaknesses:

      The production of the lysate requires special instrumentation, limiting accessibility.

      Comments on previous version:

      Thank you to the authors for addressing the concerns both textually and experimentally. This work has significant value.

    4. Reviewer #3 (Public review):

      Summary:

      The authors aimed to overcome the challenges associated with complex, conventional prokaryotic cell-free protein synthesis (CFPS) systems, which require up to thirty-five components, by developing a streamlined and efficient E. coli CFPS platform to encourage broader adoption. The main objective was to reduce the number of reaction components from thirty-five to seven, while also developing an accessible 'fast lysate' preparation protocol that eliminates time-consuming runoff and dialysis steps. The authors also sought to demonstrate the robustness and translational quality of this streamlined system by efficiently synthesising challenging functional proteins, including the cytotoxic restriction endonuclease BsaI and the self-assembling intermediate filament protein vimentin.

      Strengths:

      This study presents several key strengths of the optimised E. coli cell-free protein synthesis system in terms of its design, performance and accessibility.<br /> - The reaction mixture has been dramatically simplified, with the number of essential core components successfully reduced from up to thirty-five in conventional systems to just seven.<br /> - The "fast lysate" protocol is a significant advance in terms of procedure.<br /> - The system's ability to synthesise challenging, functional proteins is evidence of its robustness.

      Comments on previous version.

      The authors have adequately addressed my previous concerns.

    5. Author response:

      The following is the authors’ response to the previous reviews

      The revisions in this version are minor and primarily include the addition of RT-qPCR validation experiments. In addition, the benchmarking data against commercial systems have now been incorporated into the main manuscript.

      We also sincerely appreciate Reviewer 2 and Reviewer 3 for their highly encouraging evaluations and recognition of our system's robustness. To fully address the remaining mechanistic queries from Reviewer 1 and the benchmarking concerns from the editors, we have performed quantitative RT-qPCR to directly measure transcript levels and have integrated our commercial benchmarking data into the revised manuscript.

      (1) The authors have satisfactorily addressed the concerns raised by the reviewers. However, the mechanistic basis of the observed performance gain remains insufficiently substantiated. The attribution of this improvement to enhanced transcription is currently speculative. This point could be directly tested by quantifying mRNA levels, for example, using real-time PCR, in both the initial and optimized systems. Such analysis would significantly strengthen the mechanistic interpretation of the results.

      To directly validate our claims regarding transcriptional efficiency, we performed quantitative RT-qPCR to determine the transcription levels of the reporter gene in both systems.

      First, we established no-reverse-transcriptase (no-RT) controls to verify complete DNA template removal. The Ct values for these controls remained above 34, confirming the absence of plasmid DNA contamination in our RNA samples.

      Second, transcript levels were calculated using the comparative 2<sup>-ΔΔCt</sup> method, normalized to the standard initial system (100 ng/μL T7) at 30 min. The optimized system achieved a 16.56-fold increase (P < 0.001) in transcript levels. In contrast, supplementing the initial system with high concentrations of T7 RNA polymerase (400 ng/μL) only yielded a 2.87-fold increase (P < 0.01)—which is nearly 6-fold lower than our optimized system.

      These findings perfectly mirror our protein-level titration assays (Figure S3C). Supplementing the initial system with excess T7 RNA polymerase fails to rescue either transcript accumulation or protein expression. This mutual validation confirms that transcription is severely bottlenecked in traditional systems due to rapid nucleotide degradation or inhibitory reaction environments. By streamlining the reaction buffer to seven core components and omitting runoff/dialysis, our system successfully relieves these systemic bottlenecks. We have incorporated these new qPCR findings into Figure 3B, the Methods, and the Results sections of the revised manuscript.

      (2) Despite the study representing an advancement towards simplifying protein expression workflows, the evidence is solid and supports the main claims however minor weakness exists i.e. the efficiency claims about the new system needs to be supported by accurate comparisons with typical cell free expression systems...

      We appreciate the editor’s emphasis on establishing standard performance benchmarks. To address this important point, we would first like to highlight that our manuscript already contains extensive, rigorous benchmarking against typical cell-free platforms widely utilized in the literature. This includes detailed head-to-head comparisons with both our 35-component "initial" system and the classical, widely established "PEP-based" system across multiple expression kinetics and western blot analyses (as shown in Figure 4 and Figures S3–S4).

      To fully embrace the editor's valuable recommendations regarding standard commercial performance, we are very pleased to formally integrate our commercial benchmarking data into the revised manuscript as Figure S3C.

      To maintain technical neutrality, we have omitted specific brand names, presenting it generically as "a high-end commercial cell-free system." The data demonstrate that our optimized system significantly outperforms this commercial alternative in both expression speed and final absolute yield, reaching an absolute productivity of 0.46 mg/mL compared to approximately 0.21 mg/mL for the commercial kit.

      We are grateful for the guidance from the editors and reviewers, which has significantly strengthened the scientific rigor of our work.

    1. eLife Assessment

      This is an important study that comprehensively determines the consequences of DNMT3A mutations on human neuronal development in culture. The data derived from multiple different mutations in iPSC and hESC derived neurons and organoids are convincing and well controlled. This work will be of interest to researchers who study chromatin mechanisms of brain development and to those interested in DMNT3A mutations in Tatton-Brown-Rahman Syndrome.

    2. Reviewer #1 (Public review):

      Summary:

      This is an important study that describes the consequences of the DNMT3A mutation in human neuronal development for the first time. The selective impact of DNMT3A function on GABAergic interneurons is interesting and an important feature of future therapeutics. The claims made in that manuscript are supported by strong evidence for the most part. And the data are of high quality in general and presented well.

      Strengths:

      The strengths of the work include 1. Characterization of multiple DNMT3A loss-of-function alleles, including two misense variants, R882H, P904L, and a deletion allele. The missense mutation lines both include an ideal control with the same genetic background. The CRISPRi-mediated DNMT3A knockdown has also been included. The study identifies the mTOR-PI3K pathway as a factor of overgrowth issues found in the mutant organoid. In bulk mRNA sequencing and whole-genome bisulfite sequencing, identify hypomethylated genomic regions associated with gene expression repression. Again, this is more pronounced in the ventral organoid compared to the dorsal organoid. In addition, the extensive electrophysiological characterizations with a high-density microelectrode array support the more mature status of mutant interneurons.

      Weaknesses:

      Although a strong study overall, some weaknesses are noted. These include:

      (1) The lack of validation data for the generated iPSCs and hESCs, such as the chromosomal contents, ploidy, and pluripotency states

      (2) Other weaknesses relate to data interpretation and insufficient discussion of related matters, as detailed in the recommendations to the authors.

      (3) Also, some errors are noted and detailed in the recommendation section.

      Comments on the latest version:

      I have reviewed the revised manuscript and the authors' responses to the reviewers' comments. They addressed the comments adequately.

    3. Reviewer #2 (Public review):

      Summary:

      Chapman, Determan et al. investigate how pathogenic mutations in DNMT3A which cause of Tatton-Brown-Rahman Syndrome (TBRS) disrupt human cortical developmental processes using a comprehensive panel of human pluripotent stem cell models spanning DNMT3A loss-of-function severity. The authors aim to identify the cellular and molecular mechanisms underlying TBRS-associated brain overgrowth and intellectual disability, and to test whether mechanistic convergence exists between TBRS and other overgrowth-intellectual disability disorders (OGIDs) caused by mutations in EZH2 (Weaver syndrome) or PIK3CA pathway components. Their central conclusion is that GABAergic interneuron development is selectively vulnerable to DNMT3A mutation where reduced DNA methylation causes premature de-repression of neuronal and synaptic genes, driving precocious neuronal maturation and hyperactivity sufficient to disrupt neuronal network synchrony. This report adds to a growing literature supporting the vulnerability of GABAergic interneurons in NDDs and further provides a mechanistic view of this vulnerability potentially convergent across OGIDs. The mechanistic claims around H3K27me3 compensation and mTOR-based therapeutic convergence, while promising, rest on more preliminary evidence and would benefit from the distinction between correlation and mechanism being made more explicit in the text. Overall, this is a compelling study with rigorous experimental design and novel findings with potential impact across better understanding OGID pathophysiology.

      Strengths:

      (1) A major strength of this work is the breadth and rigor of the disease modeling approach. Four independent TBRS model systems are used in tandem: a patient-derived iPSC line with isogenic CRISPR-corrected control (R882H), a knock-in hESC model (P904L) with its wild-type isogenic, patient deletion iPSC lines (Del1/2), and CRISPRi knockdown models (G1/G2), collectively spanning a range of DNMT3A loss-of-function that correlates with phenotypic severity. This allelic series design substantially strengthens causal inference beyond what any single isogenic pair could provide.

      (2) The multi-omic integration across matched developmental stages provides a strong mechanistic foundation for the cellular phenotyping and provides significantly enhanced novelty. RNA-seq, whole-genome bisulfite sequencing, and H3K27me3 CUT&Tag are combined in the same cell types and timepoints show that DNMT3A loss reduces CG methylation at neuronal and synaptic gene loci, leading to premature transcriptional activation.

      (3) The selective vulnerability of ventral (GABAergic) versus dorsal (glutamatergic) progenitors is one of the study's most important findings. This lineage specificity is consistently observed across all model systems and in both 2D and organoid formats, where ventral NPCs show increased proliferation, premature neuronal gene expression, and increased neurogenesis, while dorsal NPCs are largely unaffected at the transcriptomic and cellular level despite exhibiting comparable DNA methylation changes. This adds to a body of emerging work showing GABAergic interneuron vulnerability in NDDs where ubiquitously expressed genes such as chromatin modifiers are perturbed and provides additional molecular insights into potential mechanisms of "resilience" of dorsal populations.

      (4) The functional characterization follows a logical progression from single-neuron electrophysiology (demonstrating GABAergic hyperactivity with increased action potential amplitude and firing rate) to network-level analysis using high-density multi-electrode arrays. The HD-MEA experimental design - pairing TBRS or control GABAergic neurons with a constant background of control iGlut neurons - cleanly isolates GABAergic dysfunction as the driver of network hypersynchrony.

      Weaknesses:

      (1) The concomitant induction of proliferation and differentiation in TBRS V-NPCs is conceptually striking, since these are generally considered antagonistic developmental programs. The authors clarify that neuronal and synaptic gene de-repression is the more prominent direct consequence of mCG loss, while PIK3/AKT/mTOR pathway upregulation is not itself directly linked to differentially methylated regions, suggesting an indirect relationship between DNMT3A LOF and increased proliferative signaling. This framing is reasonable, but the mechanism linking DNMT3A mutation to mTOR activation remains unresolved, and the manuscript would benefit from being explicit about this gap. Relatedly, the rapamycin rescue, while demonstrated across multiple models including 904 and Del1 (Supplementary Fig. S3e-f), remains limited to proliferation readouts. Whether mTOR inhibition also rescues the downstream neurogenesis, maturation, or network phenotypes is an important open question that the authors appropriately frame as motivation for future work.

      (2) The claim that H3K27me3 compensates for mCG loss is supported by prior work (Lii et al. 2022), which demonstrated increased PRC2 component expression and H3K27me3 gain at sites of DNA methylation loss in Dnmt3a knockout mouse neurons, and by data showing that PRC2 subunits (SUZ12, EED, EZH2) are significantly more highly expressed in D-NPCs than V-NPCs. Together, these findings provide a plausible molecular basis for why dorsal progenitors may be better equipped to maintain repression when DNA methylation is lost, and they make the EZH2 overexpression rescue in V-NPCs more interpretable. Yet, a formal distinction related to two competing, potentially underlying mechanisms, between active compensation, in which EZH2 is recruited to specific loci in response to methylation loss, and functional redundancy, in which higher baseline Polycomb occupancy in dorsal cells simply becomes the dominant repressive mark once mCG is reduced, has not been resolved.

    4. Reviewer #3 (Public review):

      Summary:

      In this manuscript, the authors investigated TBRS etiology by using new human pluripotent stem cell models, modeling varying levels of TBRS-associated loss of DNMT3A function. They identified increased lineage-specific proliferation of precursors in TBRS ventral MGE-like progenitors, which they propose was related to increased signaling through the PIK3/AKT/mTOR pathway. Furthermore, they show that reduced DNA methylation during MGE-like progenitor differentiation into GABAergic interneurons can cause a premature expression of neuronal and synaptic genes, triggering precocious neuronal maturation. In conclusion, they propose that TBRS-derived GABAergic neurons exhibit hyperactivity that can alters the development and structure of neuronal networks.

      Strengths:

      Overall, the data presented is convincing, from an early developmental point of view, given that the iPSC-derived 2D cultures or organoids used do not get to reach a mature state. Nonetheless, the data clearly show the effects that deleterious mutations in TBRS can cause during the period of neurogenesis, which was missing in the field.

      Comments on revised version.

      The authors have responded to the reviewer's comments satisfactorily, and the manuscript has been much improved.

    5. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This is an important study that describes the consequences of the DNMT3A mutation in human neuronal development for the first time. The selective impact of DNMT3A function on GABAergic interneurons is interesting and an important feature of future therapeutics. The claims made in that manuscript are supported by strong evidence for the most part. And the data are of high quality in general and presented well.

      Strengths:

      The strengths of the work include: Characterization of multiple DNMT3A loss-of-function alleles, including two misense variants, R882H, P904L, and a deletion allele. The missense mutation lines both include an ideal control with the same genetic background. The CRISPRi-mediated DNMT3A knockdown has also been included. The study identifies the mTOR-PI3K pathway as a factor of overgrowth issues found in the mutant organoid. In bulk mRNA sequencing and whole-genome bisulfite sequencing, identify hypomethylated genomic regions associated with gene expression repression. Again, this is more pronounced in the ventral organoid compared to the dorsal organoid. In addition, the extensive electrophysiological characterizations with a high-density microelectrode array support the more mature status of mutant interneurons.

      Weaknesses:

      Although a strong study overall, some weaknesses are noted. These include:

      (1) The lack of validation data for the generated iPSCs and hESCs, such as the chromosomal contents, ploidy, and pluripotency states.

      We thank the reviewer for their constructive feedback. We previously validated our 882 models with whole genome sequencing and teratoma formation upon mouse fat pad injection, while the parental human embryonic stem cell line (WA01 hESCs) used for P904L variant knock-in was validated by our Genome Engineering Stem Cell (GESC) core upon derivation of that variant knock-in model. We have now added both karyotyping and pluripotency staining (SOX2/OCT4) for all other hPSC lines as (new) Supplementary Figure S17 and included further description in our Methods section under “hPSC Model Generation and Culture” (pg. 19, lines 6-7).

      (2) Other weaknesses relate to data interpretation and insufficient discussion of related matters, as detailed in the recommendations to the authors.

      We thank the reviewer for their insightful suggestions and have detailed our responses in the “recommendations to the authors” section.

      (3) Also, some errors are noted and detailed in the recommendation section.

      We thank the reviewer for catching these errors and have since corrected them, with detailed responses below.

      Reviewer #2 (Public review):

      Summary:

      Chapman, Determan et al. investigate how pathogenic mutations in DNMT3A, which cause Tatton-Brown-Rahman Syndrome (TBRS), disrupt human cortical developmental processes using a comprehensive panel of human pluripotent stem cell models spanning DNMT3A loss-of-function severity. The authors aim to identify the cellular and molecular mechanisms underlying TBRS-associated brain overgrowth and intellectual disability, and to test whether mechanistic convergence exists between TBRS and other overgrowth-intellectual disability disorders (OGIDs) caused by mutations in EZH2 (Weaver syndrome) or PIK3CA pathway components. Their central conclusion is that GABAergic interneuron development is selectively vulnerable to DNMT3A mutation, where reduced DNA methylation causes premature de-repression of neuronal and synaptic genes, driving precocious neuronal maturation and hyperactivity sufficient to disrupt neuronal network synchrony. This report adds to a growing literature supporting the vulnerability of GABAergic interneurons in NDDs and further provides a mechanistic view of this vulnerability, potentially convergent across OGIDs. The mechanistic claims around H3K27me3 compensation and mTOR-based therapeutic convergence, while promising, rest on more preliminary evidence and would benefit from the distinction between correlation and mechanism being made more explicit in the text. Overall, this is a compelling study with a rigorous experimental design and novel findings with a potential impact on a better understanding of the OGID pathophysiology.

      Strengths:

      (1) A major strength of this work is the breadth and rigor of the disease modeling approach. Four independent TBRS model systems are used in tandem: a patient-derived iPSC line with isogenic CRISPR-corrected control (R882H), a knock-in hESC model (P904L) with its wild-type isogenic, patient deletion iPSC lines (Del1/2), and CRISPRi knockdown models (G1/G2), collectively spanning a range of DNMT3A loss-of-function that correlates with phenotypic severity. This allelic series design substantially strengthens causal inference beyond what any single isogenic pair could provide.

      (2) The multi-omic integration across matched developmental stages provides a strong mechanistic foundation for the cellular phenotyping and provides significantly enhanced novelty. RNA-seq, whole-genome bisulfite sequencing, and H3K27me3 CUT&Tag are combined in the same cell types, and timepoints show that DNMT3A loss reduces CG methylation at neuronal and synaptic gene loci, leading to premature transcriptional activation.

      (3) The selective vulnerability of ventral (GABAergic) versus dorsal (glutamatergic) progenitors is one of the study's most important findings. This lineage specificity is consistently observed across all model systems and in both 2D and organoid formats, where ventral NPCs show increased proliferation, premature neuronal gene expression, and increased neurogenesis, while dorsal NPCs are largely unaffected at the transcriptomic and cellular level despite exhibiting comparable DNA methylation changes. This adds to a body of emerging work showing GABAergic interneuron vulnerability in NDDs where ubiquitously expressed genes such as chromatin modifiers are perturbed, and provides additional molecular insights into potential mechanisms of "resilience" of dorsal populations.

      (4) The functional characterization follows a logical progression from single-neuron electrophysiology (demonstrating GABAergic hyperactivity with increased action potential amplitude and firing rate) to network-level analysis using high-density multi-electrode arrays. The HD-MEA experimental design - pairing TBRS or control GABAergic neurons with a constant background of control iGlut neurons - cleanly isolates GABAergic dysfunction as the driver of network hypersynchrony.

      Weaknesses:

      (1) The concomitant induction of proliferation and differentiation in TBRS V-NPCs is conceptually striking, since these are generally considered antagonistic developmental programs. The authors partially address this tension by noting that DNMT3A LOF alone is insufficient to initiate neuronal differentiation, i.e., V-NPCs upregulate neuronal and synaptic genes while retaining progenitor identity, implying that transcriptomic priming and commitment to differentiation are decoupled. However, the relationship between the proliferative phenotype and the epigenetic priming phenotype remains mechanistically unresolved. The manuscript documents mTOR pathway upregulation at the protein level and identifies shared DEGs that include proliferative regulators, but it does not establish whether mTOR-driven proliferation and mCG-loss-driven neuronal gene de-repression/enhanced differentiation are causally linked or represent two independent consequences of DNMT3A LOF.

      We thank the reviewer for their comment and agree that this phenotype, whereby progenitors exhibited both increased proliferation and hallmarks of gene expression associated with neuronal differentiation is striking and interesting, given that these are typically antagonistic paradigms during normal development.

      We documented that these phenotypes involve upregulated expression of both neuronal/synaptic and proliferative genes in V-NPCs (Figure 2d), with concomitant loss of repressive DNA methylation at regulatory elements associated with these genes (Figure 2f, Supplementary Data 5). In this work, DNMT3A mutation had a more prominent role in de-repressing neuronal and synaptic gene expression to promote hallmarks of neuron differentiation, while playing a relatively less central role in direct regulation of proliferation genes, as seen from the relative prominence of neuronal/synaptic- versus proliferation-related GO terms in our Supplementary Data 5 table (pg. 6, lines 16-19).

      To examine the mechanisms underlying increased V-NPC proliferation in our TBRS models, we assessed a potential relationship with the PIK3/AKT/mTOR pathway, as this is implicated in increased proliferation resulting from DNMT3A-associated mutation in myeloid leukemia (Dai et al., 2017, PMID: 28461508). In our work, DNMT3A mutation increased the expression and/or phosphorylation of mTOR signaling pathway targets specifically in V-NPCs (Figure 1q-r, Supplementary Figure S3a-d). However, while TBRS mutation directly affected repressive DNA methylation at a suite of cell proliferation-related genes, these did not include the PIK3/AKT/mTOR pathway genes themselves, suggesting an indirect relationship between altered DNA methylation and increased mTOR signaling.

      We have since incorporated discussion of how DNMT3A-mediated gene repression and levels of PIK3/AKT/mTOR pathway signaling may be interacting, providing a framework for future studies to identify how these related OGID gene mutations may converge mechanistically (pg. 5, lines 19-21; pg. 16, lines 8-10).

      (2) Relatedly, the rapamycin rescue experiment is a valuable proof-of-concept for the PIK3/AKT/mTOR convergence but is limited to a single dose in a single model (882) with a single readout (Ki67+ proliferation). Given the prominence of mTOR pathway convergence in the manuscript as a potential shared therapeutic avenue across OGIDs, the data supporting this claim are somewhat preliminary. It remains unknown whether mTOR inhibition rescues downstream phenotypes (neurogenesis, gene expression, neuronal maturation) or whether less severe TBRS models respond similarly. This might also help tackle the first comment above. e.g., if mTOR inhibition rescued proliferation but not the transcriptomic priming, that would support two independent mechanisms.

      We thank the reviewer for their comment. We explored both the overall levels and phosphorylation of proteins involved in PIK3/AKT/mTOR signaling in the 882, 904, Del1, Del2, and KO V-NPC models (Figure 1q-r, Supplementary Figure S3a-d), finding specific increases of all proteins. We showed that rapamycin addition reversed the increased proportion of KI67+ proliferating cell nuclei resulting from 882 mutation in V-NPCs in main Figure 1s, while demonstrating that rapamycin also reduced the proportion of KI67+ nuclei observed in both less severe 904 and Del1 V-NPC models (Supplementary Figure S3e-f).

      We agree that understanding whether rapamycin treatment can rescue TBRS neuronal phenotypes would be very interesting, as previous work on Tuberous Sclerosis Complex has utilized rapamycin and other mTOR inhibitors to effectively reverse TSC-related alterations of neuronal morphology and neuronal hyperexcitability (Buttermore et al., 2025, PMID: 40792287). Future studies examining convergent mechanisms and therapeutics for OGIDs should examine how similarly targeting this and related pathways rescues altered neuronal morphology, maturation, and function, as we have demonstrated that TBRS mutation has subsequent consequences for V-IN differentiation, maturation, and function. This point has been detailed in the discussion section on pages 15-16.

      (3) The claim that H3K27me3 compensates for mCG loss is an important mechanistic point, but the current data do not distinguish between active compensation, in which EZH2 is recruited in response to methylation loss, and functional redundancy, in which H3K27me3 is independently established and becomes the dominant repressive mark once DNA methylation is reduced. The EZH2 knockdown/inhibition experiments show that H3K27me3 is sufficient to maintain repression at hypo-DMR sites, but they do not establish that H3K27me3 gain is itself a response to methylation loss. Because H3K27me3 profiling was performed only in the severe 882 model, it is also unclear whether H3K27me3 gain scales with DNMT3A LOF severity, as a compensatory model would predict. Finally, the EZH2 overexpression rescue is performed in V-NPCs, whereas the compensation model is developed primarily in D-NPCs, making it difficult to assess whether the same mechanism operates in the lineage where it was originally inferred.

      We thank the reviewer for the opportunity to clarify our findings and experimental reasoning. A previous study using a conditional Dnmt3a knockout mouse model (Li et al., 2022, PMID: 35604009) demonstrated increased expression of multiple PRC2 components following the loss of Dnmt3a. This study demonstrated that sites which lost DNA methylation gained H3K27me3 in postnatal neurons upon Dnmt3a loss. Therefore, we hypothesize that the gain of H3K27me3 likely occurs in response to loss of DNMT3A methylation.

      While we did not perform CUT&Tag for H3K27me3 in our less severe models, we did validate gene expression changes following EZH2 knockdown and inhibition in both the R882H (Figure 4g-h) and P904L (Supplementary Figure S8b) models, finding that gene expression was unchanged in the model with the less severe DNMT3A mutation (P904L). Based upon these findings, we hypothesized that compensatory H3K27me3 may occur only upon severe DNMT3A loss, as seen in the dominant-negative R882H model. Furthermore, as H3K27me3 compensation was more prominent in D-NPCs, we hypothesized that this might be sufficient to prevent de-repression and aberrant neuronal gene repression upon loss of DNMT3A-mediated repression in D-NPCs. However, since TBRS mutation caused the most prominent de-repression of neuronal gene expression in V-NPCs, we also tested whether EZH2 overexpression could reverse this, finding that it partially suppressed this dysregulated neuronal gene expression. To better clarify this logic and the findings, we have made text edits to this results section and referenced Li et al., 2022 in both the results (pg. 9, lines 16-21; pg. 10, lines 4-7,10-12) and discussion (pg. 16, lines 12-15).

      (4) The narrative framing of dorsal neuron development as unaffected by DNMT3A LOF is somewhat at odds with the data presented. The 882 D-NPCs show substantial DNA methylation changes, and TBRS D-INs exhibit what the authors describe as "substantive transcriptomic differences" involving persistent expression of pluripotency and progenitor genes, which seems to be a distinct but potentially significant phenotype. The impact of DNMT3A loss between ventral and dorsal lineages might be more accurately framed as divergent in nature rather than specific to a certain population.

      We thank the reviewer for their comment. While TBRS mutations appear to have a significantly stronger effect on V-NPCs and subsequently V-INs, both transcriptomic and methylation alterations do also occur upon TBRS mutation in D-NPCs and D-INs, as noted in Supplemental Figure S4d, S11, and Supplemental Data 2. However, we observed substantially greater molecular alterations in V-NPCs/V-INs, a lack of overt cellular phenotypes in D-NPCs where assayed, and a lack of functional consequences in matured D-INs, suggesting a more significant requirement for DNMT3A in regulating the differentiation and subsequent maturation of cortical inhibitory interneurons during embryonic and early pre-natal development, the developmental periods that we can readily model in hPSC-derived neurons.

      It should also be noted that these hPSC differentiation models do not recapitulate post-natal deposition of non-CpG (mCA) DNA methylation, a mechanism disrupted postnatally by TBRS-associated mutations in our prior work in murine models (Harrison Gabel; e.g. Beard et al., 2023, PMID: 37952155), which we have now added in the results section (pg. 7, lines 8-11). Therefore, we hypothesize that if we could sufficiently mature D-INs to a state that modeled postnatal development and recapitulated this non-CpG methylation, we might be able to detect cellular and functional phenotypes in later stage D-INs. To avoid misinterpretation, we have altered the language in the results section to confirm that there are both transcriptomic and methylation changes in our D-NPCs/D-INs, but that these are not accompanied by cellular phenotypes or neuronal dysfunction (pg. 6, lines 1-2; pg. 7 line 3; pg. 7, lines 20-23; pg. 8, lines 1-2; pg. 8, lines 18-23; pg. 9, lines 1-4).

      (5) SST stainings are not entirely convincing. They appear mostly nuclear, and some instances localized to rosettes in organoids, whereas the protein is largely confined to processes and is expected to be found outside progenitor-rich zones like rosettes.

      We agree that the perinuclear SST staining detected in these young ventral telencephalic-patterned organoids at day 30 differs somewhat from the more process-localized and cytosolic signal seen in later stage organoids in other studies. This may be related to the use of different commercial SST antibodies across studies but also likely reflects SST immunoreactivity in newborn neurons near the onset of SST expression. For example, immature SST-immunoreactive neurons in the early postnatal rat cortex exhibit predominant SST staining in perinuclear cytoplasm and short processes (e.g. Fig. 3 in Lee et al, PMID: 9664223) while acquiring more cytosolic and process-localized staining as postnatal neuron maturation occurs. Evaluation of immunopositivity for other markers of neurogenesis (ASCL1) and immature neurons (TUJ1) is also congruent with these findings for SST, with TBRS-associated mutations increasing in the fraction of cells in V-NPCs/V-ORGs that express these three markers.

      Reviewer #3 (Public review):

      Summary:

      In this manuscript, the authors investigated TBRS etiology by using new human pluripotent stem cell models, modeling varying levels of TBRS-associated loss of DNMT3A function. They identified increased lineage-specific proliferation of precursors in TBRS ventral MGE-like progenitors, which they propose was related to increased signaling through the PIK3/AKT/mTOR pathway. Furthermore, they show that reduced DNA methylation during MGE-like progenitor differentiation into GABAergic interneurons can cause a premature expression of neuronal and synaptic genes, triggering precocious neuronal maturation. In conclusion, they propose that TBRS-derived GABAergic neurons exhibit hyperactivity that can alters the development and structure of neuronal networks.

      Strengths:

      Overall, the data presented is convincing, from an early developmental point of view, given that the iPSC-derived 2D cultures or organoids used do not get to reach a mature state. Nonetheless, the data clearly show the effects that deleterious mutations in TBRS can cause during the period of neurogenesis, which was missing in the field.

      Weaknesses:

      (1) Li et al., 2022 (referred to in the manuscript) seems to already show the interplay between H3K27me3 and Dnmt3a discussed in this study i.e., that in the absence of DNA methylation, there is an expansion of polycomb-like repression. These data should be better acknowledged in the paragraph 'Repressive H3K27me3 compensates for severe loss of DNA methylation' (page 9), given it supports the data presented in this manuscript and suggests this as a common mechanism in the interplay between these two repressive marks, as it is well established in the literature.

      We thank the reviewer for this suggestion. We have now added Li et al., 2022 to both the results section (pg. 9, lines 16-20) and our discussion section (pg. 16, lines 12-13).

      (2) The authors should acknowledge that the omics data come from a mixed population of cells.

      We thank the reviewer for their comment. We have validated that the established 2-D differentiation methods we used in this study generate cell populations with >85-90% enrichment for the desired progenitor and neuronal cell type, based upon marker expression, but acknowledge that these are bulk -omics data obtained from cells that may represent a mixed population and have now detailed this in the methods section under “Sequencing” (pg. 21, lines 16-18).

      (3) The authors are encouraged to further discuss whether the overgrowth observed in ventral GABAergic cultures or organoids compares to the overgrowth observed in diseased patients. One expects MRIs to have been performed in patients and that these could be harnessed to discern if overgrowth occurs in the cortex or ventral regions of the brain.

      We thank the reviewer for their suggestion and do note that at least one published study documents increased cortical thickness in the MRIs of TBRS patients (Jiménez de la Peña et al., 2024, PMID: 37795572); however, to our knowledge studies have not examined regional or cell type-selective overgrowth of cortical tissue in TBRS patients. Future clinical studies examining the nature of the neuronal progenitor overgrowth and resulting consequences for patient brain imaging would be of interest to better understand TBRS-associated etiology of brain overgrowth and its manifestations.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) A potential explanation for the ventral organoid's sensitivity compared to the dorsal can be provided. I understand that the PRC2-mediated methylation loss was more pronounced in the ventral than in the dorsal. However, this raises the question: why is this so? Are the PRC2 components more highly expressed in ventral organoids or in GABAergic neurons than in glutamatergic neurons?

      We thank the reviewer for this question. Upon examining RNA-sequencing data for control V-NPCs and D-NPCs, we found that expression of the PRC2 subunits SUZ12, EED, and EZH2 is significantly elevated in D-NPCs relative to V-NPCs (see Author response image 1). This could contribute to restoring epigenetic repression in D-NPCs in the context of DNMT3A mutation, to yield the more limited transcriptomic changes and greater gain of H3K27me3 we observed in TBRS D-NPCs, despite their comparable levels of global mCG loss). This could also account for our finding that overexpression of EZH2 in V-NPCs could normalize some gene expression changes observed in the TBRS models.

      Author response image 1.

      (2) DNMT3A is a major enzyme that mediates non-CpG methylation during synaptogenesis. There is no discussion of this phenomenon or any related analysis. I suppose the brain organoids may be too immature to detect non-CpG methylation. Even then, this can be described in the results sections.

      We thank the reviewer for this comment. While prior work has examined an important role for DNMT3A in postnatal deposition of non-CpG methylation (Christian et al. 2020; Beard et al. 2023), we found that our hPSC models, as the reviewer suggests, were too immature to detect appreciable levels of non-CpG methylation. Accordingly, we have focused our work here on neurodevelopmental requirements for DNMT3A; this work demonstrated that many of the functional alterations of neuronal network activity originate from altered neurogenesis and neuronal maturation of cortical GABAergic interneurons. We have now included a description of the important role of non CpG methylation by DNMT3A during postnatal development and have indicated that the immaturity of hPSC-derived neurons precludes our ability to study non-CpG methylation using these models in the results (pg. 7, lines 8-11) with present language in the discussion (pg. 15, lines 8-11).

      (3) Given the Figure 4 results showing H3K27 methylation and EZH2 knockdown rescue the DNMT3A mutation, the authors argue that EZH2 and DNMT3A regulate a similar set of genes and that the two diseases are related. To claim this, there should be data showing that gene misregulation upon loss of EZH2 and DNMT3A is similar.

      Given our data, it is premature to draw direct relationships with the dysregulated genes we detected in TBRS versus the molecular basis of Weaver Syndrome; therefore, we have attempted to better reflect the potential but currently unproven relationship between these OGIDs and altered the respective language in the results section (pg. 10, lines 10-12) to reflect this. Understanding convergent mechanisms of OGIDs involving both DNMT3A (TBRS) and EZH2 (Weaver Syndrome) gene mutations will be a compelling topic for future studies, as our work here suggests that they may regulate similar gene suites.

      (4) Stronger phenotypes were observed in the R882H mutant compared to the P904L mutant. However, this cannot be translated to the functionality of the mutant protein or disease severity because one is on iPSCs and another is made in hESCs. The limitation should be noted to avoid misleading.

      We agree that the severity of each pathogenic mutation can be influenced by its presence on a different iPSC vs hESC background; accordingly, our study design indicated the hPSC background for each model and employed paired isogenic controls to clearly define the consequences of each TBRS mutation relative to a model with an identical genetic background but lacking the mutation. Our finding that the R882H mutation resulted in more severe epigenomic and transcriptomic consequences than the P904L mutation is congruent with findings made in prior work (Beard et al., 2023 PMID: 37952155; Russler-Germain et al., 2014 PMID: 24656771).

      (5) It is unclear what measurements were done for DNMT3A level quantification shown in Figure 1e- f.

      Protein quantification for DNMT3A models was performed by western blot, shown in Supplemental Fig. S16. We have now included reference to whole blots in the methods section under “Cellular Phenotyping” (page 20 line 13).

      (6) The results section of the paper for Figures 6 and 7 cites the wrong figure numbers and panels.

      We thank the reviewer for catching these errors and have corrected them in the revised manuscript draft.

      (7) Some typos for PI3K (meaning PIK3) in several places, including the Abstract.

      We thank the reviewer for catching these errors and have corrected them in the revised manuscript draft.

      Reviewer #2 (Recommendations for the authors):

      (1) There is a figure numbering discrepancy in the manuscript - the text references six main figures, but the figure pages include seven, with the MEA network data apparently mislabeled.

      We thank the reviewer for catching these errors and have corrected them in the new manuscript draft.

      (2) The nomenclature "D-IN" for dorsal immature neurons is potentially misleading, as "IN" conventionally denotes interneurons, which these glutamatergic cells are not.

      While we agree that IN could be interpreted as interneurons, this nomenclature is defined at an early point in the manuscript and was used to allow readers to easily identify compare findings made in D-NPCs and D-INs (and, as a counterpart, findings made in V-NPCs versus V-INs).

    1. eLife Assessment

      This important work presents a novel computational framework for modeling macroscopic traveling waves in the mouse cortex by integrating open-source connectomic and transcriptomic data into a spiking network model. This approach allows the computational model to assign excitatory/inhibitory connections based on neurotransmitter profiles and extends simulations to the 3D domain. The authors present results that demonstrate how spatiotemporal dynamics such as slow oscillations (0.5-4 Hz) emerge and self-organize at the whole-brain scale. This study provides convincing initial insights into the structural basis of traveling waves at the whole-brain scale, and allows future links between connectome-driven spiking neural networks and whole-hemisphere imaging in the mouse.

    2. Reviewer #1 (Public review):

      I thank the authors for their thoughtful and thorough responses, which address my concerns. Their two methodological changes: (1) the switch to Poisson stimulation and (2) the new LFP estimation pipeline, together with the expanded parameter-grid sweep and Kuramoto synchrony analysis, substantially strengthen the manuscript. The Poisson spike train better approximates the stochastic subcortical drive cortex receives in vivo and removes the artificiality of the original protocol (Point 1.2). The LFP pipeline directly resolves my concern about the disconnect between simulated voltages and experimental signals; showing that the macroscopic wave structure persists in the LFP-like proxy clarifies the framework's practical relevance (Point 1.6). The expanded per-band sweep addresses my worry that the Allen-connectivity advantage was confined to a narrow regime, and acknowledging the small delta-band difference is a more convincing presentation (Point 1.5). The Kuramoto analysis connects dynamics across scales and gives a clear, quantitative account of the non-monotonic coupling dependence (Points 1.4, 1.7). Finally, I appreciate that the remaining connectivity-realism issues (Points 1.3, 1.8) are now stated explicitly as limitations with concrete future directions. I agree that incorporating them is beyond the scope of the present study, and their upfront acknowledgement is appropriate.

    3. Reviewer #2 (Public review):

      Summary:

      This work presents a spiking network model of traveling waves at the whole-brain scale in mouse neocortex. The authors use data from the Allen Institute to re-construct connectivity between different neocortical sites. They then quantify macroscopic traveling waves following stimulation of all layer 4 neurons in neocortex.

      Strengths:

      Overall, the results are interesting and shed new light on the dynamic organization of activity across neocortex of the mouse. The paper uses realistic neuron models specifically fit to intracellular recordings, demonstrating that traveling waves occur in the mouse neocortex with both realistic connectivity and realistic single-neuron dynamics. The paper is also well-written in general. For these reasons, the authors have generally achieved their aims in this work.

      Weaknesses:

      (1) Description of Algorithm 1: While the Methods section clearly explains the density parameter \rho, the statement on line 358 concerning the "ideal" average number of connections is a little unclear. The authors should explicitly clarify that \rho is a free parameter that can be adjusted to balance computational feasibility (for a given set of computational resources) and biological fidelity.

      (2) Lines 102-103: The \rho parameter used here results in approximately 300 connections per neuron on average. The authors should state clearly that the number of connections per cell is the key determinant of computational feasibility (cf. Morrison et al., Neural Computation, 2005). The authors should also review neuronal density and synaptic connectivity in mouse neocortex and clearly reference density and connectivity in their model to the biological scales found in the mouse.

      (3) Line 131: From the plots in Figure 2, it is not clear that the stimulus response is necessarily a rhythmic oscillation, in the sense of a single narrowband frequency.

      (4) Line 217: Can the authors clarify how these findings relate to the results from Mohajerani et al. (Nature Neuroscience, 2013), or differ from them?

      (5) Line 230: Because higher temporal frequency activity also tends to be more spatially localized, a correlation between PGD and temporal frequency could be an inherent consequence of this relationship, rather than a meaningful result.

      (6) Line 247-248: It is not clear that the algorithm for generating connections between neurons presented here really relates to those for community detections. For example, in the case of the Allen Institute data, the communities are essentially in the data already.

      (7) Line 284-285: The relationship between conduction delay is more direct than this sentence suggests. Conduction delay is fundamentally determined by the time required for action potentials to propagate along axons, making it intrinsically linked to anatomical distance.

      (8) Line 287-288: The authors suggest at this point that they do not have enough information to estimate time delays due to axonal conduction along white matter fibers. However, experimental data from white matter connections typically includes information about fiber length, which does enable estimating conduction delays. These estimations have been previously implemented for Allen Institute connectome data in the mouse (Choi and Mihalas, PLoS Comput Biology, 2019) and human connectome data (Budzinski et al., Physical Review Research, 2023).

      (9) Lines 294-295: Several methods do exist for detecting and characterizing wave dynamics in three-dimensional data (Budzinski et al., Physical Review Research, 2023).

      Comments on revised version.

      In this response and revised manuscript, the authors have addressed all points raised in the first round of review. In response to Point 2.7, however, is it not the case that the Allen dataset has the axonal lengths?

    4. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      (1.1) The manuscript “Realistic coupling enables flexible macroscopic traveling waves in the mouse cortex” by Sun, Forger, and colleagues presents a novel computational framework for studying macroscopic traveling waves in the mouse cortex by integrating realistic brain connectivity data with large-scale neural simulations.

      The key contributions include: (1) developing an algorithm that combines spatial transcriptomic data (providing detailed neuron positions and molecular properties) with voxelized connectivity data from the Allen Brain Atlas to construct neuron-to-neuron connections across 300,000 cortical neurons; (2) building a GPU-accelerated simulation platform capable of modeling this large-scale network with both excitatory and inhibitory HodgkinHuxley neurons; (3) extending phase-based analysis methods from 2D to 3D to quantify traveling wave activity in the realistic brain geometry; and (4) demonstrating that realistic Allen connectivity generates significantly higher levels of macroscopic traveling waves compared to simplified local or uniform connectivity patterns.

      The study reveals that wave activity depends non-monotonically on coupling strength and that slow oscillations (0.5-4 Hz) are particularly conducive to large-scale wave propagation, providing new insights into how anatomical connectivity enables flexible spatiotemporal dynamics across the cortex.

      The authors leverage two existing dense datasets of spatial transcriptomic data and connection strength between pairwise voxels in the mouse cortex in a novel way, allowing for the computational model to capture molecular and functional properties of neurons as determined by their neurotransmitter profiles, rather than making arbitrary assignments of excitatory/inhibitory roles. Additionally, the author’s expansion of 2D phase dynamics to 3D phase gradient analysis methods is important and can be widely applied to calcium imaging, LFP recordings, and likely other electrophysiological recordings.

      Thank you for the accurate summary of our manuscript and list of strengths.

      (1.2) The model’s Allen connectivity approach overlooks critical aspects of real cortical dynamics. Most importantly, it excludes subcortical structures, especially the thalamus, which drives cortical traveling waves through thalamocortical interactions. The authors’ method of electrically stimulating all layer 4 neurons simultaneously to initiate waves is artificially crude and bears little resemblance to natural wave generation mechanisms.

      We agree that excluding subcortical structures, especially the thalamus, is an important limitation of the current model. Because adding these structures would substantially expand the scope and complexity of the model, we now state this limitation more explicitly in the Discussion and leave it as a future extension:

      “For simulations, we choose to randomly stimulate the total population of layer 4 neurons as a way to mimic subcortical input and generate traveling waves, which can be unrealistic. Subcortical structures, such as the thalamus, are vital to cortical dynamics like slow-wave activity [1] and are known to regulate traveling waves [2]. Therefore, a direct and important future improvement would be adding subcortical structures to the model.”

      We also agree that the original constant-current stimulation was too artificial. We therefore replaced it with a 10Hz Poisson spike train delivered to layer-4 excitatory neurons across the isocortex, which more closely mimics the stochastic input that cortex receives from subcortical regions such as the thalamus. The revised stimulation protocol is described in the Results:

      “Therefore, to mimic the stochastic drive that cortex receives from subcortical regions like thalamus, we deliver a 10Hz Poisson spike train to layer-4 excitatory neurons (Figure 2a), since layer 4 is the canonical thalamocortical input layer [3]; each Poisson event applies a fixed voltage bump V<sub>stim</sub>, and no other external input is applied.”

      as well as in the Methods 4.5 Simulation.

      The revised stimulation protocol allowed us to rerun the parameter sweep under stochastic drive. A direct comparison of alternative wave-generating mechanisms remains an important direction for future work.

      (1.3) The model handles voxel-to-voxel connections crudely when neurons have mixed excitatory/inhibitory properties and varying synaptic strengths. Real connectivity differs dramatically between neuron types (pyramidal cells vs. interneurons, across cortical layers), but the model only distinguishes excitatory and inhibitory neurons. Additionally, uniform synaptic weights ignore natural variations in connection strength based on neuron type, distance, and functional role. Integrating the updated thalamocortical dataset mentioned by the authors, even at regional resolution, would substantially improve the model.

      We thank the reviewer for raising these important points regarding cell-type-specific connectivity and heterogeneity of synaptic weights. We agree that cell-type-specific connectivity and heterogeneous synaptic weights are important limitations. Because the current voxelized projectome is not cell-type specific, we now state these limitations explicitly and outline how future versions of the model could incorporate improved density and synaptic weight assumptions in the Discussion (Construction of the Allen connectivity). Specifically, we now write:

      “First, the voxelized projection data is not cell-type specific. In our model, we distinguish only two neuronal populations: glutamatergic (excitatory) and GABAergic (inhibitory), based on the Zhuang-ABCA-1 transcriptomic dataset. However, real cortical connectivity differs dramatically between more refined cell types: pyramidal neurons and interneurons have distinct connection targets, and connectivity is strongly layer-specific. At the same time, we use a single global density parameter ρ to set the average number of connections per neuron across the entire cortex. A region-specific ρ could better capture the known variation in local synaptic density across cortical areas, for example the higher synaptic density of primary sensory regions relative to higher-order association areas. Future versions of the model could allow a region- and cell-type-specific ρ derived from region-level synapse-density atlases together with cell-type-resolved connectivity [4], which would simultaneously address the limitations noted above.

      Also, our model uses uniform synaptic weights within each synapse type: every AMPA synapse has conductance g<sub>AMPA</sub> and every GABA synapse has conductance g<sub>GABA</sub>. In reality, synaptic strength varies with presynaptic and postsynaptic cell type, anatomical distance, and the distribution of synaptic weights is typically heavy-tailed (e.g. log-normal [5]). Extending our algorithm to sample synaptic weights from realistic distributions would be a natural next step and is likely necessary for quantitative comparisons with electrophysiological recordings.”

      These additions clarify which aspects of the current model are constrained by the available voxelized projectome and which extensions would require cell-type-resolved connectivity data, region-specific density information, or more detailed synaptic-weight estimates.

      (1.4) While the authors bridge microscopic (single neuron) and mesoscopic (regional connectivity) data to study macroscopic (whole-cortex) waves, they don’t integrate the distinct mechanisms operating at each scale. The framework demonstrates that realistic connectivity enables macroscopic waves but fails to connect how wave dynamics emerge and interact across spatial scales systematically.

      We thank the reviewer for this insightful comment. In the revision, we added the Kuramoto synchrony analysis as a first step toward connecting scales: changes in the microscopic coupling parameter alter network synchrony, and this synchrony measure is closely associated with the macroscopic PGD observable (new Fig. 5). We agree that a full separation of layer-specific microcircuit mechanisms, region-specific connectivity motifs, and whole-cortex wave propagation remains beyond the scope of the present study. The revised text now frames the synchrony-to-wave relationship as one concrete cross-scale link that can be investigated further with this framework.

      (1.5) Claims that Allen connectivity produces higher phase gradient directionality (PGD) than local connectivity appear limited to delta oscillations at very specific coupling strengths and applied currents. Few parameter combinations show significantly higher PGD for Allen connectivity, and these are generally low PGD values overall.

      We agree that the original comparison did not sufficiently establish whether the Allen-connectivity advantage held beyond a small number of parameter choices.

      In the revised manuscript, we expanded the analysis to a 10 × 12 grid of excitatory coupling strengths and Poisson stimulus magnitudes for Allen, local, and uniform connectivity. Across this full (g<sub>AMPA</sub>,V<sub>stim</sub>) grid, Allen connectivity shows higher per-band maximum PGD than local or uniform connectivity, especially in the theta, alpha, and beta bands (Fig. 5b). The delta-band difference is small, consistent with the reviewer’s observation that the original delta-band result did not clearly separate Allen from local connectivity.

      In the revised manuscript, Fig. 5c and Fig. 5d show how PGD varies with stimulus magnitude and coupling strength individually. The full per-band PGD heatmaps for all three connectivities, together with the per-band Allen-minus-opponent gap bars, are provided in Figure 4-figure supplement 6. These additions show that the Allen-connectivity trend is not limited to a single representative operating point.

      (1.6) Broadly, it’s unclear how this computational framework can study memory, learning, sleep, sensory processing, or disease states, given the disconnect between simulated intracellular voltages and the local field potentials or other electrophysiological measurements typically used to study cortical traveling waves. While computationally impressive, the practical research applications remain vague.

      We thank the reviewer for this important point. To bridge the gap between our simulations and experimentally measured data, such as local field potentials (LFP), we used a post-hoc LFP estimation pipeline and added a dedicated Methods subsection describing it (LFP estimation from intracellular voltage). The key idea is that LFP primarily reflects the net transmembrane synaptic current in a local population, which we can reconstruct directly from the voltage traces and the connectivity used in the simulation. Please see Methods section LFP estimation from intracellular voltage.

      This pipeline allows us to test whether the macroscopic traveling-wave structure identified in the voltage traces is also present in an LFP-like signal. In the revised manuscript, we added a new Results section and show, in Fig. 6, side-by-side voltage- and LFP-based snapshots, the voltage-LFP PGD scatter (Pearson r = 0.78-0.89 per band), and per-band PGD comparisons across Allen, local, and uniform connectivity on the LFP signal. These results indicate that our conclusions are not restricted to intracellular voltage and provide a closer bridge to LFP-based experimental measurements.

      (1.7) The paper needs a clearer explanation for why medium coupling (100%) eliminates waves in Allen connectivity (Figure 6) while stronger coupling (150%) restores them.

      Thank you for requesting this clarification. The revised analysis suggests that the non-monotonic relationship between coupling strength and wave activity reflects an interaction between network synchrony and spatial organization. At weak coupling, the network has enough coordination to support propagating waves. At medium coupling, increased synaptic drive pushes the network into an asynchronous irregular state that disrupts coherent wave fronts. At strong coupling, rhythmic synchronization is re-established and again supports wave propagation.

      To substantiate this explanation quantitatively, we measured the Kuramoto order parameter R(t)=|〈e<sup>iϕ(x,t)</sup>〉<sub>x</sub>| from the generalized phase ϕ(x, t) of the band passed voltage field (see updated Quantitative measurement of neuronal activity in Methods) and reduced it to the maximum over each 1-s recording window. We then swept the same ten g<sub>AMPA</sub> values (0.005-0.050 nS) and twelve stimulus magnitudes used for the PGD sweeps, for Allen, local and uniform connectivity, in all five frequency bands. The new analysis is presented in Fig. 5: panel (e) shows synchrony versus g<sub>AMPA</sub> for the three connectivities, panel (d) shows the matching PGD curve, and panel (f) shows the synchrony-PGD scatter with Pearson r per connectivity.

      The synchrony curve captures the main peak-dip-recovery structure of the PGD curve. For Allen connectivity in the alpha band, mean synchrony peaks at weak coupling, collapses to a 75 % suppression in the medium-coupling window g<sub>AMPA</sub> = 0.020-0.035 nS, and recovers near unity at g<sub>AMPA</sub> ≥ 0.040 nS. Uniform connectivity follows the same U-shape with a slightly earlier dip. Per-band versions of the PGD- and synchrony-vs-coupling curves are provided in Figure 5-figure supplement 1, showing that the peak-dip-recovery profile holds across all five canonical bands for Allen and uniform connectivity. Across the 120- point (g<sub>AMPA</sub>, V<sub>stim</sub>) grid, synchrony and PGD are positively correlated in every band and every connectivity (Pearson 0r ranging from ≈ 0.30 to ≈ 0.92 across band-connectivity combinations; per-band scatters in Figure 5-figure supplement 2). Local connectivity follows a different trajectory: its synchrony is moderate at weak coupling and decays monotonically with gAMPA without recovering at strong coupling. This is consistent with local connectivity not supporting large-scale propagating waves at strong coupling, so the peak-dip-recovery interpretation applies mainly to Allen and uniform connectivity. We have added this synchrony analysis to the revised manuscript:

      “To diagnose the mechanism behind this profile, we measured the Kuramoto order parameter R(t) = |⟨e<sup>iϕ(x,t)</sup>⟩<sub>x</sub>| from the generalized phase field ϕ(x, t) of each frequency band and recorded its maximum over each simulation window (Quantitative measurement of neuronal activity). The resulting synchrony curve (Figure 5e) resembles the trend of the maximum PGD well (Figure 5d). For Allen connectivity in the alpha band, mean synchrony peaks at weak coupling, collapses in the medium-coupling window, and recovers at strong coupling. Across the full 120-point (g<sub>AMPA</sub>, V<sub>stim</sub>) grid, synchrony and PGD are positively correlated in every band and every connectivity (Pearson r = 0.30- 0.92, all p < 10−3; Figure 5f, with per-band scatters in figure Supplement 2). Therefore, the PGD trough at medium coupling may be a synchrony trough: increased synaptic drive pushes the network into an asynchronous irregular state, while strong coupling re-establishes rhythmic synchronization that supports wave propagation. This synchrony-to-wave bottleneck seems more significant to the networks with long-range connectivity (Allen, uniform), partly because long-range connections can augment synchrony in the coupled neuronal network [6]. We also note that although uniform connectivity is able to achieve almost an identical level of synchrony to that of Allen connectivity, the PGD remains much lower due to the loss of spatial organization within.”

      (1.8) Does using a single connectivity parameter (ρ = 300) across all regions miss important regional differences in cortical connectivity density?

      We thank the reviewer for raising this important point. We agree that a single global density parameter misses region-to-region variation in local synaptic density, and we have extended the Discussion (Construction of the Allen connectivity) to state this limitation alongside the cell-type-specific connectivity and synaptic-weight limitations discussed in response to Point 1.3. Specifically, we now write:

      “At the same time, we use a single global density parameter ρ to set the average number of connections per neuron across the entire cortex. A region-specific ρ could better capture the known variation in local synaptic density across cortical areas, for example the higher synaptic density of primary sensory regions relative to higher-order association areas. Future versions of the model could allow a region- and cell-type-specific ρ derived from region-level synapse density atlases together with cell-type-resolved connectivity [4], which would simultaneously address the limitations noted above.”

      This paragraph is placed directly after the cell-type and weight-heterogeneity limitations added in response to Point 1.3.

      Reviewer #2 (Public review):

      (2.1) This work presents a spiking network model of traveling waves at the whole-brain scale in the mouse neocortex. The authors use data from the Allen Institute to reconstruct connectivity between different neocortical sites. They then quantify macroscopic traveling waves following stimulation of all layer 4 neurons in the neocortex.

      Overall, the results are interesting and shed new light on the dynamic organization of activity across the neocortex of the mouse. The paper uses realistic neuron models specifically fit to intracellular recordings, demonstrating that traveling waves occur in the mouse neocortex with both realistic connectivity and realistic single-neuron dynamics. The paper is also well-written in general. For these reasons, the authors have generally achieved their aims in this work.

      We thank Reviewer 2 for the positive assessment and accurate summary of our work.

      (2.2) Description of Algorithm 1: While the Methods section clearly explains the density parameter ρ, the statement on line 358 concerning the “ideal” average number of connections is a little unclear. The authors should explicitly clarify that ρ is a free parameter that can be adjusted to balance computational feasibility (for a given set of computational resources) and biological fidelity. The ρ parameter used here results in approximately 300 connections per neuron on average. The authors should state clearly that the number of connections per cell is the key determinant of computational feasibility (cf. Morrison et al., Neural Computation, 2005). The authors should also review neuronal density and synaptic connectivity in the mouse neocortex and clearly reference density and connectivity in their model to the biological scales found in the mouse.

      We thank the reviewer for raising this important point about our connectivity algorithm and simulation. We have clarified in the revised Methods (Use of the voxelized connectivity data, Methods 4.2) that ρ is a free parameter that controls the trade-off between computational feasibility and biological fidelity. Specifically, we now write:

      “The density parameter ρ is a free parameter that controls the trade-off between computational feasibility and biological fidelity: higher values of ρ yield more connections per neuron and thus higher biological realism, at the cost of greater memory and runtime [7].”

      A careful accounting of the biological scales involved (synapse density per neuron, total cortical population) and incorporating region or cell-type-specific connection density is left as a future direction. We have noted this in the Discussion subsection:

      “First, the voxelized projection data is not cell-type specific. In our model, we distinguish only two neuronal populations: glutamatergic (excitatory) and GABAergic (inhibitory), based on the Zhuang-ABCA-1 transcriptomic dataset. However, real cortical connectivity differs dramatically between more refined cell types: pyramidal neurons and interneurons have distinct connection targets, and connectivity is strongly layer-specific. At the same time, we use a single global density parameter ρ to set the average number of connections per neuron across the entire cortex. A region-specific ρ could better capture the known variation in local synaptic density across cortical areas, for example the higher synaptic density of primary sensory regions relative to higher-order association areas. Future versions of the model could allow a region- and cell-type-specific ρ derived from region-level synapse-density atlases together with cell-type-resolved connectivity [4], which would simultaneously address the limitations noted above.”

      (2.3) Line 131: From the plots in Figure 2, it is not clear that the stimulus response is necessarily a rhythmic oscillation, in the sense of a single narrowband frequency.

      The reviewer is correct, and we are grateful for the prompt to be more precise. The Results phrasing around Fig. 2 has been softened to avoid any implication of narrowband rhythmicity, and now reads:

      “Under this protocol, we immediately observe macroscopic traveling waves emerge across the cortex (Figure 2 and Videos). The global mean voltage and the region-sorted raster (Figure 2c, d) reveal oscillatory activity that is well synchronized across regions, while the local-mean intracellular voltage maps over a representative 50ms window (Figure 2b) reveal a coherent wavefront sweeping along the anterior-posterior axis, consistent with previously reported cortex-wide waves [8, 9, 10]. The corresponding single-neuron-resolution view of the same simulation, with no spatial averaging, is shown in figure Supplement 1.”

      Moreover, we have revised the Introduction to describe [11] as demonstrating traveling waves in broadband (5-40Hz) activity, making clear that traveling waves can occur without requiring narrowband oscillations (see also our response to your related point below). In the revised manuscript, we also separate the broadband activity into canonical frequency bands and analyze the wave activity in each band independently.

      (2.4) Line 217: The authors should clarify how these findings relate to the results from Mohajerani et al. (Nature Neuroscience, 2013) or differ from them.

      We thank the reviewer for this suggestion. We agree that [12] is an important experimental benchmark. A direct quantitative comparison is difficult because the original data were not aligned to the Allen Brain Atlas CCF used in our simulations. We therefore revised the Discussion to identify this comparison as a future direction, alongside the visual-cortex bidirectional-wave data of [13]:

      “A more detailed quantitative comparison with experimental cortical-wave studies, such as the cortex-wide voltage-imaging data of [12] or the bidirectional visual-evoked waves reported by [13], is left as a future direction.”

      (2.5) Line 230: Because higher temporal frequency activity also tends to be more spatially localized, a correlation between PGD and temporal frequency could be an inherent consequence of this relationship, rather than a meaningful result.

      We thank the reviewer for raising this important point. The reviewer is correct that higher-frequency oscillations tend to be more spatially localized, which can inherently reduce PGD when measured globally. We therefore revised this analysis by separating the broadband signals into canonical frequency bands and comparing PGD within each band.

      In the revised manuscript, we no longer interpret cross-band PGD differences as evidence for a frequency-to-spatial-scale relationship. Instead, we report the level of macroscopic wave activity within each canonical band. This per-band comparison is summarized in Fig. 5b, which reports the maximum PGD in each band (mean ± SEM across the entire (g<sub>AMPA</sub>,V<sub>stim</sub>) grid) for the three connectivities. Allen connectivity shows higher per-band PGD than local or uniform connectivity in the theta, alpha, and beta bands, without requiring an interpretation of PGD differences across frequency bands.

      For completeness, we also computed the per-band mean(Allen) − mean(opponent) PGD gap across the full parameter grid (10 g<sub>AMPA</sub> × 12 V<sub>stim</sub> = 120 points per connectivity). The result is presented in Figure 4-figure supplement 6: panel (a) gives the per-band maxPGD heatmaps over the (gAMPA, Vstim) grid for each connectivity, and panel (b) gives the per-band Allen-minus-opponent mean-PGD gap. The mean gap against Local is −0.004 in delta, +0.027 in theta, +0.036 in alpha, +0.036 in beta and +0.004 in gamma, and against Uniform is +0.015, +0.044, +0.052, +0.040 and +0.011, respectively. The gap is largest in alpha and broadly concentrated in theta-alpha-beta, rather than in delta as we had originally written. We report these per-band gaps descriptively because they summarize one parameter sweep per connectivity rather than independent biological or simulation replications, and we do not interpret the differences across bands as a meaningful frequency dependence. We have added this analysis to the revised manuscript.

      (2.6)Line 247-248: It is not clear that the algorithm for generating connections between neurons presented here really relates to those for community detection. For example, in the case of the Allen Institute data, the communities are essentially in the data already.

      We agree with the reviewer that the relevant anatomical blocks are already present in the data. Our intent was to relate the sampling procedure to stochastic block models for network generation, not to community-detection algorithms. We have revised this passage in the Discussion subsection Construction of the Allen connectivity: ”In essence, our algorithm belongs to the family of stochastic block models [14], where the block structure is given by the voxelization of the Allen Brain Atlas and the inter-block connection probabilities are set by the voxelized projection strengths.”

      (2.7) Line 284-285: The relationship between conduction delay is more direct than this sentence suggests. Conduction delay is fundamentally determined by the time required for action potentials to propagate along axons, making it intrinsically linked to anatomical distance.

      Thank you for raising this important point. We agree that conduction delay is directly tied to axonal propagation time and therefore to anatomical path length. Our original wording was intended to note that the relevant path length cannot be approximated reliably by Euclidean distance in the 3-D coordinate space. We have revised this passage in the Discussion sub-section Cortical model and simulation to state that conduction delay is linked to the white-matter path length of the connection, while noting that our model lacks the actual axonal path geometry through the cortical manifold:

      “However, we did not include conduction delay in our study. Conduction delay is thought to have a proportional relationship with the white matter path length of the connection between two neurons [15]. In our model, although we have the 3D positions of neurons, we do not have geometric information about the cortical manifold. Two neurons can be very close in the Cartesian coordinates measured by the Euclidean distance, but very far in terms of the length of the actual connection in the brain. Therefore, incorporating accurate conduction delay in the model is an important future direction.”

      (2.8) Lines 294-295: Several methods do exist for detecting and characterizing wave dynamics in three-dimensional data (Budzinski et al., Physical Review Research, 2023).

      Thank you for this reference. We have added a citation to [16] in the Discussion subsection Quantitative measurements of 3-D traveling waves, acknowledging that methods for 3-D wave analysis do exist while noting that most published algorithms are designed for 2-D data:

      “There are many techniques available for identifying and measuring large-scale neuronal spatiotemporal patterns [17, 18]. While methods for detecting wave dynamics in three-dimensional data do exist [16], most published algorithms are designed to analyze 2-D data, so our simulation data, which is intrinsically 3-D, presents new challenges for measurement.”

      (2.9) Line 28: It is important to note that the Davis et al. (2020) reference is not actually in the beta band, but instead in the broadband (5-40 Hz). This distinction is important because it demonstrates that waves can occur in neural data without requiring narrowband oscillations.

      Thank you for this correction. We have moved the [11] citation out of the beta-band group in the Introduction and reframed it as a broadband (5-40Hz) reference, so that the sentence now reads:

      “These waves are observed at different frequencies during various brain activities, ranging from slow-wave activity [8, 9], sleep spindles [19], to faster oscillations in alpha [20, 21], beta [22, 23], and gamma [21, 13] frequency bands, as well as in broadband (5-40Hz) activity [11].”

      This distinction is important because it makes clear that traveling waves do not require narrowband oscillations. The revised manuscript therefore separates the broadband activity into canonical frequency bands and compares wave activity within each band.

      (2.10) Line 46-49: This sentence could be clearer, for example, by specifying “certain dynamics” in more precise terms.

      We apologize for the imprecision in the original manuscript. We have clarified the sentence in the Introduction to specify that the dynamics of interest are coexistence patterns of local and global activity in the network, as described in the cited reference [24]:

      “It has been shown that in a coupled neuronal network, the coexistence of global wave activity with locally asynchronous states only arises when the number of oscillators is high enough, where local and global activities can coexist [24].”

      (2.11) Figure 1b(i): Small typo in the label for this panel.

      This has been corrected. Thank you for catching it.

      (2.12) Lines 121-122: It may be important to note that spiking neural networks can also generate self-sustained activity (Vogels and Abbott, JNeurosci, 2005; Kumar et al., Neural Computation, 2008). This self-sustained activity is a form of internally generated “frozen” noise that is fundamentally different from externally imposed noise sources (such as Poisson external input) (Destexhe and Contreras, Science, 2006). Waves appear in this self-sustained activity, as well (Davis et al., Nature Communications, 2021), supporting the generality of this phenomenon.

      Thank you for pointing out these important references. We have added citations to [25], [26], [27], and [24] in the Results subsection Macroscopic traveling waves emerge from random stimulation through realistic connectivity:

      “Spiking networks of this scale can also generate self-sustained activity in similar regimes [25, 26], which differs fundamentally from externally imposed noise [27], and traveling waves have been reported under such conditions [24].”

      In the revised manuscript, we also changed the stimulation protocol from constant current to Poisson input, which is more similar to the stochastic input that cortical neurons receive in vivo. Traveling waves observed in our model under this Poisson drive therefore complement, rather than depend on, the self-sustained-activity regime emphasized by the cited works.

      (2.13) Lines 139-140: “anterior-posterior macroscopic waves in both directions” and ”in the reverse direction right after each other” could be clearer. In addition, the study from Aggarwal et al. (Nature Communications, 2022) could be relevant to note at this point.

      We thank the reviewer for these wording suggestions and the relevant reference. In the revised manuscript, we reran the simulations and updated Figure 2 accordingly. The new representative simulation shown in Fig. 2 emphasizes a coherent wavefront sweeping along the anterior-to-posterior axis, and the original passages describing consecutive opposite-direction waves have been removed from the Results subsection Macroscopic traveling waves emerge from random stimulation through realistic connectivity, which now reads:

      “Under this protocol, we immediately observe macroscopic traveling waves emerge across the cortex (Figure 2 and Videos). The global mean voltage and the region-sorted raster (Figure 2c, d) reveal oscillatory activity that is well synchronized across regions, while the local-mean intracellular voltage maps over a representative 50ms window (Figure 2b) reveal a coherent wavefront sweeping along the anterior-posterior axis, consistent with previously reported cortex-wide waves [8, 9, 10]. The corresponding single-neuron-resolution view of the same simulation, with no spatial averaging, is shown in figure Supplement 1.”

      Bidirectional propagation can still occur in the model, but we have chosen not to present it as a focal result of this revised manuscript; consequently, the specific phrasings flagged by the reviewer no longer appear in the Results text.

      We agree that [13] is an important reference, and we now discuss it in the Comparing with experimental data subsection as a potential benchmark for future quantitative comparison with our model.

      (2.14) Figure 7: The caption for this figure could be clearer.

      Line 281: typo ”Ermentrou”.

      Line 418: typo ”excitatory”.

      The typos (“Ermentrou” → “Ermentrout”, and the “excitatory” typo at line 418) have been corrected. The original Figure 7 has been removed from the revised manuscript; its content (PGD versus stimulus magnitude and coupling strength, and PGD by dominant frequency bucket) has been reorganized across the new Figs. 4-6.

      References

      (1) Steriade M, Mccormick DA, Sejnowski TJ. Thalamocortical Oscillations in the Sleeping and Aroused Brain. Science. 1993;262(5134):679-85. Available from: <GotoISI>:// WOS:A1993MD95200029.

      (2) Ye Z, Bull MS, Li A, Birman D, Daigle TL, Tasic B, et al. Brain-wide topographic coordination of traveling spiral waves. BioRxiv. 2023:2023-12.

      (3) Rockland KS.What do we know about laminar connectivity? Neuroimage. 2019;197:772-84.

      (4) Harris JA, Mihalas S, Hirokawa KE, Whitesell JD, Choi H, Bernard A, et al. Hierarchical organization of cortical and thalamic connectivity. Nature. 2019;575(7781):195+. Available from: <GotoISI>://WOS:000496159900061https://www.nature.com/articles/ s41586-019-1716-z.pdf.

      (5) Buzs´aki G, Mizuseki K. The log-dynamic brain: how skewed distributions affect network operations. Nature Reviews Neuroscience. 2014;15(4):264-78.

      (6) Bazhenov M, Rulkov NF, Timofeev I. Effect of synaptic connectivity on long-range synchronization of fast cortical oscillations. Journal of neurophysiology. 2008;100(3):156275.

      (7) Morrison A, Aertsen A, Diesmann M. Spike-timing-dependent plasticity in balanced random networks. Neural computation. 2007;19(6):1437-67.

      (8) Massimini M. The Sleep Slow Oscillation as a Traveling Wave. Journal of Neuroscience. 2004;24(31):6862-70. Available from: https://dx.doi.org/10.1523/jneurosci. 1318-04.2004 https://www.ncbi.nlm.nih.gov/pmc/articles/PMC6729597/pdf/ 0246862.pdf.

      (9) Liang Y, Song C, Liu M, Gong P, Zhou C, Kno¨pfel T.Cortex-Wide Dynamics ofIntrinsic Electrical Activities: Propagating Waves and Their Interactions. The Journal of Neuroscience. 2021;41(16):3665-78. Available from: https://www.jneurosci.org/ content/jneuro/41/16/3665.full.pdf.

      (10) Aggarwal A, Luo J, Chung H, Contreras D, Kelz MB, Proekt A. Neural assemblies coordinated by cortical waves are associated with waking and hallucinatory brain states. Cell Reports. 2024;43(4):114017. Available from: https://www.sciencedirect.com/ science/article/pii/S2211124724003450.

      (11) Davis ZW, Muller L, Martinez-Trujillo J, Sejnowski T, Reynolds JH. Spontaneous travelling cortical waves gate perception in behaving primates. Nature. 2020;587(7834):4326. Available from: https://doi.org/10.1038/s41586-020-2802-y.

      (12) Mohajerani MH, Chan AW, Mohsenvand M, LeDue J, Liu R, McVea DA, et al. Spontaneous cortical activity alternates between motifs defined by regional axonal projections. Nature Neuroscience. 2013;16(10):1426-35. Available from: https://doi.org/ 10.1038/nn.3499https://www.nature.com/articles/nn.3499.pdf.

      (13) Aggarwal A, Brennan C, Luo J, Chung H, Contreras D, Kelz MB, et al. Visual evoked feedforward–feedback traveling waves organize neural activity across the cortical hierarchy in mice. Nature Communications. 2022;13(1):4754. Available from: https://doi.org/10.1038/s41467-022-32378-xhttps://www.nature. com/articles/s41467-022-32378-x.pdf.

      (14) Holland PW, Laskey KB, Leinhardt S. Stochastic blockmodels: First steps. Social Networks. 1983;5(2):109-37. Available from: https://www.sciencedirect.com/science/ article/pii/0378873383900217.

      (15) Lemar´echal JD, Jedynak M, Trebaul L, Boyer A, Tadel F, Bhattacharjee M, et al. A brain atlas of axonal and synaptic delays based on modelling of cortico-cortical evoked potentials. Brain. 2022;145(5):1653-67.

      (16) Budzinski RC, Nguyen TT, Min´o-Calero J, Davis ZW, Muller LE. Analyzing transientevoked neural activity in three-dimensional cortical recordings. Physical Review Research. 2023;5(1):013012.

      (17) Townsend RG, Gong P. Detection and analysis of spatiotemporal patterns in brain activity. PLOS Computational Biology. 2018;14(12):e1006643. Available from: https: //doi.org/10.1371/journal.pcbi.1006643.

      (18) Gutzen R, De Bonis G, De Luca C, Pastorelli E, Capone C, Allegra Mascaro AL, et al. A modular and adaptable analysis pipeline to compare slow cerebral rhythms across heterogeneous datasets. Cell Reports Methods. 2024;4(1). Available from: https: //doi.org/10.1016/j.crmeth.2023.100681.

      (19) Muller L, Piantoni G, Koller D, Cash SS, Halgren E, Sejnowski TJ. Rotating waves during human sleep spindles organize global patterns of activity that repeat precisely through the night. Elife. 2016;5. Available from: https://www.ncbi.nlm.nih.gov/pubmed/27855061https://www.ncbi.nlm. nih.gov/pmc/articles/PMC5114016/pdf/elife-17267.pdf.

      (20) Zhang H, Watrous AJ, Patel A, Jacobs J. Theta and alpha oscillations are traveling waves in the human neocortex. Neuron. 2018;98(6):1269-81. e4.

      (21) van Kerkoerle T, Self MW, Dagnino B, Gariel-Mathis MA, Poort J, van der Togt C, et al. Alpha and gamma oscillations characterize feedback and feedforward processing in monkey visual cortex. Proc Natl Acad Sci U S A. 2014;111(40):14332-41. Available from: https://www.ncbi.nlm.nih.gov/pubmed/25205811.

      (22) Bhattacharya S, Brincat SL, Lundqvist M, Miller EK. Traveling waves in the prefrontal cortex during working memory. PLoS Comput Biol. 2022;18(1):e1009827.

      (23) Rubino D, Robbins KA, Hatsopoulos NG. Propagating waves mediate information transfer in the motor cortex. Nature Neuroscience. 2006;9(12):1549-57. Available from: https://doi.org/10.1038/nn1802.

      (24) Davis ZW, Benigno GB, Fletterman C, Desbordes T, Steward C, Sejnowski TJ, et al. Spontaneous travelling waves naturally emerge from horizontal fiber time delays and travel through locally asynchronous-irregular states. Nature Communications. 2021;12(1):6057.

      (25) Vogels TP, Abbott LF. Signal propagation and logic gating in networks of integrate-and-fire neurons. Journal of Neuroscience. 2005;25(46):10786-95.

      (26) Kumar A, Schrader S, Aertsen A, Rotter S. The high-conductance state of cortical networks. Neural Computation. 2008;20(1):1-43.

      (27) Destexhe A, Contreras D. Neuronal computations with stochastic network states. Science. 2006;314(5796):85-90.

    1. eLife Assessment

      This study shows that combining forced cell cycle re-entry with Rbpj deletion enhances Müller glia dedifferentiation and promotes their conversion into retinal neuron-like cells in the uninjured mouse retina. It provides a valuable strategy for improving Müller glia-mediated neurogenesis and advancing regenerative potential in the mammalian retina. Overall, the data are convincing. The authors have also addressed concerns regarding Müller glia function, cell survival, and the limitations of neuronal maturation and integration, further strengthening the conclusions of the study.

    2. Reviewer #2 (Public review):

      Summary:

      The inability of the mammalian retina to regenerate poses a major clinical challenge. Much has been learned about the regenerative potential of the retina from teleost fish, where Müller glia (MG) are able proliferate and produce new neurons after injury. However, MG do not retain this potential in the mammalian retina. The authors showed previously that that forcing MG to re-enter the cell cycle by downregulating p27 and upregulating cyclin D1 could induce MG to dedifferentiate, but the results were transient, and these cells eventually reverted back to MG and did not form neurons. Here they expand on this to show that in MG, coupling forced cell cycle re-entry with deletion of Rbpj, which inhibits of the transcriptional effects of Notch signaling, induces some MG to proliferate and take on features of multiple cell types, including MG precursor cells, amacrine-like cells, and bipolar-like cells. This work lends valuable insight into the regenerative potential of mammalian MG, particularly when Notch signaling is manipulated.

      Strengths:

      The major claims of the authors are well-supported. They show convincingly and through multiple methods, including immunostaining, single nucleus RNA sequencing, and in situ hybridization, that coupling notch inhibition with cell cycle re-activation induces the expression of neuronal markers in mammalian MG. The sn-RNA-seq data is particularly valuable in demonstrating the induction of bipolar-cell subtypes. Edu labeling is effective in demonstrating the induction of proliferation, and the long-term viability of the generated neuron-like cells is intriguing.

      Comments on revised version:

      The authors sufficiently addressed all concerns. I particularly appreciate the additional experiments to demonstrate retinal function, and the edits to the text regarding retinal and cell function and retinal organization.

    3. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This study examines Müller glia (MG) reprogramming in the uninjured mouse retina through a combination of Notch signaling inhibition and AAV-induced proliferation. Building on their prior work showing that Cyclin D1 overexpression and p27^Kip1^ knockdown (CCA) promotes MG proliferation with very limited neurogenesis, the authors now demonstrate that Rbpj deletion alone induces a modest degree of MG-to-neuron conversion without proliferation, in agreement with recent work in the field. However, combining Rbpj deletion with CCA-mediated proliferation substantially enhances MG dedifferentiation and the generation of retinal neuron-like cells. Through genetic lineage tracing, histological analyses, and single-cell transcriptomics, the authors provide evidence that MG-derived cells acquire molecular features of bipolar (ON, OFF, and rod bipolar) and amacrine neurons. Most MG-derived cells appear to survive long-term (up to 9 months).

      Strengths:

      Overall, the study is carefully designed and executed, and the manuscript is clearly written with well-presented figures. While the work does not significantly expand the repertoire of neuronal types generated from mammalian MG beyond what has been previously reported in the field, it provides a valuable and improved strategy for inducing robust MG proliferation and neurogenesis in the mammalian retina.

      Weaknesses:

      (1) It would be better to include a negative control AAV when evaluating the effect of CCA AAV in the Rbpj KO background. This could help distinguish the specific contribution of the CCA construct from potential effects of intravitreal AAV injection itself, which can induce mild inflammation, known to influence MG reprogramming.

      To address this concern, in the revised manuscript we included the result from Rbpj KO eyes injected with a negative control AAV (AAV<sub>7m8</sub>-GFAP-GFP) (Fig. S14a). MG reprogramming efficiency, quantified as the proportion of tdT<sup>+</sup>Otx2<sup>+</sup> cells among total tdT<sup>+</sup> cells, was then compared between the AAV-GFP–treated and Rbpj KO–only eyes. At 4 months post-injection, the percentage of tdT<sup>+</sup>Otx2<sup>+</sup> cells in the AAV-GFP–treated eyes was comparable to that of Rbpj KO alone (Fig. S14b–c), and substantially lower than in the CCA-treated eyes. Together, these results indicate that the enhanced MG reprogramming observed in the Rbpj KO+CCA group is driven by transgenes expressed rather than by nonspecific effects of AAV or injection.

      (2) The extent of MG transduction by the CCA AAV is not clear. As quantifications are normalized to total MG (GFP^+^ or TdTomato^+^) or retinal length, it would be useful to clarify whether near-complete transduction is assumed, or if additional information on transduction efficiency can be provided.

      In our previous study (Wu, Liao, et al., 2025, eLife), we have demonstrated that high-dose (4E10vg/injection) AAV7m8 effectively transduced the whole retina, with near-complete MG transduction observed in the vicinity of the injection site, as evidenced by virtually all MG expressing GFP in these regions. In the revised manuscript, we clarified the transduction efficiency in Line 108-110 on Page 5 and Line 625-626 on Page 27.

      (3) In Figure S10, the reduced MG proliferation observed in the CCA + Rbpj deletion group could also potentially reflect decreased GFAP promoter activity in dedifferentiated MG following Rbpj deletion. Alternatively, MG-derived cells may be more fragile under these conditions.

      We thank the reviewer for these excellent insights. We agree that a down-regulation of GFAP promoter activity following Rbpj-mediated dedifferentiation is a highly plausible explanation for the moderate reduction in proliferation, as lower promoter activity would diminish AAV transgene expression. We have included this possibility in the data interpretation (Line 176-178, page 8). Regarding the alternative possibility of increased cell fragility, we agree that cell death cannot be ruled out, but occasional apoptotic cells over a long period of time are difficult to capture experimentally.

      (4) In the CCA + Rbpj deletion condition, do MG undergo single or multiple rounds of cell division?

      We have previously demonstrated that MG typically undergo a single round of cell division in wild type mouse retina following CCA treatment (Wu, Liao, et al., 2025, eLife). Given our observation that Rbpj deletion suppresses CCA-induced MG proliferation (Fig. S11), it is unlikely that the addition of Rbpj deletion would trigger multiple or continuous rounds of cell division beyond the single-round baseline established by CCA alone. While we did not re-evaluate cell division kinetics in the current study, we reason that CCA similarly drives MG to undergo a single round of division in the Rbpj KO context.

      (5) What fraction of neuron-like cells (bipolar- and amacrine-like) arises from proliferation versus direct transdifferentiation? Quantification of MG-derived cells expressing neuronal markers (e.g., Otx2, HuC/D), with and without EdU labeling, would help distinguish these mechanisms.

      The percentages of MG-derived cells expressing neuronal markers with and without EdU labeling, were shown in Fig 3d-e and Fig S19d-e. In the Rbpj KO-only group, neuron-like cells arise exclusively through direct transdifferentiation without cell division, as no EdU incorporation was detected in Rbpj-deficient MG. In this group, a small fraction of MG-derived cells expressed the neuronal marker Otx2 or HuC/D (Fig 3e, Fig S19e). In contrast, the Rbpj KO+CCA group achieved a substantially higher neurogenesis rate, with a significant proportion of Otx2+ or HuC/D+ MG-derived cells also being EdU+ (Fig. 3d, Fig. S19d), indicating that they arose through de novo neurogenesis. By subtracting the contribution of direct transdifferentiation observed in the Rbpj KO-only group, we estimate that majority of MG-derived neuron-like cells in the Rbpj KO+CCA group were generated through proliferation-mediated de novo neurogenesis.

      (6) In Figure S18a, the authors state that "while the neuron-like clusters were best classified as BC-like and AC-like based on their distinct marker gene expression, they also exhibited mixed expression of genes associated with other retinal neuronal types, including RGC markers (e.g., Tubb3, Myt1l, Grin1) and photoreceptor markers (e.g., Crx, Prom1, Epha10, Gucy2e, Scg3) (Fig. S18a), suggesting that the regenerated cells exist in a hybrid state" and "MG derived neuron like cells also expressed genes characteristic of RGCs and photoreceptors, indicating enhanced lineage". However, many of these genes are not specific to RGCs or photoreceptors and are instead broadly expressed in retinal neurons or enriched in bipolar/amacrine populations. Therefore, it is unclear whether these cells exhibit hybrid RGC or photoreceptor identity.

      We thank the reviewer for this insightful comment and for pointing out the need for greater precision in our terminology regarding these markers. While individual markers may lack absolute, 100% cell-type exclusivity, genes such as Tubb3 and Gucy2e serve as widely accepted lineage-associated genes that characterize RGC and photoreceptor programs, respectively (Soto et al., 2008; Sato et al., 2018; Sotani et al., 2024). We have revised the manuscript to replace terms "RGC-specific genes" and "photoreceptor-specific genes" with "RGC signature genes" and "photoreceptor signature genes", respectively. Furthermore, these RGC- and photoreceptor-signature genes are co-expressed across the entire Otx2+ MG population rather than being segregated into distinct, specialized subpopulations (Fig. 4d, Fig. S20). This uniform distribution indicates that these cells possess a hybrid transcriptional program that concurrently incorporates elements of both RGC and photoreceptor identities.

      (7) The authors provide a thorough molecular characterization of MG-derived cells through immunostaining and single-cell sequencing. However, their morphological features, synaptic connectivity (e.g., synaptic marker expression), and electrophysiological properties remain largely uncharacterized. While these experiments may be technically challenging, this limitation should be discussed.

      We agree with the reviewer that characterizing the precise morphological features, synaptic connectivity, and electrophysiological properties of MG-derived cells is a crucial step for any neuronal regeneration study, and we acknowledge that this represents an important limitation of our current study.

      As demonstrated by snRNA-seq data, the MG-derived neuron-like cells exhibit an incompletely mature state, characterized by hybrid transcriptomic signatures. By immunostaining, we did not observe any MG-derived cells with photoreceptor outer segment or typical RGC morphology. Therefore, it is highly likely that these cells have not established functional synaptic connectivity or acquired mature electrophysiological properties. Performing functional or circuitry assessments at this stage would be premature.

      We have added a comprehensive discussion regarding this limitation, along with future directions for long-term functional validation, in the revised manuscript (Line 502-515 on Page 22).

      (8) The conclusion that CCA + Rbpj deletion induces neurogenesis without compromising MG supportive functions or retinal homeostasis appears somewhat oversold. This claim is primarily based on gross retinal morphology and ZO-1 staining. Given the extent of MG dedifferentiation and ectopic cell generation in the ONL and INL, it is likely that retinal function is affected. Functional assessments (e.g., ERG) would be required to support this conclusion. The authors should consider tempering this statement.

      To address the concern raised by the reviewer, we performed electroretinography (ERG) to evaluate both scotopic and photopic retinal function in the Rbpj KO+CCA-treated eyes compared to contralateral untreated controls (Supplementary figure S23e-h). In addition, we conducted optomotor response testing to assess whether visual behavior is affected following treatment (Supplementary Figure S23d). The results demonstrate that combined Rbpj KO and CCA treatment achieves neurogenesis without compromising retinal function.

      (9) Regarding the mechanism by which CCA-induced proliferation enhances MG reprogramming in the Rbpj knockout background, one plausible explanation is that chromatin states (e.g., histone modifications and DNA methylation) are transiently reset during DNA replication and cell division. While this alone may be insufficient to activate neurogenic programs, it could synergize with Rbpj deletion to allow neurogenic transcription factors (such as Ascl1, Otx2, NeuroD1, and NeuroD2) to access previously inaccessible chromatin regions, thereby promoting MG reprogramming.

      We thank the reviewer for the insightful suggestion on the model, which aligns well with our experimental findings. Our snATAC-seq data demonstrate that CCA-induced proliferation broadly increases chromatin accessibility at key neurogenic loci, including Neurod2, Dll1, and Otx2, in active MG compared to resting MG (Figure 6f–h). This chromatin remodeling alone is insufficient to drive neurogenesis, as CCA-only treated MG largely revert to a quiescent glial state. However, when combined with Rbpj deletion, which derepresses downstream neurogenic transcription factors such as Ascl1 and Neurog2 by relieving Notch-mediated transcriptional repression, these newly accessible chromatin regions can be effectively occupied and activated by the available neurogenic factors. The concept that cell division facilitates epigenetic resetting to enhance reprogramming efficiency is well established in the somatic cell reprogramming field, where proliferation rate is directly proportional to reprogramming success by promoting the erasure of lineage-restrictive epigenetic marks and the re-establishment of new transcriptional circuits. In the revised manuscript, we incorporated this mechanistic discussion to provide a more comprehensive interpretation of how proliferation and Notch inhibition converge to promote MG neurogenesis in Line 462-477 on Page 20-21.

      Reviewer #2 (Public review):

      Summary:

      The inability of the mammalian retina to regenerate poses a major clinical challenge. Much has been learned about the regenerative potential of the retina from teleost fish, where Müller glia (MG) are able to proliferate and produce new neurons after injury. However, MG do not retain this potential in the mammalian retina. The authors showed previously that forcing MG to re-enter the cell cycle by downregulating p27 and upregulating cyclin D1 could induce MG to dedifferentiate, but the results were transient, and these cells eventually reverted back to MG and did not form neurons. Here, they expand on this to show that in MG, coupling forced cell cycle re-entry with deletion of Rbpj, which inhibits the transcriptional effects of Notch signaling, induces some MG to proliferate and take on features of multiple cell types, including MG precursor cells, amacrine-like cells, and bipolar-like cells. This work lends valuable insight into the regenerative potential of mammalian MG, particularly when Notch signaling is manipulated.

      Strengths:

      The major claims of the authors are well-supported. They show convincingly - and through multiple methods including immunostaining, single-nucleus RNA sequencing, and in situ hybridization - that coupling notch inhibition with cell cycle reactivation induces the expression of neuronal markers in mammalian MG. The snRNA-seq data are particularly valuable in demonstrating the induction of bipolar-cell subtypes. Edu labeling is effective in demonstrating the induction of proliferation, and the long-term viability of the generated neuron-like cells is intriguing.

      Weaknesses:

      Whether the newly generated neurons are functionally integrated remains unclear, and the effect of the manipulation on the function of the retina was not tested. Imaging data suggests that many of the newly generated neurons persist for months, but often appear mislocalized. It is also not clear if the manipulation of MG affects long-term MG function. Cell death was not evaluated, and although the authors evaluated the long-term effect on tight junctions, this data was not quantified, and further analysis on morphology or function was not done. Control eyes were untreated, not vehicle-injected.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) The transgenic line may be Glast-CreERT, not Glast-CreERT2.

      We appreciate the reviewer for bringing this to our attention. The formal allele symbol for this transgenic line is Tg(Slc1a3-cre/ERT)1Nat, while this strain is generically classified as "Cre/ERT2" by the Jackson Laboratory. Some published studies referred to this line as Glast-CreERT and others as Glast-CreERT2. To maintain consistency with the formal allele symbol, we have adopted "Glast-CreERT" throughout the revised manuscript.

      (2) For snATAC data in Figure 6 e,f, and Figure 19b. It is most likely gene activity, not gene expression, since these are snATAC, not snRNA data.

      For this inaccurate terminology, we have corrected all relevant figure labels and associated text in the revised manuscript to clearly state "gene activity" instead of "gene expression."

      (3) Some text in Figure 6 is a bit too small to read.

      We have increased the font size of the text elements in Figure 6 to ensure readability and have also reviewed all other figures for consistency. Revised figures with improved legibility have been included in the updated manuscript.

      Reviewer #2 (Recommendations for the authors):

      (1) There are multiple instances where further elaboration of methods or tools in the test would improve readability and comprehension by a broader audience. It would be helpful to (early, often, and clearly) explain precisely which cell types are labeled in your mouse line and how. Someone unfamiliar with the mouse line may struggle to understand what is labeled by the tdT or GFP. Likewise, it would help to consistently define what cell types are labeled by tdT+ vs. Sox9+, tdT+, etc.

      We have added a clear and detailed description of the mouse lines and labeling strategy early in the Results section, specifying which cell types are labeled by tdT and GFP and how the labeling is achieved. We have also ensured that the definitions of cell type identifiers (e.g., tdT<sup>+</sup> for MG-derived cells, Sox9<sup>+</sup>/tdT<sup>+</sup> for MG remaining in a glial state) are consistently stated upon first use and maintained throughout the manuscript to improve readability for a broader audience. In addition, we added headings for the quantification graphs to improve readability in all quantification figures.

      (2) It is unclear what the difference is between Figure 1c and S1c, and these should be quantified as the % of positive cells, as described in the text.

      We have removed Figure S1c and moved Figure 1c to supplementary figure 1. The MG labeled by EdU and Sox9 or Otx2 were quantified as % of the EdU+ MG.

      (3) S2e: Clarify what pixel level means, is this pixel intensity?

      Yes, "pixel level" in Figure S2e refers to pixel intensity. We apologize for the ambiguous wording and replaced "pixel level" with "pixel intensity" in the revised figure legend to ensure clarity.

      (4) Figure 2: In the magnified image of the GFP+, Sox9- cell, the GFP is also very faint. Could these cells be dying? Analysis of the expression profile of these cells (or ruling out apoptosis) would better support a dedifferentiation argument.

      The faint GFP signal observed in GFP<sup>+</sup> Sox9<sup>-</sup> cells is a sign of ongoing dedifferentiation rather than cell death. This is likely due to chromatin remodeling during reprogramming. A similar decrease in reporter signal intensity during MG dedifferentiation has been previously reported by Le et al. 2024, 2025, supporting the interpretation that reduced fluorescence is a characteristic feature of this process. It is possible that a small fraction of GFP<sup>+</sup> Sox9<sup>-</sup> cells may undergo cell death over an extended period, which would be difficult to detect using apoptosis assays. Our long-term survival experiments demonstrate that more than 80% of MG-derived neuron-like cells survive for at least 9 months following treatment (Figure 7), indicating that majority of these cells are viable. The discussion is included in line 112-114 on page 5.

      (5) Figure 3: The Crx labeling appears everywhere except the identified cell. This seems the opposite of the point you are making.

      Crx signal of the MG-derived cell (tdT<sup>+</sup> Crx<sup>+</sup>), which is pointed out by arrowhead, is in a ring-like pattern. This pattern is consistent with the euchromatin region in inverted nucleus of rod. Crx labeling appears in other cells in the image as Crx is highly expressed in native photoreceptors.

      (6) I think it would be nice to address, in the discussion, the apparent disorganization and mislocalization of cells in the long-term images.

      We thank the reviewer for highlighting this critical observation. During retinal development, precise laminar positioning of neurons is guided by a coordinated interplay of cell-intrinsic transcriptional programs and extrinsic cues including cell adhesion molecules, guidance factors, and interactions with neighboring cells. In the adult retina, many of these developmental cues are no longer present or active, which likely contributes to the failure of MG-derived neurons to migrate to their appropriate laminar positions. Interestingly, the vast majority of our divided MG cells remained localized within the outer nuclear layer (ONL). Because the ONL is the physiological location of photoreceptors, this preferential position could serve as an advantageous baseline layout for driving targeted photoreceptor differentiation in future work. To address reviewer’s feedback, we have expanded our discussion section (Line 543-559, page 23-34) to cover the mechanisms underlying this structural disorganization and its downstream implications for functional circuit integration.

      (7) I'm not convinced that ZO1 alone is sufficient to suggest MG function normally or that retinal homeostasis is maintained. I suggest tempering that conclusion in the text.

      For the revision, we have performed additional experiments to address this concern. The optical coherence tomography (OCT) images revealed that retinal layer organization and ONL thickness were comparable among the uninjected eyes, GFP AAV-injected control eyes, and CCA-treated eyes, demonstrating that overall retinal architecture was well-preserved (Fig. S23a–c). Optomotor response testing revealed no significant differences in visual acuity across groups, suggesting that visual function remained intact (Fig. S23d). Furthermore, electroretinography (ERG) demonstrated that scotopic and photopic a- and b-wave amplitudes were unaffected by the treatment, confirming that light responses from photoreceptor and inner retinal neuron were preserved (Fig. S23e–h). Taken together, these findings demonstrate that combined Rbpj KO and CCA treatment achieves neurogenesis without compromising retinal structure and functional visual circuitry.

    1. eLife Assessment

      This study investigates the role of the Z-disc protein Zasp52 in Drosophila flight muscles and provides evidence that an intrinsically disordered region (IDR) helps to stabilize and promote the localization of the protein to the Z-disc. Overall, this represents an important study that provides insights into Z-disc function and maintenance. The data are convincing, supported by strong genetic evidence and behavioral tests, well-controlled experiments, and detailed statistical analyses. Characterization of a new actin-binding motif mutant and FRAP analysis further support the functional importance of the Zasp52 IDR.

    2. Reviewer #1 (Public review):

      [Editors' note: this version has been assessed by the Reviewing Editor without further input from the original reviewers. The authors have addressed the comments raised in the previous round of review.]

      The manuscript by Ho and Schock investigates the role of the Z-disc protein Zasp52 during Drosophila flight muscle development. It was known before, mainly by findings from this group, that Zasp52 is required for normal sarcomere morphogenesis, specifically Z-disc morphogenesis in indirect flight muscles. But the exact molecular mechanism by which Zasp52 contributes, apart from the fact that it is localised there and is somehow involved in multimerization/cross-linking, was not clear. This paper proposes that an intrinsically disordered region (IDR) in Zasp52 is needed for some of its functions, by stabilising Zasp52 localisation at the Z-disc. Specifically, the IDR in Zasp52 is proposed to be required for Z-disc maintenance during the mechanical challenges of flight, while being dispensable for the initial morphogenesis during development. This hypothesis is supported by strong genetic evidence and behavioural tests, deleting Zasp's IDR impairs flight from mid-age onwards, while a block in flight activity lifts the phenotype.

      Strengths:

      (1) The linker in the alternatively spliced exon 15 of Zasp52 was deleted with a state-of-the-art genetic editing strategy. Surprisingly, flies are homozygous viable, showing that this long part of the Zasp52 protein is not essential for animal survival or sarcomere morphogenesis.

      (2) The observed sarcomere phenotypes with age, especially the bending Z-discs, are new and exciting.

      (3) The displayed EM images document interesting phenotypes.

      (4) Most of the observed phenotypes can be rescued by re-expression of the long Zasp52 isoform, which does contain the IDR region, but not by a shorter one without it, suggesting that IDR is important.

      (5) FRAP data measure the local turnover of a short-ZaspGFP and show that this increased in the Zasp mutant lacking the IDR domain, suggesting that Zasp-IDR might stabilise Zasp at the Z-disc.

      (6) Interestingly, flight and sarcomere morphology phenotypes can be rescued by preventing the flies from flying, suggesting that they are mechanically induced.

    3. Reviewer #2 (Public review):

      Summary and Strengths:

      This in-depth genetic analysis of Zasp52 function in Drosophila indirect flight muscle (IFM) provides an interesting perspective regarding the role of a partially disordered region (IDR) in exon 15e. This exon seems to be exclusively present in IFM and contributes to the prevention of myofibril disintegration during aging, likely due to interactions of this region with Z-disc insertion and/or stability. The addition of an isoform (PR) that lacks exon 15e serves as a nice control to illustrate the necessity of exon 15e in muscle structure and function. Overall, the manuscript is exceptionally well-written, logical, with nicely controlled experiments and detailed statistical analysis that largely support the conclusions drawn by the authors. While exon 15e is clearly involved in preventing muscle degeneration, a solid role for thin filament stability is not clearly shown (as mentioned in the abstract). In addition, which regions/how the proteins of the IDR may contribute are unclear.

    4. Author response:

      The following is the authors’ response to the original reviews.

      Thank you for the helpful comments and criticisms. We provide exciting additional data, in particular a CRISPR actin-binding motif mutant and FRAP analysis of an exon15e-GFP transgene, both further supporting the importance of the IDR in thin filament stability. We believe that these additional experiments provide compelling evidence supporting our conclusion and substantially advance the current limited body of knowledge surrounding the role of IDRs in structural proteins.

      Public Reviews:

      Reviewer #1 (Public review):

      The manuscript by Ho and Schock investigates the role of the Z-disc protein Zasp52 during Drosophila flight muscle development. It was known before, mainly by findings from this group, that Zasp52 is required for normal sarcomere morphogenesis, specifically Z-disc morphogenesis in indirect flight muscles. But the exact molecular mechanism by which Zasp52 contributes, apart from the fact that it is localised there and is somehow involved in multimerization/cross-linking, was not clear. This paper proposes that an intrinsically disordered region (IDR) in Zasp52 is needed for some of its functions, by stabilising Zasp52 localisation at the Z-disc. Specifically, the IDR in Zasp52 is proposed to be required for Z-disc maintenance during the mechanical challenges of flight, while being dispensable for the initial morphogenesis during development. This hypothesis is supported by strong genetic evidence and behavioural tests, deleting Zasp's IDR impairs flight from mid-age onwards, while a block in flight activity lifts the phenotype.

      However, some of the phenotypic analysis, in particular the bending of the sarcomere, likely upon mechanical challenge by muscle contractions, needs more detailed investigations to be fully convincing.

      Strengths:

      (1) The linker in the alternatively spliced exon 15 of Zasp52 was deleted with a state-of-the-art genetic editing strategy. Surprisingly, flies are homozygous viable, showing that this long part of the Zasp52 protein is not essential for animal survival or sarcomere morphogenesis.

      (2) The observed sarcomere phenotypes with age, especially the bending Z-discs, are new and exciting.

      (3) The displayed EM images document interesting phenotypes.

      (4) Most of the observed phenotypes can be rescued by re-expression of the long Zasp52 isoform, which does contain the IDR region, but not by a shorter one without it, suggesting that IDR is important.

      (5) FRAP data measure the local turnover of a short-ZaspGFP and show that this increased in the Zasp mutant lacking the IDR domain, suggesting that Zasp-IDR might stabilise Zasp at the Z-disc.

      (6) Interestingly, flight and sarcomere morphology phenotypes can be rescued by preventing the flies from flying, suggesting that they are mechanically induced.

      Weaknesses:

      (1) The western blot quantifications of Zasp isoform expression are weak. No error bars are indicated in the quantifications; the quantifications appear to be more qualitative than quantitative. According to band intensities, the long Zasp isoforms seem to be less present compared to the shorter ones, even in the flight muscles.

      We have now included quantifications with error bars for the Western blots in our resubmission. It is important to keep in mind that the main point in figure 1B is that there are plenty of exon15e-containing isoforms in IFM, in contrast to other tissues with very limited exon15e-containing isoforms. This is confirmed by the analysis of RNA-seq data in figure 1C, and of course, by the flightless phenotype of the exon15e mutant.

      (2) The phenotypic analysis of the sarcomere appears somewhat superficial throughout the paper. Only Zasp52 and phalloidin are shown; no other Z-disc or thick filament proteins. At least myosin stainings and overview images are important to better judge the phenotypic variations. Are the variants between individuals or regional in the same muscle?

      Our images are representative of the observed phenotypes. Phenotypes are consistently present across all individuals, as reflected in our replicates. Interestingly, they appear not to be randomly interspersed among the sarcomeres but concentrated in certain regions of muscle more than others. Full images are available in the online repository FigShare.

      (3) EM images would benefit from better quantification.

      We do not believe that EM images can be meaningfully quantified, because of the many selection steps preceding image acquisition.

      (4) Other proteins were not analysed with the FRAP-based turnover assay for comparison in wild type and mutant. All Z-proteins might turn over faster in the mutant with the defective Z-disc.

      This is the point we are trying to make. The Zasp52 IDR appears to stabilize the Z-disc and is likely involved in fastening a variety of proteins to it.

      Reviewer #2 (Public review):

      Summary and Strengths:

      This in-depth genetic analysis of Zasp52 function in Drosophila indirect flight muscle (IFM) provides an interesting perspective regarding the role of a partially disordered region (IDR) in exon 15e. This exon seems to be exclusively present in IFM and contributes to the prevention of myofibril disintegration during aging, likely due to interactions of this region with Z-disc insertion and/or stability. The addition of an isoform (PR) that lacks exon 15e serves as a nice control to illustrate the necessity of exon 15e in muscle structure and function. Overall, the manuscript is exceptionally well-written, logical, with nicely controlled experiments and detailed statistical analysis that largely support the conclusions drawn by the authors. While exon 15e is clearly involved in preventing muscle degeneration, a solid role for thin filament stability is not clearly shown (as mentioned in the abstract). In addition, which regions/how the proteins of the IDR may contribute are unclear.

      Weaknesses:

      (1) It is not clear in Figure S1A where exon 15e fits within the Zasp52 locus schematic. This is important as a premise of this paper describes this region to be key, and proof from multiple prediction programs would lend more weight to the prediction of the exon being largely disordered. Inclusion of the discussed short linear motifs, comparison with Canoe or LBD3 for similarities and/or an Alphafold structure would help make the authors' point (colorized with known domains).

      We added a bar below figure S2A to show the region corresponding to exon 15e. We used three disorder prediction programs and one structure (order) prediction program. The majority of exon15e is completely disordered and of very low confidence score, and thus uninformative to display as an AlphaFold structure. Likewise, IDR’s are very difficult to classify, therefore we cannot say much more than that LDB3, Zasp52, and Canoe contain IDRs, with Zasp52 and Canoe both having a putative actin-binding domain within the IDR. We now provide data on the function of the ABD in this resubmission.

      (2) Interesting that immobilization rescues the deterioration phenotypes. The authors should explain in more detail how this was done to avoid dehydration/starvation of the flies.

      We provided more details in materials and methods.

      (3) There is a lot of discussion about the potential function of the IDR region, specifically a putative actin binding motif or other 'ordered' regions that may contain short linear motifs. It would strengthen the findings to show which of these may be essential for Zasp52 function in the IFM. The ability to bind actin could be tested biochemically, and/or smaller deletions could be made to unequivocally test the role of the ABD vs other predicted motifs using genetics. If some of these regions are more ordered, where do they lie within, and do they form a predicted fold or structure that gives insight into function?

      We now provide data on the function of the ABD showing that deleting it has almost no phenotypic defects. That means the IDR is largely/entirely responsible for the observed phenotypes.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) Western blot in Figure 1B needs proper quantification. A ratio between long and short isoforms in the same muscle type might be informative. Is it known which epitope the antibody recognises? Can a GFP insertion that also labels all isoforms be used as verification? Quantifications are also needed in Figure 2A.

      We have added quantifications of all Western blots (Fig. 1B, 2A, and 2A’). The ratio between exon 15e-containing and total Zasp52 in the same muscle type is included in the lowest bar graph in Fig. 1B. The full-length antibody is polyclonal and was raised against Zasp52-PR which contains all ordered domains; the anti-LIM antibody was raised against the last three LIM domains (both are described or referenced in the materials and methods section). Such a GFP insertion cannot exist due to the complex splicing patterns of Zasp52.

      (2) The name of the deletion allele could be specifically indicated in Figure 1A below the red bar.

      Done.

      (3) It would be useful to indicate the order group names in Figure S2B since species names are hard to read.

      For the version of record we provided high-resolution images, where species names can be read. Drosophilids, Ephemeroptera and Odonata are indicated.

      (4) The inverted spelling of the numbers for the control in Figures 2C and 5H is strange.

      Changed to normal spelling.

      (5) The bending of the myofibril at the Z-disc is a really interesting phenotype. However, it seems it is not always visible; at least it is visible in many myofibrils shown in Figure 3B, but in none in Figure 3E, same genotype, just different staining. Hence, I wonder if this bending could be force-induced by the cutting of the thorax during tissue preparation. It would be useful to display some overview images to allow the reader to judge the quality of the tissue preparation, indicating from where the high magnification view shown was taken. The same is true for Figure 5.

      Overview images are available on FigShare. Note that you can see some “H-zone actin” sarcomeres in Fig. 3B, as well as some mildly bent ones in Fig. 3E. We generally selected images that best demonstrated the phenotype described. Furthermore, neither phenotype is fully penetrant so we cannot expect to see it everywhere. Lastly, it is always possible that phenotypes are affected by preparation, since it is impossible to know what the myofibrils look like in situ. However, all samples were prepared using the same protocol with replicates, and since we see a phenotype in our mutants and not in the control, this indicates that something is different between the two.

      (6) The same applies to the visualisation of the "hyper-contracted" phenotype; again, it seems to be an all-or-nothing phenotype in the zoom shown. An overview image should be shown. The zoom in Figure 4E would benefit from displaying phalloidin in a separate channel. Are actin filaments pulled out of the Z-disc? The latter is often seen in non-perfect cuts in wild-type, but the accumulation at the M is curious. It would be informative to locate the ends of the thick filaments in these cases or quantify thick filament lengths; do these invade the Z-discs? This can easily be done by a myosin staining.

      Is this a regional effect or does it depend on the individual or on the preparation? I am surprised to also see the "hyper-contraction" in 10% of wild-type 5-day adults.

      See previous response where we include overview images. Single-channel images are available; it is visible that actin filaments are not pulled out of the Z-disc. Phenotypes are consistent across individuals as evidenced in our replicates but do tend to be concentrated in certain regions of muscle.

      (7) The EM images would benefit from more overview images. At the moment, we only see a single sarcomere from wild type and mutant, with no quantification of the phenotype. Can the authors see the invading thin filaments into the M-band? The disrupted Z-disc phenotypes are impressive. What is the age of the animal shown in Figure 4?

      We have a panel displaying several mutant sarcomeres. Due to the selectivity and challenges of the EM preparation process, we do not believe we can perform meaningful statistics on them. It sometimes looks like myosin heads are visible in the H zone which may support the presence of thin filaments in the H zone (Fig. 4B and C). However, the quality of these particular EM images is not high enough to identify thin filaments. All phenotypes shown are from 3-week-old animals.

      (8) Is UH-3 GAL4 expressed at the adult stage?

      Yes, from 36 h APF into adulthood (Singh et al. 2014). Now mentioned in the results section.

      (9) Figure 6 would strongly benefit from a myosin staining. Do thin and thick filament lengths scale? It seems that overlap is reduced in the double hets. How can this be envisioned with Z-disc stability? Is myofibril diameter reduced?

      We searched for non-additive differences in myofibril diameter but were unable to detect any.

      (10) What is the FRAP turnover rate of a long Zasp-GFP compared to a short one in wild type? A difference would indicate that it is really the IDR domain that keeps Zasp52 longer at the Z-disc, instead of an indirect effect caused by Z-disc morphology

      We have newly added FRAP data of a GFP-tagged exon 15e construct which displays much lower turnover. This indicates that the IDR does indeed retain Zasp52 at the Z-disc.

      Reviewer #2 (Recommendations for the authors):

      (1) The total protein stain should also be included if it is used for quantitation in Figures 1B and 2A-A'.

      These are available on FigShare.

      (2) It is a bit confusing that the Alphafold plot is inversely correlated with the other 3 prediction programs, although this is explained in the legend. Maybe an Alphafold structure would help make the authors' point (colorized with known domains).

      The AlphaFold structure is almost entirely low-confidence disordered region except for the structured domains so we do not believe it would be helpful to include.

      (3) The title of Figure 8 says 'Certain ex15e defects are rescued by immobilization.' What other defects are not rescued? If true, these should be shown.

      There was a full rescue. We deleted the word “certain”

      (4) Please include a brief explanation of the spatiotemporal expression of UH3-Gal4.

      From 36h APF into adulthood (Singh et al. 2014). Now mentioned in the results section.

      (5) Statistics should be added to Figure 8E.

      Figure 8E (now 9E) has statistics.

      (6) The dark blue color used for integrin staining in Figure S3 is difficult to see. Changing this color may help visualize differences. Also, pointing them out with arrows, etc., will help clarify abnormalities.

      We have described these differences in the figure caption. Single-channel images are available for viewing in any color in FigShare.

    1. eLife Assessment

      This study reports important findings regarding social influence on charitable donations, showing that giving is shaped by the statistical properties of others' donations in a manner that can be captured by a reinforcement learning model. The evidence for the conclusions is solid, although the computational modelling could be better motivated and described, the individual differences analyses could be more robust, and some design choices could be better motivated. Overall, the core effect appears robust and is supported by multiple well-designed experiments, but the conclusions that rely upon computational modelling and individual differences may require further support.

    2. Reviewer #1 (Public review):

      This manuscript investigates how people use sequential social information when deciding how much to donate to charity. Across four preregistered experiments, participants first made baseline donations to a set of charities, then observed a sequence of donations from five other people whose mean and variability were experimentally manipulated, and finally made a second donation to the same charities. The authors ask whether the mean and variability of others' donations affect the mean and variability of participants' own donations, and whether individual differences in psychopathy and empathy are associated with responsiveness to social information.

      The main behavioral finding is that participants shifted their second donations toward the mean of the donations they observed: generous social information increased donations, whereas stingy social information decreased donations. In contrast, the variability of observed donations had little effect on the mean donation shift, but did affect the variability of participants' subsequent donations, with more consistent social information producing stronger reductions in variability. The authors also fit several computational models and conclude that a hybrid model, in which second donations reflect both participants' initial donations and learned predictions of others' donations, best accounts for the data. Finally, they report that psychopathic traits are positively associated with donation change and with model-derived social-information use, and that this association generalizes to a perceptual social-influence task in Experiment 4.

      The paper addresses an interesting question and has several strengths, especially the repeated experimental design, the direct manipulation of social-information statistics, and the attempt to connect descriptive behavior with computational modeling and individual-difference measures. However, several aspects of the design and analysis currently block some of the major conclusions. The behavioral results provide convincing evidence that observed donation levels affect later donation decisions. The current evidence is less decisive for the stronger claims that the winning computational model identifies the underlying mechanism, that individual-level model parameters are robust phenotypes, and that psychopathy specifically increases susceptibility to social information.

      Strengths:

      A major strength of the manuscript is that it investigates social influence in charitable giving across four preregistered experiments with relatively large samples. The core mean-effect result is replicated across different donation scales, across hypothetical and incentivized settings, and across student and more general online samples. This gives the descriptive behavioral finding substantially more credibility than would be available from a single experiment.

      The experimental manipulation is also valuable. Rather than presenting only a single prior donation or a simple group average, the authors expose participants to sequences of donations and independently manipulate the mean and variability of this social information. This design allows the authors to ask not only whether social information changes donation levels, but also whether the distributional structure of that information changes the variability of participants' own responses.

      Another strength is the combination of traditional statistical analyses with computational modeling. The hybrid model is a reasonable descriptive candidate because it formalizes the intuitive idea that second donations may depend both on participants' initial preferences and on learned expectations about others' donations. This modeling approach has the potential to clarify mechanisms of social-information use, especially if the validation of the model and its individual-level parameters is strengthened.

      Experiment 4 is a sensible extension because it uses an incentivized design, includes a more diverse sample, examines transfer to novel charities, and adds a perceptual social-influence task. These features broaden the empirical scope of the manuscript and make the psychopathy-related findings more interesting, although the perceptual-task result should still be treated as requiring replication.

      Weaknesses

      The first limitation concerns causal interpretation of the phase effects. Participants always make baseline donations first, then observe social information, and then make second donations to the same charities. There is no non-social repeated-donation control condition. This type of design does support the conclusion that donation changes differ as a function of the observed social-information condition, especially the mean of others' donations. However, it does not by itself fully isolate social influence from other processes that could also occur between a first and second donation to the same item, such as repeated exposure to the charities, slider familiarity, memory of the first donation, regression to the mean, reduced uncertainty, fatigue, or "the experiment clearly wants me to update" demand effects. This issue is especially relevant for the claim that observing others' donations generally reduces the variability of individual donations. The variability effect may well be socially driven, but the absence of a non-social or irrelevant-information repeated-donation control means that this cannot be decisively demonstrated.

      The second limitation concerns the trial-level mixed models. The primary mixed-effects models include random intercepts for participants and items, but do not appear to include random slopes for within-participant or within-item phase effects. Since phase is repeatedly manipulated within participants and items, random-intercept-only models may underestimate uncertainty for some phase interactions, resulting in anti-conservative p-values. The convergent participant-level ANOVA analyses are reassuring, but the trial-level inferential claims would be stronger if the authors reported additional analyses using fuller random-effects structures or other methods that better reflect the repeated-measures structure.

      The third limitation concerns model comparison and model validation. The computational models are fit separately to each participant, and model comparison is based on summed information criteria and protected exceedance probabilities derived from those participant-level fits. This is informative about relative conditional fit within the tested sample and model set. However, the manuscript uses the winning model to support broader claims about latent computational mechanisms, individual computational phenotypes, psychopathy-related susceptibility, and potential intervention relevance. For these claims, the relevant prediction target is generalization to new participants, whose individual parameters are not known in advance. The current model-comparison approach is not well aligned with that target. Additionally, the loss appears to combine prediction trials and donation outcomes, so the selected model may more strongly reflect performance at predicting participants' guesses about others rather than specifically predicting their own donation decisions.

      The fourth limitation concerns the model adequacy checks and recovery analyses. The analyses described as posterior predictive checks do not appear to be posterior predictive checks, because the models are not Bayesian and there consequently isn't a posterior to check. Instead, the analyses appear closer to some sort of in-sample fitted-value reconstruction checks. Such checks provide limited evidence of model adequacy, especially because the same second-donation data used to estimate individual parameters are then used to assess whether the fitted model reproduces the main behavioral patterns. In addition, the reported model and parameter recovery analyses use extremely favorable response-noise assumptions that are not expected to be met in real data. The analyses establish that the models and parameters are mathematically distinguishable in principle, but they do not establish that the individual-level parameters are reliably recoverable under realistic empirical noise levels to the extent required for the analyses performed in the manuscript.

      The fifth limitation concerns the interpretation of the psychopathy results. The association between psychopathic traits and donation change is interesting and appears directionally consistent across experiments. However, the interpretation that psychopathy increases susceptibility to social information is vulnerable to biasing by baseline-distance. The manuscript reports that psychopathy is negatively associated with baseline donations in Experiments 1-3. Participants with lower baseline donations have more room to move toward generous social information, and absolute donation change is partly a function of the distance between the initial donation and the observed social mean for mechanical reasons. Thus, an association between psychopathy and absolute donation change could theoretically arise even if psychopathy does not directly increase social susceptibility.

      A sixth limitation is that we could not find the links to the preregistration. The authors state when preregistered hypotheses were or were not supported, but it is unclear how these hypotheses were phrased. Most notably, it is unclear how variance in the observed donation choices was supposed to influence participants. As a side note, it was not quite clear if the variance in the observations was higher or lower across charities, across observed persons, or across both.

      Several more minor suggestions can also be made regarding the modelling and the presentation of the task, etc.

    3. Reviewer #2 (Public review):

      Summary:

      This manuscript examines how the statistical properties of others' charitable donations shape subsequent giving using four preregistered experiments and computational modelling. The authors find that both the average level and variability of observed donations influence donation behaviour, and that individual differences in social information use are associated with psychopathic traits.

      Strengths:

      This is a well-executed paper on the important question of how social information shapes charitable giving. In my view, the combination of preregistered experiments, large sample sizes, computational modelling, and a multi-paradigm approach makes for convincing evidence. The progression across experiments, the use of real donation data rather than deception, the incentivized experiment 4, and the generalization to a second paradigm are all notable strengths. The introduction is clearly written and well-motivated - an enjoyable read. The experimental paradigm is thoughtfully designed, and the methods and supplementary materials are described in considerable detail. The computational modelling provides useful additional insights beyond the behavioural analyses.

      As far as I could tell, the manuscript also adheres closely to the preregistrations. The primary hypotheses, experimental designs, exclusion criteria, and key analyses are all consistent with the preregistered plans. Deviations seem to consist of methodological improvements (e.g., mixed-effects models replacing ANOVAs), additional computational and robustness analyses, and therefore strengthen rather than weaken the manuscript. (NB: for transparency, I would appreciate a clearer distinction between preregistered and post hoc analyses, as well as a brief explanation for why some preregistered secondary analyses are no longer reported; see minor comments below).

      Overall, I enjoyed reading this paper. I believe it will make a valuable contribution. My comments below are intended to further strengthen an already solid manuscript.

      Weaknesses:

      (1) The rationale for the social-information phase could be clarified further. Given the research question, I wondered why participants observed the five donations sequentially (and only briefly) rather than simultaneously. In particular, variance is arguably more difficult than the mean to encode and remember, and a sequential presentation may both obscure distributional differences and introduce primacy or recency effects. It would be helpful if the authors could better motivate this design choice, and indicate whether they examined possible order effects.

      Relatedly, I felt somewhat uncertain about the purpose of asking participants to predict each donation before observing it. The prediction phase appears to play an important role in the computational model, but its theoretical role is not clearly introduced. Is it intended as a measure of participants' evolving beliefs about the descriptive norm, or primarily as a modelling device? Finally, were these predictions incentivized (e.g., for accuracy), and if not, how should readers interpret them?

      (2) I would appreciate having the full experimental materials reproduced in the Supplementary Information. This would make it easier to understand what participants experienced during the task, including what they were told about the "other participants" whose donations they observed.

      Minor points:

      (1) The interpretations around domain-generality would be strengthened by reporting the association between social information use in the charitable giving task and in the BEAST. Currently, both measures are shown to correlate with psychopathy, but it remains unclear whether individuals who rely strongly on social information in one task also do so in the other. Reporting this correlation (or explaining why it cannot be meaningfully computed) would provide a nice and direct test of a domain-general tendency to use social information.

      (2) It would help to explain more explicitly why the standard deviation of donations is theoretically interesting in its own right. The motivation for studying the mean seems immediately intuitive, whereas the motivation for focusing on variability could be elaborated on further in the Introduction.

      (3) As I said above, I think the manuscript follows the preregistrations closely. Maybe I missed it, but it seems that prediction accuracy and reaction-time analyses were omitted. It would improve transparency further if the authors would briefly mention the preregistered secondary analyses that are no longer reported (and explain why they were omitted).

    4. Reviewer #3 (Public review):

      Summary:

      In this manuscript, the authors aimed to assess the mechanisms of social influence on charitable giving, particularly by separating the role of donation magnitude and variability in others' donations, and by examining the role of incremental social information in a learning framework. They additionally investigated individual differences in the magnitude effects in relation to self-reported psychopathy and empathy. The main findings suggest that magnitude and variability of others' donation impacted the magnitude and variability of the participants' donations, respectively, and that the weight of social information on individual decisions correlates positively with psychopathy, but not with empathy.

      Strengths:

      (1) The findings extend previous evidence for social influence on charitable giving to contexts where social information is provided incrementally, and to effects on the variability in social information (in addition to the mean).

      (2) Individual differences suggest a role for psychopathy, but not empathy.

      (3) Findings are replicated across all 4 (or for some findings 3 out of the 4) experiments, which helps strengthen the claims.

      (4) Multiple experiments are a strength, especially Experiment 4, which helped address concerns/potential confounds in the previous experiments, increase representativeness of the sample, add incentive compatibility, and generalize to another task domain (perceptual).

      (5) For modelling, strong model and parameter recovery was obtained, thus validating the modelling pipelines.

      (6) The experiments were pre-registered, though it's unclear whether only planned analyses were pre-registered, or specific directional hypotheses. It would help if the manuscript took the reader through the pre-registration (and any deviation from it), instead of expecting the reader to do the comparison between the pre-registrations and actual manuscripts.

      (7) The studies are appropriately powered, and power analyses are provided.

      Weaknesses

      (1) Lack of rationale and justification for the between-subjects design.

      While this design may be appropriate in some cases (for example, for the generalization of donation to new charities or as a potential "intervention"), it would have been great to know if the findings related to social influence extend to a within-subjects design, especially given the weak results related to the effects of standard deviation in others' donations. It is possible that variability in others' responses would have a stronger effect if manipulated within individuals, since the same individual exposed to both high-SD and low-SD social information may weight low-SD information more, but this effect may lack when individuals are only exposed to the same variability across trials.

      (2) Motivation for the RL framework.

      The use of reinforcement learning (RL) isn't very well motivated, both in the introduction and methods/results (given the task). In particular, why is RL relevant to studying the problem of social influence, which isn't inherently a learning problem? This should be better motivated in the introduction. Second, when taking the task into account, it's unclear why RL is an appropriate model, given that from the perspective of the participant, the 5 others are different individuals, so the model shouldn't assume that predicting an individual's donation should be related to the previous individual's donation. Unless participants are informed that there is some dependency between the 5 donors they observe on each trial? If so, this should be made clear.

      (3) Specifics of modelling analyses, and separability between prediction and second donation data.

      Does the RL-based model (either prediction-only or hybrid) explain more variance in second donations than a simple linear regression model predicting second donation from initial donation and the mean of others' donations (or each individual other's donation)? It could be helpful to add some models that include social influence (i.e., integration of social and individual information) but no learning mechanisms per se. If this is not done, I do not believe that current results show that participants combine "their initial self-donation tendencies with their predictions of observed others' giving to guide their second individual donations". While participants may update their predictions, the authors should test multiple models of prediction update (fit only on the prediction data to understand the specific mechanisms of prediction update independently of second donation - for example, is it RL, or could it just be a running average, or some other heuristic? In parallel, it would be helpful to test whether it's the learned predictions (or whatever other prediction update mechanism was found to best explain the prediction data) or the actual others' donation information that best explains second donation - when combined with initial donation. These latter models would be fit on second donation data only in order to be comparable. If it's not possible to separate people's predictions from the actual social information (others' donations) then this should be acknowledged as a limitation. Ultimately, separating the modelling by data type (prediction only vs second donation data only) would help provide more insights into the learning mechanisms (if any) and whether it's learned prediction, or just social information, which influences second donation.

      (4) Missing statistics in generalization to novel donation results.

      On page 13, in the generalization effect, the authors mention that "Compared with participants exposed to High-SD social information, those exposed to Low-SD social information exhibited less variability in their novel donations, with this effect being especially pronounced in the Low-Mean condition." Was this supported by a significant interaction between SD and Mean condition? If so, please report the statistics of the interaction; if not, it's probably better to refrain from making this claim.

      (5) Behavioral index of social influence individual differences.

      For the first analysis reported on the association with psychopathy (Figure S9), as well as empathy (Figure S10), the absolute change between first and second donation does not seem like the appropriate marker of social influence. While I understand from Figure 2 that most participants changed their donation in a direction consistent with the social information, it would appear more appropriate to calculate an index of donation change consistent with influence, so calculated as D2 - D1 for the high mean groups and D1 - D2 for the low mean groups. This would be a better measure to interpret high values as an index of social influence.

      (6) Interpretation of psychopathy effects.

      a) The general idea that high psychopathy would be associated with increased social influence seems counterintuitive. While I appreciate that the authors controlled for additional variables such as age, gender, condition, and other model parameters, is it possible that this effect could be instead explained by the availability heuristic (the social information is more readily available to participants than their individual choice from the baseline trials), lower memory for their own choice, or lower IQ/cognitive abilities? These appear to be important confounds to address to be able to interpret the findings.

      b) Related to this, and given that psychopathy/empathy were negatively/positively related to baseline donation amounts, it would be good to account for baseline mean donation amount in the individual difference analyses.

      c) Finally, the authors interpret this association in line with other studies that have shown strategic social blending in psychopathy - while this seems possible in contexts where others are present, it doesn't really seem to be the case in this task. Did participants believe the other donors were watching them somehow? It also appears contradictory for the incentivized experiment, whereby if high psychopathy participants would no longer be able to "maintain a favorable social image while still pursuing their own self-interests" (p.23), since as soon as incentivization is added, participants' own self-interests are directly in conflict with the social image. Was participants' understanding of the incentive compatibility tested in Experiment 4?

      (7) Asymmetry between generous vs stingy social influence and link with psychopathy.

      a) Was such an asymmetry present - in other words, were people more strongly influenced by generous others or stingy others, or were the two comparable? I believe some analyses could be added to test this, and this is also where a within-subject design could help (e.g., different parameters for the two directions of social influence at the individual levels).

      b) Related to that, does the correlation with psychopathy vary between conditions? It appears important to test if the increased social susceptibility is general or specific to increases (~high mean group, generous social influence) or decreases (~low mean group, stingy social influence) in donation. I understand that the main effect of psychopathy survived controlling for conditions, but it would still be interesting to test for an interaction between psychopathy and condition in predicting donation changes (calculated as suggested in point 5 above) or social influence weight.

      (8) Perceptual task in Experiment 4.

      a) While it is good to show that there was no correlation between psychopathy and initial estimate in the perceptual task, were there differences in initial estimate accuracy (i.e., difference between initial estimate and correct answer) along psychopathology? If so, this should be controlled for in the analyses. Given that social influence is always in the direction of the true value, the proportional deviations between initial estimate and social information could yield larger numerical differences and induce larger changes in estimate.

      b) Even if previous studies have excluded rounds in which participants update their estimate in the opposite direction of the social information or move beyond it, I believe analyses that include those rounds should be included, especially in the context of individual difference analyses. Could it be that individuals who are high in psychopathy or low in empathy have a higher proportion of rounds where they go against the social influence? The same question applies to the main 4 experiments (in case this criterion was applied to) as well as the perceptual task.

      c) Because the perceptual task was completed by the same participants as Experiment 4, were the two social influence measures correlated across tasks? Was psychopathy better predicted by a combination of predictors across the two tasks?

      (9) Were individual difference measures examined in relation to the variability effect?

      (10) Discussion.

      The authors argue against a role for opportunistic conformity. While I tend to agree with their interpretation, I believe that it could be strengthened as follows:

      a) First, it relies on a null result (the absence of a difference in decreases between low-mean low-SD and low-mean high-SD groups), which I do not believe was explicitly tested; and even if it was, it should ideally be corroborated by Bayesian statistics to provide strength of evidence for the null effect.

      b) Second, this could be a great opportunity to dive into the mechanisms of social influence in the model, by testing the theory that only the lowest (or highest) donation from the group (rather than the mean, or the learned prediction) influences donation. Could a subset of participants be better fitted by such a model?

      (11) Methods. Maybe I missed it, but it's unclear what participants were told about the other donors they are observing. It is mentioned that they were fully debriefed after the experiment, but what they were told in the instructions appears important. Was believability tested (this also relates to my comment #1 about the rationale for a between-subjects design, which creates fairly biased sets of social information from the perspective of a single participant)? And related to my comment #2, what participants were told about the donors could help justify the rationale for the RL framework.

    5. Author response:

      We thank the editors and reviewers for their thoughtful and constructive comments on our manuscript. We are pleased that they considered the core behavioral findings important and robust, especially the results showing that the magnitude and variability of others’ donations affected the magnitude and variability of participants' donations, respectively. We also appreciate their acknowledgement of the strengths of the experimental design, large sample sizes, the incentive-compatible and across-domain measures included in Experiment 4, and combined behavioral and computational approaches.

      We agree that the manuscript would benefit from greater clarification in several areas, further analyses, and more cautious interpretations. In the revised manuscript, we plan to clarify the rationale for sequentially presenting social information, the role of prediction responses, the theoretical motivation of examining the variability of others’ donation, the use of the between-subjects design, and the motivation for the RL framework. We also agree that the lack of a non-social repeated-donation control condition limits the interpretation of the phase effects. Our design permits strong inferences about differences in donation changes across different conditions, but it cannot establish that the phase-related changes are exclusively attributable to social information exposure. We will revise the wording accordingly, moderate the causal language, and explicitly discuss the limitations of our design.

      To strengthen the behavioral analyses, we plan to supplement the current mixed-effects linear models with models that reflect the repeated-measure structure of the task, including random slopes for the phase. We will also add statistics in the generalization results section and test the asymmetry between generous vs. stingy social influence. Moreover, the reviewers raised an important concern regarding the associations between psychopathy and the donation change. Because psychopathy is negatively correlated with initial donations in several experiments, absolute donation changes may partly reflect the distance between initial donation and the observed donation mean. We therefore plan to reanalyze the psychopathy effects by using signed donation changes and trial-level discrepancies between participants’ initial donations and the observed social information. These additional analyses will enable a more direct and precise assessment of whether psychopathy is associated with greater susceptibility to social influence.

      We further agree that the comparison and validation of the computational models should be strengthened. In the revised manuscript, we plan to clarify that the learning models are intended to describe the updating beliefs about a group-level donation norm from sequential social information, rather than learning about a single donor. We will expand the candidate model set to include non-learning models, such as models based on the actual social mean, a running average. We will also model the prediction phase and the second donation phase separately. This will help identify the models that provide explanatory values for both predictions of others’ donations and individual donation behaviors. In addition, because the models were not estimated via a Bayesian framework, we agree that the term “posterior predictive checks” is inappropriate. We will rename these analyses. We will also rerun the parameter and model recovery analyses using empirically informed noise levels separately for the prediction and donation phases. In addition, we will implement model-evaluation processes, such as cross-validation, that better reflect prediction for new participants.

      In Experiment 4, we plan to directly report the association between social-information-use measures in the perceptual and the donation task to strengthen the domain-generality effect. We will additionally examine whether social susceptibility in the perceptual task is associated with psychopathy by including all trials, including those in which participants moved away from or beyond the social value.

      Finally, we will correct the reporting and presentation issues identified by the reviewers, including the social information use equation in the perceptual task, the pseudo-SD of individual donations formula, supplementary figure captions, and task duration. We will also provide fuller experimental materials and make the preregistration links more prominent. In addition, we intend to make the analysis code, model-fitting scripts, and data available during the revision process.

      We greatly appreciate the editors’ and reviewers’ thoughtful suggestions, which will help us substantially strengthen the manuscript. We are grateful for the opportunity to address these important points and believe that the planned revisions will enhance the manuscript’s clarity, robustness, and its contribution to the understanding of social influence in donation behaviors.

    1. eLife Assessment

      In this important study, Lau et al. identify non-conserved nucleotides within the common binding motifs of BLIMP1 and IRF4 that provide a molecular mechanism for their distinct roles as crucial transcription factors during the antibody-secreting cell differentiation. The major strength of this manuscript is the solid and detailed characterization of human in vitro plasma cell differentiation. However, several overstatements exist, therefore requiring careful revision to improve the manuscript.

    2. Reviewer #1 (Public review):

      Summary:

      This manuscript investigates how IRF4 and BLIMP1 coordinate human plasma cell differentiation. Using a stepwise in vitro culture system starting from primary human naïve B cells, the authors define a developmental window enriched for plasma cell precursors and use stage-specific CRISPR/Cas9 perturbation to examine the roles of IRF4 and PRDM1/BLIMP1 during the transition from plasmablast-like precursors to plasma cells. Single-cell transcriptomic analyses suggest that IRF4 acts early to license plasma cell differentiation, whereas BLIMP1 contributes more prominently to consolidation of the terminal plasma cell program. The authors further combine multiome profiling, CUT&RUN, motif modeling, and EMSA assays to propose the sublet nucleotide variation within ISRE/EICE-like motifs contributes to differential or shared binding by IRF4 and BLIMP1.

      Overall, this is a carefully performed and conceptually interesting study. It provides a useful experimental platform for dissecting human plasma cell differentiation and offers a mechanistic model for how two closely connected transcription factors can exert distinct and coordinated genomic functions during terminal B cell differentiation.

      Strengths:

      A major strength of the study is the establishment and detailed characterization of a human in vitro plasma cell differentiation system. The authors combine phenotypic, functional, and single-cell transcriptomic analyses to define the transition from activated B cells to plasmablst/plasma cell precursor-like cells and then to more mature plasma cells. This system is very useful for future perturbation studies of human plasma cell differentiation.

      A second strength is the stage-specific perturbation strategy. By targeting IRF4 or PRDM1 at the precursor-enriched stage, the authors avoid some of the interpretive limitations associated with earlier perturbations that would affect B cell activation, proliferation, and plasma cell commitment simultaneously. The distinct phenotypes observed after IRF4 versus PRDM1 perturbation provide support for a model in which these two factors act in a temporally ordered manner.

      A third strength is the integration of multiple genomic and biochemical approaches. The combination of single-cell RNA-seq, chromatin accessibility profiling, CUT&RUN, computational motif analysis, and EMSA assays provides a rich dataset and supports the idea that ISRE/EICE sequence variation contributes to differential IRF4 and BLIMP1 occupancy.

      Weaknesses:

      While the multi-omic approach and computational modeling are highly impressive, several major assumptions regarding the cellular differentiation model and genomic linkages require more rigorous validation.

      First, because CRISPR editing was performed on heterogeneous bulk Day 7 cells rather than purified precursor populations, it remains ambiguous whether the observed developmental blocks are truly specific to the prePC window.

      Second, given that IRF4 and BLIMP1 operate within a mutually reinforcing positive feedback loop, the phenotypic divergence between IRF4 KO and PRDM1 KO may reflect differences in protein degradation kinetics or hierarchical dominance rather than a strictly ordered "sequential function".

      Lastly, the motif-lexicon model is elegant and supported by biochemical DNA-binding assays, but the link between motif variation and gene regulation in cells remains partly correlative. Direct testing of selected regulatory elements would make the causal claim stronger. Alternatively, the authors should temper the language and present the motif lexicon as a predictive model for differential occupancy rather than as a fully demonstrated mechanism of gene regulation.

    3. Reviewer #2 (Public review):

      Summary:

      The manuscript by Lau et al. investigates the mechanisms underlying IRF4 and BLIMP1 transcriptional activities during antibody-secreting cell fate decision. Both master regulators of plasma cell differentiation, these two transcription factors have distinct targets and non-overlapping roles. The authors used an in vitro culture system to generate antibody-secreting cells from human naïve B cells, and scRNA-seq, Crispr Cas9 editing, and Cut&Run to dissect the molecular mechanisms defining their specificity.

      Strengths:

      The experiments are overall well executed, and the manuscript is well written. The in vitro culture model appears to generate genuine human antibody-secreting cells. The identification of non-conserved nucleotides within the binding motifs that induce the specific binding of IRF4 or BLIMP1 is convincing, novel, and exciting.

      Weaknesses:

      The authors need to correct some overstatements and flaws to improve the manuscript.

      In Figure 1f, the authors aimed to determine whether in their culture system the plasma cells emerged from the plasmablasts or directly from the activated B cells. First, it is noticeable that the distinction between plasmablasts and plasma cells relies here only on the expression of CD138. It does not include a higher capacity to secrete antibody or their proliferative state. In Figure 1e, the authors could have strengthened their distinction by showing the Ki67 staining at day 21 for both subpopulations. Second, this question does not seem to be related to IRF4 or Blimp1 activity, and thus one could wonder if it is relevant to this study. Finally, and most importantly, the design of the experiment appears flawed to me. The authors sorted cells at day 7 of culture based on their expression of CD20 and put the two subpopulations back for 14 more days. This culture system is a stepwise system, and it is not specified if the CD20+ cells were put back in the day 7 condition or the day 0 condition with the CD40L stimulation. Have both conditions been tested? This experiment also assumes that all B cells have equal potential to differentiate into antibody-secreting cells. What if it is not the case and some are anergic or have committed to the memory B cell fate during the first 7 days? Then the day 7 CD20+ fraction would be enriched in these cells. Moreover, this experiment didn't show that the plasma cell derived from the plasmablasts in the strict sense of the term, as the CD138+CD20- cells could be a mix of proliferative plasmablasts and immature plasma cells.

      In Figure 3a and thereafter, the authors claimed that IRF4 acted earlier than BLIMP1, but both deletions strongly affected differentiation at day 7. IRF4 might have a stronger effect, but it does not mean that it had an earlier effect. To substantiate their claim, the authors would need to demonstrate that, at an earlier time point, deletion of IRF4, but not BLIMP1, results in defective differentiation.

      In Figure 3b, the authors stated that in each individual KO the expression of the other transcription factor was lower. Given that there were no cells in the gate, it is puzzling to figure out how these expressions were compared.

      In Figure 3c, on the UMAP the bottom right part of the activated B cell cluster does not appear to be attributed to any condition. How can it be? Besides, it is highly surprising that at D9 we cannot see any plasmablast on these UMAP, even in the control. Based on the G1/S and G2/M scores, none of the ASC represented were proliferating. Could the authors explain this strong discrepancy with Figure 1?

      Another discrepancy exists between Figure 3b and c: Figure 3b depicted no IRF4- or BLIMP1-expressing cells in either KO, so what were the stunted PC and the BLIMP-KO PC reported in Figure 3c? What are the signature genes defining pre-PC and the score depicted in Supplementary Figure 3d, as the materials and methods only state that they are intermediate between PC and B cells? Could the authors show IRF4, BLIMP1 and some of their known target expression in these populations?

      The authors claim that BLIMP1 is not needed to initiate the transition from pre-PC to PC, but in Figure 1, the intracellular staining showed that at day 7 the antibody secreting cells already expressed BLIMP1. This would rather suggest that BLIMP1, unlike IRF4, does not need to be maintained once the cell reaches a certain point.

    4. Author response:

      Reviewer #1 (Public review):

      Summary:

      This manuscript investigates how IRF4 and BLIMP1 coordinate human plasma cell differentiation. Using a stepwise in vitro culture system starting from primary human naïve B cells, the authors define a developmental window enriched for plasma cell precursors and use stage-specific CRISPR/Cas9 perturbation to examine the roles of IRF4 and PRDM1/BLIMP1 during the transition from plasmablast-like precursors to plasma cells. Single-cell transcriptomic analyses suggest that IRF4 acts early to license plasma cell differentiation, whereas BLIMP1 contributes more prominently to consolidation of the terminal plasma cell program. The authors further combine multiome profiling, CUT&RUN, motif modeling, and EMSA assays to propose the sublet nucleotide variation within ISRE/EICE-like motifs contributes to differential or shared binding by IRF4 and BLIMP1.

      Overall, this is a carefully performed and conceptually interesting study. It provides a useful experimental platform for dissecting human plasma cell differentiation and offers a mechanistic model for how two closely connected transcription factors can exert distinct and coordinated genomic functions during terminal B cell differentiation.

      Strengths:

      A major strength of the study is the establishment and detailed characterization of a human in vitro plasma cell differentiation system. The authors combine phenotypic, functional, and single-cell transcriptomic analyses to define the transition from activated B cells to plasmablast/plasma cell precursor-like cells and then to more mature plasma cells. This system is very useful for future perturbation studies of human plasma cell differentiation.

      A second strength is the stage-specific perturbation strategy. By targeting IRF4 or PRDM1 at the precursor-enriched stage, the authors avoid some of the interpretive limitations associated with earlier perturbations that would affect B cell activation, proliferation, and plasma cell commitment simultaneously. The distinct phenotypes observed after IRF4 versus PRDM1 perturbation provide support for a model in which these two factors act in a temporally ordered manner.

      A third strength is the integration of multiple genomic and biochemical approaches. The combination of single-cell RNA-seq, chromatin accessibility profiling, CUT&RUN, computational motif analysis, and EMSA assays provides a rich dataset and supports the idea that ISRE/EICE sequence variation contributes to differential IRF4 and BLIMP1 occupancy.

      Weaknesses:

      While the multi-omic approach and computational modeling are highly impressive, several major assumptions regarding the cellular differentiation model and genomic linkages require more rigorous validation.

      First, because CRISPR editing was performed on heterogeneous bulk Day 7 cells rather than purified precursor populations, it remains ambiguous whether the observed developmental blocks are truly specific to the prePC window.

      We agree that CRISPR/Cas9 editing of bulk D7 cultures complicates interpretation because this population contains both activated B cells and PB/prePCs. We will therefore revise the text to distinguish phenotypic effects measured across the bulk D7 culture from the downstream single-cell analysis focused on cells along the prePC-to-PC trajectory. In particular, our interpretation of IRF4 and BLIMP1 function in prePCs is based primarily on the D9 scRNA-seq analysis, in which cells arrested in the activated B cell compartment are not used to define the perturbed PC-trajectory states. We will clarify this analytic design in a future revision and temper language implying that all effects arise exclusively within prePCs.

      Second, given that IRF4 and BLIMP1 operate within a mutually reinforcing positive feedback loop, the phenotypic divergence between IRF4 KO and PRDM1 KO may reflect differences in protein degradation kinetics or hierarchical dominance rather than a strictly ordered "sequential function".

      We agree that the divergence between IRF4 and PRDM1 perturbations could reflect differences in protein turnover, or hierarchical dominance, in addition to developmental timing. We will revise the Discussion to state that our data support a temporally ordered model in which IRF4 acts early to license the prePC-to-PC transition and BLIMP1 consolidates the terminal state, but that the current experiments do not exclude alternative explanations related to hierarchical dominance or degradation kinetics. We will also note in the revised Discussion that degron-based perturbations, rescue experiments, and gain-of-function analyses would be needed to resolve the functional ordering of IRF4 and BLIMP1 with higher temporal precision.

      Lastly, the motif-lexicon model is elegant and supported by biochemical DNA-binding assays, but the link between motif variation and gene regulation in cells remains partly correlative. (1) Direct testing of selected regulatory elements would make the causal claim stronger. (2) Alternatively, the authors should temper the language and present the motif lexicon as a predictive model for differential occupancy rather than as a fudlly demonstrated mechanism of gene regulation.

      We agree that the current data support the motif lexicon primarily as a predictive model for differential TF occupancy rather than as a fully causal mechanism of gene regulation. We will therefore revise the relevant text in the Results and Discussion. The EMSA data directly test nucleotide-dependent binding preferences, and the CUT&RUN/multiome analyses show that these motif variants are differentially associated with IRF4- or BLIMP1-bound DEG-linked OCRs. However, direct causal testing of endogenous regulatory elements, for example by base editing of selected ISRE/EICE variants, will be required to determine whether these variants are sufficient to predictably alter gene activity in differentiating plasma cells.

      Reviewer #2 (Public review):

      Summary:

      The manuscript by Lau et al. investigates the mechanisms underlying IRF4 and BLIMP1 transcriptional activities during antibody-secreting cell fate decision. Both master regulators of plasma cell differentiation, these two transcription factors have distinct targets and non-overlapping roles. The authors used an in vitro culture system to generate antibody-secreting cells from human naïve B cells, and scRNA-seq, Crispr Cas9 editing, and Cut&Run to dissect the molecular mechanisms defining their specificity.

      Strengths:

      The experiments are overall well executed, and the manuscript is well written. The in vitro culture model appears to generate genuine human antibody-secreting cells. The identification of non-conserved nucleotides within the binding motifs that induce the specific binding of IRF4 or BLIMP1 is convincing, novel, and exciting.

      Weaknesses:

      The authors need to correct some overstatements and flaws to improve the manuscript.

      In Figure 1f, the authors aimed to determine whether in their culture system the plasma cells emerged from the plasmablasts or directly from the activated B cells. First, it is noticeable that the distinction between plasmablasts and plasma cells relies here only on the expression of CD138. It does not include a higher capacity to secrete antibody or their proliferative state. In Figure 1e, the authors could have strengthened their distinction by showing the Ki67 staining at day 21 for both subpopulations.

      Second, this question does not seem to be related to IRF4 or Blimp1 activity, and thus one could wonder if it is relevant to this study.

      Finally, and most importantly, the design of the experiment appears flawed to me. The authors sorted cells at day 7 of culture based on their expression of CD20 and put the two subpopulations back for 14 more days. This culture system is a stepwise system, and it is not specified if the CD20<sup>+</sup> cells were put back in the day 7 condition or the day 0 condition with the CD40L stimulation

      We agree that CD138 alone does not fully define terminal PC maturation. In the revised manuscript, we will clarify that CD138 was interpreted in the context of a broader maturation profile, including CD20 downregulation, ICAM2 upregulation, IRF8 loss, IRF4/BLIMP1 expression, Ki-67 loss, and antibody secretion. The D7 PB population was proliferative and CD138<sup>-</sup>, whereas D21 CD20<sup>-</sup> cells were largely Ki-67<sup>-</sup> and included CD138<sup>+</sup> cells, supporting their progressive maturation. We will include Ki67 analysis in CD138<sup>-</sup> and CD138<sup>+</sup> cells at D21 in the revision.

      We agree that the motivation and culture conditions for this experiment required a clearer explanation. The purpose of the D7 sort-and-reculture experiments was to identify the developmental window enriched for cells competent to generate PCs, thereby defining the stage at which IRF4 and PRDM1 should be perturbed. Sorted D7 CD20<sup>+</sup> actB cells and CD20<sup>-</sup>CD38<sup>+</sup>CD27<sup>+</sup> PBs were both placed into the same D7-D14 differentiation conditions, allowing a direct comparison of their PC-generating competence under identical culture conditions. We will clarify this design in the Results and Methods. We have not tested whether returning D7 CD20<sup>+</sup> cells to D0 conditions involving CD40L stimulation restores PC differentiation, and we will now acknowledge in the revision that the CD20<sup>+</sup> fraction may contain cells with distinct intrinsic differentiation potential, including cells differentiating into non-PC states.

      What if it is not the case and some are anergic or have committed to the memory B cell fate during the first 7 days? Then the day 7 CD20<sup>+</sup> fraction would be enriched in these cells.

      We agree that the CD20<sup>+</sup> D7 cells may contain anergic or memory B cell precursors. However, this does not alter the interpretation that the CD20<sup>-</sup> (CD38<sup>+</sup>/CD27<sup>+</sup>) PBs contain a PC precursor population. Even if memory B cells are generated in the CD20<sup>+</sup> fraction by D7, based on the sorting experiments, their presence would have little-to-no effect on developing PC precursor populations.

      Moreover, this experiment didn't show that the plasma cell derived from the plasmablasts in the strict sense of the term, as the CD138<sup>+</sup>CD20- cells could be a mix of proliferative plasmablasts and immature plasma cells.

      We agree that the heterogeneous nature of CD20<sup>-</sup> cells complicates the interpretation. However, we would like to emphasize that all D7 CD20<sup>-</sup> cells are Ki67<sup>+</sup> whereas all D21 CD20<sup>-</sup> cells are nearly all Ki67<sup>-</sup>. Though we concede these could include recently proliferated PCs at D21, it is consistent with this population becoming quiescent. To better address this question in a future revised version, we will include direct measurements of Ki67 levels in CD138<sup>+</sup> and CD138<sup>-</sup> cells at D21.

      In Figure 3a and thereafter, the authors claimed that IRF4 acted earlier than BLIMP1, but both deletions strongly affected differentiation at day 7. IRF4 might have a stronger effect, but it does not mean that it had an earlier effect. To substantiate their claim, the authors would need to demonstrate that, at an earlier time point, deletion of IRF4, but not BLIMP1, results in defective differentiation.

      We agree that the current data do not by themselves prove that IRF4 acts earlier than BLIMP1 in developmental time. We will revise the text to state that the data are consistent with a temporally ordered model, rather than demonstrating strict sequential action. The latter interpretation is based on the distinct IRF4 KO stunted PC state observed by D9 scRNA-seq (Fig. 3C), together with the stronger early phenotypic effect of IRF4 loss (Fig. 3A). However, because both factors are mutually reinforcing and because perturbations were not performed across multiple time points, alternative explanations remain possible, including differences in editing efficiency, protein stability, and feedback-dependent TF decay. We will modify the text to better explain the rationale behind this interpretation while also acknowledging alternative interpretations that do not involve sequential IRF4-BLIMP1 functions (see response to Reviewer 1).

      In Figure 3b, the authors stated that in each individual KO the expression of the other transcription factor was lower. Given that there were no cells in the gate, it is puzzling to figure out how these expressions were compared.

      In Figure 3c, on the UMAP the bottom right part of the activated B cell cluster does not appear to be attributed to any condition. How can it be? Besides, it is highly surprising that at D9 we cannot see any plasmablast on these UMAP, even in the control. Based on the G1/S and G2/M scores, none of the ASC represented were proliferating. Could the authors explain this strong discrepancy with Figure 1?

      Another discrepancy exists between Figure 3b and c: Figure 3b depicted no IRF4- or BLIMP1-expressing cells in either KO, so what were the stunted PC and the BLIMP-KO PC reported in Figure 3c?

      We thank the reviewer for identifying these points of confusion. We will revise Fig. 3B to display outlier events and frequencies more clearly and revise the figure legend to clarify the donor origin of the displayed UMAPs and corresponding supplemental analyses. We will also clarify that Fig. 3B and Fig. 3C represent distinct readouts: flow cytometry measures IRF4 and BLIMP1 protein abundance, whereas scRNA-seq resolves transcriptional states after perturbation. Thus, the “stunted PC” state in IRF4 KO cells is defined at a transcriptional level as a population positioned between prePCs and PCs, not as a population retaining normal IRF4 or BLIMP1 protein expression. We further clarify that the apparent reduction in proliferative plasmablast-like cells at D9 likely reflects both the later timepoint relative to D7 and differences between transcriptional cell-cycle gene scores and Ki-67 protein persistence.

      What are the signature genes defining pre-PC and the score depicted in Supplementary Figure 3d, as the materials and methods only state that they are intermediate between PC and B cells?

      The signature genes defining the scores in Fig. S3D are listed in Table S2. We will clarify the source of these genes in the figure legend of a future revised version.

      Could the authors show IRF4, BLIMP1 and some of their known target expression in these populations?

      We thank the reviewer for this suggestion to highlight IRF4, BLIMP1 and exemplar target genes in the various populations. We will update Fig. S3 to show transcript levels of IRF4, PRDM1 and an example of one of each of their target genes in unperturbed cells to demarcate their normal expression pattern.

      The authors claim that BLIMP1 is not needed to initiate the transition from pre-PC to PC, but in Figure 1, the intracellular staining showed that at day 7 the antibody secreting cells already expressed BLIMP1. This would rather suggest that BLIMP1, unlike IRF4, does not need to be maintained once the cell reaches a certain point.

      We agree that the data do not exclude the possibility that BLIMP1 is required before the perturbation window but is less continuously required once cells have progressed beyond a defined prePC stage. Our statement that BLIMP1 is not required to initiate the prePC-to-PC transition is based on the observation that PRDM1 KO cells did not accumulate in the prePC or stunted PC intermediate states observed after IRF4 loss. We will revise the text to make this interpretation more precise by stating that BLIMP1 appears less important than IRF4 for progression into a PC-like transcriptional state but is required for efficient consolidation of the mature PC program. We will also acknowledge that differences in editing efficiency, protein persistence, and timing of BLIMP1 action could contribute to the observed differences in phenotypes.

    1. eLife Assessment

      This important study shows that cave-adapted Astyanax mexicanus have shifted from avoiding alarm and decay odors to being attracted by them, alongside sex-specific responses to social odors. The evidence for the behavioral and heritability claims is convincing, supported by analyses of different cave populations and F2 hybrids, as well as starvation-induced plasticity experiments. Whole-brain pERK mapping offers suggestive mechanistic insight, although uncertainties in anatomical assignments and aspects of statistical and ethogram reporting temper the strength of the neurobiological conclusions. The work will be of interest to anyone working on the evolution of behavior.

    2. Reviewer #1 (Public review):

      In their submitted manuscript, Harkinish-Murray and colleagues from the Kozol lab present convincing evidence for a genetically encoded shift in the odor perception of cavefish compared to their surface ancestors. Surface Astyanax, just as zebrafish, are attracted to food odors and are repelled by death odors and the alarm substance Schreckstoff (released from damaged skin by specialized club cells). Based on the experimental evidence in this manuscript, however, their cavefish counterparts are attracted to these odors as well. This would make sense, in an evolutionary framework, as predation is less likely in cave settings and decaying fish are a valuable source of nutrients for their living counterparts.

      Using an F2 hybrid cross scheme between surface fish and cavefish, authors also provide compelling evidence that genetic factors are behind this behavioral shift. Furthermore, they also show that this behavior (i.e., attraction to skin and decay extracts) can be observed in surface fish given long enough food deprivation. This latter observation also makes sense in the light of evolution and is genuinely interesting as it also provides a plausible roadmap to the shift in behavior through Waddingtonian genetic assimilation.

      The manuscript is generally well written and clear, we have identified only few weaknesses, some regarding the presentation of the data.

      (1) For Figure 3, on the x-axis of panels b, e, and h, supposedly we see surface fish vs. different cavefish populations. This is currently missing and makes the figure harder to interpret. Also, two populations (panel e) show a bimodal distribution upon indirect white light exposure, suggesting that some fish still acted as if they were exposed to direct light, while others acted as if they were in darkness (infrared light). We believe this warrants more consideration as it could tell us something about the existing (and relevant) genetic variance within this population. It is also notable that the third cavefish population also showed increased odor indices under indirect white light and infrared light conditions, suggesting that increasing the number of observations could have yielded a statistically significant result.

      (2) Some extra details about the methods could also be provided to enhance the reproducibility of the experiments.

      (3) A more serious concern is about the anatomical designation of particular brain regions in Figure 7d and consequently Figure 7f. Whereas we would agree with the positioning of the medial pallium (Dm), we think the region depicting the thalamus is in fact still part of the telencephalon, and the real thalamus should be more posteriorly. On the other hand, we think that the preoptic areas should be under the pallium and not posterior to it (see PMID: 22586363 for corresponding zebrafish anatomy). We would suggest, therefore, that the authors revisit this issue (a minor one, considering the depth of the results presented in the manuscript), and provide a better anatomical annotation - e.g., the identity of particular brain regions could be backed up by Hybridization Chain Reaction experiments for region-specific transcripts. (Disclaimer: we do not consider ourselves experts in adult cavefish neuroanatomy; therefore, we consulted in this case a colleague with much more knowledge on this topic.)

      (4) It would also be useful to expand the brain imaging data displaying results for similar tests in surface fish, to see if skin and decay extracts trigger different or similar brain activity in those fish.

      Further work will surely be able to discern the more precise genetic changes that made the shift in behavior possible. Once these causative variants (or at least linked markers) are determined, it will be quite revealing to see if these variants are indeed already present in the surface population (as hinted by the authors), and also, if besides the Surface x Tinaja F2 hybrids, crosses between other cave populations and surface fish can be performed, we could also see how much evolutionary convergence happened in the parallel evolution of different cave morphs. Were there multiple possible pathways for similar behaviors in different cave populations, or - as in freshwater stickleback populations - do we see broadly the same genetic playbook repeated each time?

      Another outstanding question, also demonstrated and discussed, albeit briefly, in this paper relates to the behavior-modulating effect of light in cavefish. What is the physiological relevance for a dark-dwelling animal to have this capacity? Is this just the chance result of occasional gene flow from surface populations, or does it have a genuine evolutionary significance?

    3. Reviewer #2 (Public review):

      Summary:

      The authors tested whether the olfactory cues that drive attraction or avoidance behavior have diverged between surface‑dwelling and cave‑adapted strains of the Mexican cavefish Astyanax mexicanus. They use high‑throughput odor‑discrimination assays between known attractants and repellents by calculating an "odor index" per fish (=the difference in time spent in an odor zone versus a control zone). Further, hybrid crosses to probe heritability, starvation experiments to assess plasticity of odor perception, and whole‑brain pERK detection/mapping to link behavioral changes with known localized neural activity. The results support the hypothesis that the extreme cave environment has selected for an approach response to stimuli that are ancestrally aversive (like alarm or death odors) but in harsh environments can be used as guidance to the rare food sources in this ecosystem.

      Strengths:

      The odor index analysis is convincing, and the experiments for odor attraction/avoidance are robustly performed. The light-to-darkness shift reflected by avoidance to attraction in cavefish towards skin odors is compelling and carefully analyzed. The analysis of odor indices of three cave-dwelling populations in comparison to surface fish highlights a similar regime, yet with differences among the different populations, suggesting population-specific genetic variation.

      Another strength of the paper is exactly this genetic inheritance study by generating F2 hybrids of cave-dwelling and surface-living individuals. The hybrids displayed a continuous range of odor indices for social, alarm, and death odors, indicating that these traits are heritable and likely based on additive genetic markers. Further, the authors uncovered a sexual dimorphism: only female cavefish exhibited approach behavior to social odors, whereas males remained neutral. This result aligns with known differences in olfactory organ morphology between sexes of other species from harsh environments.

      Although limited in number, the neurophysiological correlation using whole‑brain pERK mapping after 10 min of odor exposure is convincing. The data revealed overlapping activation in the thalamus and pre‑optic region for food and decay odors, suggesting that these brain areas mediate the evolved attraction response to previously repellent stimuli.

      Overall, the manuscript presents a concise story: cavefish have evolved attraction to alarm and death odors as a result of shifting from ancestral avoidance-driven to attraction by genetic changes and physiologically similar activation of specific neural circuits. The evidence is robust, with multiple independent experiments (behavioral assays, hybrid genetics, starvation experiments, and brain mapping) that collectively support the conclusions.

      Furthermore, exposure to unpleasant odors can not only be tolerated but can even serve as a trigger for foraging. This plasticity demonstrates that genetic predispositions can be put into practice through active changes in physiology in species or organisms confronted with (drastically) changing environmental conditions.

      Weaknesses:

      I value that the authors are critical of their own data, indicating low numbers in the pERK/brain experiments. Yet this is a weak point as the statistical power is thus limited. However, their reasoning is careful, based on the results and not over-interpreting.

      The layout/design of the ethograms (bout category plots) for both individual and population-wise are not easy to follow. Reworking these display items to convey the information is necessary.

      Taken together, the manuscript uses odor perception and attraction/avoidance behavior studies to show that environmental changes (light-to-darkness) have an immediate impact on smell perception and behavior. Attraction to otherwise repellent odors is used by cavefish to likely adapt to harsh environments with low food sources. The manuscript convincingly demonstrates this plasticity, which is an interesting idea to follow up for other traits spreading among a population. This also underlines that a genome may be fixed and the blueprint for behavioral traits, but extrinsic cues can readily be adapted to change wired behavior even to the extreme as reported here: changing avoidance to attraction.

    4. Author response:

      The authors thank the reviewers for their thorough and fair assessment of our manuscript. We are currently working to edit the manuscript based on the critiques and guidance offered by the reviewers. This will consist of fixing grammar and typos, expanding the material and methods section to include more information on the behavioral assays, modifying graphs for clarity between visuals and interpretations, and correcting our mistakes in neuroanatomical labeling.

      Public Reviews:

      Reviewer #1 (Public review):

      In their submitted manuscript, Harkinish-Murray and colleagues from the Kozol lab present convincing evidence for a genetically encoded shift in the odor perception of cavefish compared to their surface ancestors. Surface Astyanax, just as zebrafish, are attracted to food odors and are repelled by death odors and the alarm substance Schreckstoff (released from damaged skin by specialized club cells). Based on the experimental evidence in this manuscript, however, their cavefish counterparts are attracted to these odors as well. This would make sense, in an evolutionary framework, as predation is less likely in cave settings and decaying fish are a valuable source of nutrients for their living counterparts.

      Using an F2 hybrid cross scheme between surface fish and cavefish, authors also provide compelling evidence that genetic factors are behind this behavioral shift. Furthermore, they also show that this behavior (i.e., attraction to skin and decay extracts) can be observed in surface fish given long enough food deprivation. This latter observation also makes sense in the light of evolution and is genuinely interesting as it also provides a plausible roadmap to the shift in behavior through Waddingtonian genetic assimilation.

      The manuscript is generally well written and clear, we have identified only few weaknesses, some regarding the presentation of the data.

      (1) For Figure 3, on the x-axis of panels b, e, and h, supposedly we see surface fish vs. different cavefish populations. This is currently missing and makes the figure harder to interpret. Also, two populations (panel e) show a bimodal distribution upon indirect white light exposure, suggesting that some fish still acted as if they were exposed to direct light, while others acted as if they were in darkness (infrared light). We believe this warrants more consideration as it could tell us something about the existing (and relevant) genetic variance within this population. It is also notable that the third cavefish population also showed increased odor indices under indirect white light and infrared light conditions, suggesting that increasing the number of observations could have yielded a statistically significant result.

      We agree with the reviewer that our light testing data suggests complexity in the response to indirect white light within certain cave populations. In addition, an expanded sample size would likely provide clarity on whether individuals fall within two groups, behavior that looks like direct light or infra-red light, that could relate to genetic variation within cavefish populations. We are currently working to reassess the current data and determining the best course of action for continued studies related to light exposure.

      (2) Some extra details about the methods could also be provided to enhance the reproducibility of the experiments.

      We agree with both reviewers that the methodological section on behavior needs to be expanded. We are currently editing our methods section to include more detail on water exchanges, odor preparation, timing, biological replicates, and binning.

      (3) A more serious concern is about the anatomical designation of particular brain regions in Figure 7d and consequently Figure 7f. Whereas we would agree with the positioning of the medial pallium (Dm), we think the region depicting the thalamus is in fact still part of the telencephalon, and the real thalamus should be more posteriorly. On the other hand, we think that the preoptic areas should be under the pallium and not posterior to it (see PMID: 22586363 for corresponding zebrafish anatomy). We would suggest, therefore, that the authors revisit this issue (a minor one, considering the depth of the results presented in the manuscript), and provide a better anatomical annotation - e.g., the identity of particular brain regions could be backed up by Hybridization Chain Reaction experiments for region-specific transcripts. (Disclaimer: we do not consider ourselves experts in adult cavefish neuroanatomy; therefore, we consulted in this case a colleague with much more knowledge on this topic.)

      We agree that our annotation was incorrect or more accurately mislabeled in our write-up of the preprint and submitted manuscript. Therefore, we have now re-assessed the regions using the tissue cleared and light sheet collected zebrafish atlas, Adult Zebrafish Brain Atlas (AZBA; doi: 10.7554/eLife.69988). We are now editing the resubmission in the following manner: our initial labeling of the ventromedial thalamus will be changed to the lateral olfactory tract (nLOT) of the pallium and the preoptic region to the ventromedial thalamus (VM). We will provide a comparable z-slice of the AZBA segmentation file to illustrate the similarity in position. This would also support a known continuous circuit of olfactory integration, with information flowing from the lateral olfactory tract-to the piriform cortex-to the thalamus. We also agree that a more accurate assessment in Astyanax would require HCR in situ hybridization of markers for those specific brain regions or a neurocomputational brain atlas for adult Astyanax populations. Finally, we assert that this small dataset is preliminary at best and only provides regions of shared activity that could explain anything from perception related processes to relay of odor signaling unrelated to perception. Further work with larger sample sizes and additional populations are currently underway for a follow-up study on the neurobiological basis of olfactory processing and perception in adult cavefish.

      (4) It would also be useful to expand the brain imaging data displaying results for similar tests in surface fish, to see if skin and decay extracts trigger different or similar brain activity in those fish.

      We agree with the reviewer that the brain mapping section lacks a sufficient sample size and no control group for comparison (surface fish). However, we found the variation in pERK intensity (notably the putative nLOT) to be informative and decided to include the dataset in the manuscript. We are currently working to fill in these data gaps by sampling all populations and increasing the Pachon cavefish sample size. This will be a follow-up study as mentioned above in the last rebuttal paragraph.

      Further work will surely be able to discern the more precise genetic changes that made the shift in behavior possible. Once these causative variants (or at least linked markers) are determined, it will be quite revealing to see if these variants are indeed already present in the surface population (as hinted by the authors), and also, if besides the Surface x Tinaja F2 hybrids, crosses between other cave populations and surface fish can be performed, we could also see how much evolutionary convergence happened in the parallel evolution of different cave morphs. Were there multiple possible pathways for similar behaviors in different cave populations, or - as in freshwater stickleback populations - do we see broadly the same genetic playbook repeated each time?

      We agree with the reviewer that the hybrid results setup a promising follow up project to map these traits genetically. We are continuing to test odor perception in other hybrid populations and have started Quantitative Trait Locus mapping experiments.

      Another outstanding question, also demonstrated and discussed, albeit briefly, in this paper relates to the behavior-modulating effect of light in cavefish. What is the physiological relevance for a dark-dwelling animal to have this capacity? Is this just the chance result of occasional gene flow from surface populations, or does it have a genuine evolutionary significance?

      Reviewer #2 (Public review):

      Summary:

      The authors tested whether the olfactory cues that drive attraction or avoidance behavior have diverged between surface‑dwelling and cave‑adapted strains of the Mexican cavefish Astyanax mexicanus. They use high‑throughput odor‑discrimination assays between known attractants and repellents by calculating an "odor index" per fish (=the difference in time spent in an odor zone versus a control zone). Further, hybrid crosses to probe heritability, starvation experiments to assess plasticity of odor perception, and whole‑brain pERK detection/mapping to link behavioral changes with known localized neural activity. The results support the hypothesis that the extreme cave environment has selected for an approach response to stimuli that are ancestrally aversive (like alarm or death odors) but in harsh environments can be used as guidance to the rare food sources in this ecosystem.

      Strengths:

      The odor index analysis is convincing, and the experiments for odor attraction/avoidance are robustly performed. The light-to-darkness shift reflected by avoidance to attraction in cavefish towards skin odors is compelling and carefully analyzed. The analysis of odor indices of three cave-dwelling populations in comparison to surface fish highlights a similar regime, yet with differences among the different populations, suggesting population-specific genetic variation.

      Another strength of the paper is exactly this genetic inheritance study by generating F2 hybrids of cave-dwelling and surface-living individuals. The hybrids displayed a continuous range of odor indices for social, alarm, and death odors, indicating that these traits are heritable and likely based on additive genetic markers. Further, the authors uncovered a sexual dimorphism: only female cavefish exhibited approach behavior to social odors, whereas males remained neutral. This result aligns with known differences in olfactory organ morphology between sexes of other species from harsh environments.

      Although limited in number, the neurophysiological correlation using whole‑brain pERK mapping after 10 min of odor exposure is convincing. The data revealed overlapping activation in the thalamus and pre‑optic region for food and decay odors, suggesting that these brain areas mediate the evolved attraction response to previously repellent stimuli.

      Overall, the manuscript presents a concise story: cavefish have evolved attraction to alarm and death odors as a result of shifting from ancestral avoidance-driven to attraction by genetic changes and physiologically similar activation of specific neural circuits. The evidence is robust, with multiple independent experiments (behavioral assays, hybrid genetics, starvation experiments, and brain mapping) that collectively support the conclusions.

      Furthermore, exposure to unpleasant odors can not only be tolerated but can even serve as a trigger for foraging. This plasticity demonstrates that genetic predispositions can be put into practice through active changes in physiology in species or organisms confronted with (drastically) changing environmental conditions.

      Weaknesses:

      I value that the authors are critical of their own data, indicating low numbers in the pERK/brain experiments. Yet this is a weak point as the statistical power is thus limited. However, their reasoning is careful, based on the results and not over-interpreting.

      We agree with the reviewer and direct their attention to the same critique by reviewer 1. We believe this is predominantly preliminary data that was included due to the conspicuous increase in pERK signal from the putative lateral olfactory tract (nLOT) for food and decay exposed cavefish. We are continuing to work on odor stimulated brain mapping and look forward to publishing a comprehensive dataset across wildtype and hybrid populations.

      The layout/design of the ethograms (bout category plots) for both individual and population-wise are not easy to follow. Reworking these display items to convey the information is necessary.

      We agree with the reviewer that the ethograms are challenging to read, especially due to our use of different colors for odor categories. We are currently preparing alternative graphs for displaying ethograms that reduce confusion and make following bout transitions for individual traces and bout probabilities for populations easier on the eyes.

      Taken together, the manuscript uses odor perception and attraction/avoidance behavior studies to show that environmental changes (light-to-darkness) have an immediate impact on smell perception and behavior. Attraction to otherwise repellent odors is used by cavefish to likely adapt to harsh environments with low food sources. The manuscript convincingly demonstrates this plasticity, which is an interesting idea to follow up for other traits spreading among a population. This also underlines that a genome may be fixed and the blueprint for behavioral traits, but extrinsic cues can readily be adapted to change wired behavior even to the extreme as reported here: changing avoidance to attraction.

    1. eLife Assessment

      Using sci-L3-Strand-seq, this study shows genome-wide, single-cell evidence that DSBs triggered by CRISPR/Cas9 are frequently resolved through sister chromatid exchange (SCE), which is undetectable with conventional whole-genome sequencing approaches. This important work provides helpful metrics quantifying the occurrence of SCE events at targeted and repetitive genomic loci that are of particular interest to those working in gene editing, DNA repair, genome instability, or repetitive genomic loci. However, the evidence supporting the proposed involvement of under-replicated region/replication-termination-zone resolution and TRAIP/URR-like pathways is currently incomplete and could be strengthened with an increased number of reciprocal daughter-cell pairs and by genetic or molecular perturbation, or alternatively, this can be addressed by changing the discussion.

    2. Reviewer #1 (Public review):

      Summary:

      This manuscript uses sci-L3-Strand-seq to map sister chromatid exchange events following CRISPR/Cas9-induced DNA damage. Because exchanges between identical sister chromatids are largely invisible to conventional sequencing, the study addresses an important blind spot in the assessment of genome editing outcomes. The authors compare single-locus Cas9 cleavage, simultaneous targeting of 237 repetitive genomic sites, and Cas9 nickase variants. They further use reciprocal daughter-cell pair analysis to ask whether Cas9-associated SCEs are copy-neutral or linked to larger structural alterations. Overall, this is a valuable study that introduces an important additional layer to the analysis of CRISPR/Cas9 repair outcomes. The central finding that Cas9-induced DSBs can trigger frequent local SCE is well supported and likely to be of broad interest. The evidence for structural complexity associated with some induced SCEs is intriguing, but the mechanistic interpretation should either be tested directly or presented more cautiously.

      Strengths:

      The major strength of the manuscript is the application of a strand-resolved, single-cell method to a question that is difficult to address with standard genome sequencing. The evidence that a single Cas9-induced DSB can trigger strong local SCE is compelling in concept and supported by multiple guide RNAs targeting distinct loci. The reported on-target SCE frequencies, reaching up to 41%, suggest that inter-sister exchange is a substantial and underappreciated outcome of Cas9 cleavage.

      Of particular interest is the comparison between single-site and multi-site targeting. The finding that 237 programmed Cas9 targets produce only mild bulk enrichment of on-target SCE but stronger enrichment in a subpopulation of cells with elevated SCE burden is interesting and may have wider biological implications, particularly if the findings extend beyond Cas9-induced SCE to spontaneous SCEs. Given that potential, the current manuscript would benefit greatly from any experiments characterizing this sub-population: are these cells in a particular cell cycle state, experiencing changes in gene expression, or do they have other unique biological properties?

      The reciprocal daughter-cell pair analysis is another notable feature of the study. The observation that some Cas9-associated SCEs are accompanied by structural alterations could challenge the assumption that SCE after a programmed break reflects error-free homologous recombination.

      Weaknesses:

      The number of informative RDCPs is limited, and the mechanistic interpretation of the "WWC-or-WCC/deletion" signature is more suggestive than definitive. In particular, the manuscript invokes (even though only in the Discussion section) URR or replication-termination-zone resolution and discusses TRAIP-dependent CMG unloading, nuclease cleavage, and polymerase theta-mediated joining, but these pathway components are not directly tested herein. A more conservative conclusion that some Cas9-associated SCEs coincide with structural alterations is more appropriate, particularly in the Discussion and Conclusion. For example, the statement that this work provides "direct genetic evidence" for a URR-type mechanism is overstated unless supported by additional experiments or a more extensive analysis of alternative models. Similarly, while the authors explain the limitations of acute Cas9 disruption of LIG3, LIG4, XRCC1, and XRCC4, the manuscript should clarify what biological questions this experiment can and cannot answer.

    3. Reviewer #2 (Public review):

      Summary:

      In this short paper, a clever single-cell Strand-seq method was used to study the number and location of sister chromatid exchange events (SCEs) in cells after CRISPR/Cas9-induced DNA double-strand breaks (DSBs). Unique as well as multiple genomic loci were targeted. Cas9-induced cuts at unique genomic locations led to statistical enrichment of SCEs at the target site, whereas Cas9 targeted at repetitive targets revealed only mild enrichment of on-target SCEs unless analysis was restricted to a subset of cells with >8 SCEs per cell. Interestingly, reciprocal daughter-cell pair analysis revealed large-scale structural alterations on some chromosomes. Whereas disruption of DNA repair genes, including LIG3, LIG4, XRCC1, and XRCC4, did not measurably alter SCE frequency per cell within 24 hrs, consistent with delayed functional loss following editing and selection against essential genes. Together, these findings demonstrate that Cas9-induced DSBs are potent local triggers of SCE at unique loci and can be associated with structural alterations, highlighting the influence of lesion type and genomic context on recombination outcomes during genome editing.

      Strengths:

      The data in this paper represent a very rich resource of how parental DNA template strands are distributed in paired daughter cells after various treatments. Abnormalities observed in only one of such paired daughter cells provide a novel and exciting approach to study mechanisms of DNA instability and DNA repair at a genome-wide level in general and following Cas9-induced DSB in particular.

      Weaknesses:

      The effect of Cas9-induced DSBs in the cells that are used will depend on the cell cycle stage of the cells that are targeted, as well as the number of times cuts are made. The latter could happen before, during, and after DNA repair reactions on one or both alleles in a diploid cell. As a result, it is very difficult to extrapolate the mechanisms of DNA instability and DNA repair from the observed genomic rearrangements. Novel approaches are needed to limit the number and timing of Cas9-induced breaks to overcome some of these limitations. The language and logic in the paper can be improved, and some of the claims seem incorrect. For example, the abstract reads "A single Cas9 cut at a unique genomic locus led to strong local enrichment of SCE at the break site, reaching up to 41% in the same cell cycle and 17% in the subsequent division, indicating that DSB repair frequently engages non-local inter-sister repair." The evidence that only a single Cas9 cut was made is lacking (see my earlier comment); it is not clear how local enrichment or non-local inter-sister repair are defined.

    4. Reviewer #3 (Public review):

      Summary:

      Chovanec and Yin used their newly developed sci-L3-Strand-seq powerful method to characterize SCE after Cas9 cleavage in a human cell line, using either a single target site or an element repeated 237 times in the genome. SCE are often neglected in DNA repair analyses since they are “genetically silent”. Interestingly, the authors found enrichment of SCE at unique Cas9 sites, but only a modest enrichment of SCE when Cas9 targets 237 sites in the genome. The genetic control of SCE formation at Cas9 sites is not deliberately addressed in this paper. However, the authors found that targeted SCE seem to be enriched in a subpopulation of cells, particularly “permissive” for SCE, but the determinants of such a population are unknown. Finally, the power of the sci-L3-Strand-seq allowed the authors to characterize a specific type of SCE based on the analysis of reciprocal daughter-cell pairs' genomes that is associated with a specific type of chromosomal rearrangement compatible with the ones observed in HR defective BRCA1/2 deficient cells.

      Strengths:

      This is an interesting paper that molecularly explores sister chromatid exchanges, which represent an important challenge in molecular biology since they are genetically silent.

      Weaknesses:

      A complexity of the current paper is that it heavily relies on a recently published paper (Chovanec et al 2026, NAR) describing the powerful but complex technique sci-L3-Strand-seq. Knowledge of this paper is a prerequisite to understanding the current manuscript because no reminder is provided. In addition, the current manuscript presents the use of the sci-L3-Strand-seq technique in the study of SCE after Cas9-induced DSBs, while a companion study is referred to several times for containing results about SCE in XRCC1 KO. At some point, one questions the relevance of splitting the use of sci-L3-Strand-seq in different papers instead of making a single integrated one.

    5. Author response:

      We thank the editors and reviewers for their thoughtful evaluation and constructive feedback. We are pleased that all three reviewers recognize the importance of mapping Cas9-induced sister chromatid exchanges (SCE) as a previously invisible repair outcome, and that the RDCP analysis is a notable feature of the study.

      We note that since our manuscript was posted, two companion studies in Science have provided direct biochemical evidence for the TRAIP-dependent pathway we discussed:

      (1) Fujisawa & Labib (Science, 2026; DOI: 10.1126/science.aeh2300) showed that TTF2 bridges CDK1-phosphorylated TRAIP to DNA Polymerase epsilon in the replisome, triggering mitotic CMG helicase disassembly, fork cleavage, and repair via SCE. Loss of this pathway reduced replication stress-induced SCE approximately two-fold in mouse ES cells.

      (2) Can et al. (Science, 2026; DOI: 10.1126/science.aeh1834) independently identified the same CDK1-TTF2-TRAIP axis in Xenopus egg extracts and validated it in HCT116 cells, showing that disrupting the TRAIP-TTF2 interaction reduced common fragile site deletions.

      We will incorporate these references in the revised discussion while still framing our RDCP observations as consistent with, rather than definitive proof of, this pathway.

      Below we briefly address the main points raised in the public reviews.

      Reviewer #1:

      We agree that the mechanistic interpretation of the RDCP signature should be presented more cautiously. We will reframe the URR/TRAIP discussion as a model, replacing language such as "direct genetic evidence" with "consistent with." We will add a summary table of RDCP data. We will also expand the description of rescued SCE calls. We will clarify what the DNA repair gene targeting experiment can and cannot answer (delayed protein loss, essential-gene selection) - we think that there is a notable difference at the bulk vs. at the single-cell level depending on the nature of the assay. Fig.1 fonts, labels, and pileup plot descriptions will be improved.

      We agree that the high-SCE subpopulation is particularly interesting and we cannot currently distinguish higher RNP uptake, a permissive cell-cycle state, altered expression, or stochastic variation. This may be better explored by future co-assays with sci-L3-Strand-seq.

      Reviewer #2:

      We agree with the limitations that Cas9-induced DSBs can be dependent on the cell cycle stage and the number of times cuts are made. We will add a brief discussion on this limitation in extrapolating the mechanisms of DNA instability and DNA repair from the observed genomic rearrangements.

      We agree that "a single Cas9 cut" should be revised to "Cas9 targeting of a single genomic locus" to accurately reflect the experimental design. We will clarify the possibility of multiple rounds of cutting at the same sites. We will also clarify "non-local" by modifying Fig.1 - we used this term specifically to include the possibility of inter-sister NHEJ.

      Reviewer #3:

      We will improve the self-contained nature of the manuscript so that readers need not consult the earlier NAR papers, and the companion preprint to understand the key results.

    1. eLife Assessment

      This valuable study provides insights into the developmental regulation of the unusual form of holometabolous metamorphosis that occurs in the black soldier fly, in which a distinct prepupal stage is interposed between the final larval instar and pupation. This type of life history strategy is similar to what is seen in insects at the hemimetabolous-holometabolous boundary, but, given the phylogenetic position of dipterans, it is likely to be a derived trait. The combination of developmental characterization, gene expression profiling, and RNAi-mediated functional analyses provides solid evidence supporting the authors' conclusions and will be of interest to researchers studying insect development and the evolution of metamorphosis.

    2. Joint Public Review

      Summary:

      In this study, the authors investigated the developmental and molecular basis of the unusual metamorphic program of the black soldier fly, Hermetia illucens, which differs from the canonical holometabolous life cycle by inserting a distinct, non-feeding prepupal instar between the final larval stage and pupation. Most insects that undergo complete metamorphosis molt to the final instar and then develop into the prepupal stage without molting. H. illucens, however, undergoes a molt before entering a non-feeding prepupal stage. Thus, it is an unusual, novel developmental strategy, and its regulation has remained a mystery. Through an integrated approach combining detailed morphological characterization, developmental gene expression profiling, and RNAi-mediated functional analyses of the core components of the Metamorphic Gene Network (MGN), the authors examine the developmental identity of this prepupal stage and how the temporal deployment of conserved metamorphic regulators has been reorganized to accommodate this atypical developmental program. In particular, they show that the prepupal stage expresses a distinct combination of the key genes known to regulate life history transitions, including unusually high levels of Br-C expression.

      Strengths:

      The study represents a valuable contribution to insect developmental biology. A major strength is the comprehensive characterization of postembryonic development, which establishes a robust developmental framework for H. illucens. This is complemented by detailed expression profiling and RNAi-based functional analyses of the Metamorphic Gene Network (MGN), comprising the temporal specifier factors, Kr-h1, chinmo, Br-C, and E93. The results show that these conserved regulators are deployed in a modified temporal sequence that accommodates the distinctive prepupal stage while largely preserving their canonical developmental functions. Together, the morphological, molecular, and functional data support the conclusion that the prepupal stage of H. illucens is a distinct developmental transition associated with a characteristic configuration of the metamorphic gene network. The results are supported by solid methodology and approaches and will serve as valuable resources for future investigations into insect development, the evolution of metamorphosis, and the diversification of insect life-history strategies.

      Weaknesses:

      While the study successfully establishes the developmental identity of the prepupal stage and its association with a modified temporal deployment of the MGN, some aspects of the proposed regulatory model are less directly supported by the experimental evidence.

      (1) Several regulatory interactions within the MGN remain inferential rather than experimentally demonstrated in H. illucens. In particular, the proposed relationship between juvenile hormone (JH), Kr-h1, and chinmo is based primarily on expression dynamics and RNAi-induced transcriptional changes. Although these observations are consistent with the proposed model, they do not directly demonstrate that JH induces chinmo expression or establish the regulatory relationship between Kr-h1 and chinmo in this species. As a result, the corresponding regulatory interactions presented in the final model should be regarded as plausible hypotheses rather than experimentally validated mechanisms.

      (2) A second limitation concerns the developmental role assigned to Br-C and E93 during the larval-to-prepupal transition. The authors conclude that sustained Br-C expression is a defining molecular feature of the prepupal stage and discuss its potential role in prepupal specification. However, the functional analyses of both Br-C and E93 were initiated only after larvae had already entered the prepupal stage. Consequently, while the RNAi experiments convincingly demonstrate essential roles for Br-C during the prepupal-to-pupal transition and for E93 during adult differentiation, they do not directly address whether either factor is required to trigger the formation of the prepupal stage itself. Therefore, the molecular mechanisms governing the initiation of this distinctive developmental transition remain unresolved. In particular, the proposed lack of repression of E93 by Br-C is only weakly supported, yet may be an essential feature of the prepupal stage of Hermetia illucens.

      (3) Although knockdowns of Kr-h1 and chinmo knockdowns look superficially similar, it would be good to confirm this with higher-magnification views of the cuticles for all three treatments (control, Kr-h1 RNAi, and chinmo RNAi). In other species, Kr-h1 knockdown leads to premature adult cuticle development, whereas chinmo knockdown typically leads to premature appearance of pupal characteristics. Similarly, in Fig. 4A and 4D, higher-magnification images of the cuticle would be helpful.

      (4) (Relating to Line 336 and Figure 7): "This low but persistent prepupal Kr-h1 expression, together with modest chinmo expression from PPD0 to PPD8, may be correlated to a JH-dependent antimetamorphic effect that maintains the prepupal stage." However, we are not aware of a function of JH in extending the prepupal stage. In addition, in most insects, the prepupal stage expresses high Kr-h1 expression; this peak likely prevents the animal from turning into an adult instead of the pupa. We presume the same holds true for H. illucens (although the lower expression of Kr-h1 during that stage is curious). As a result, we suggest that Fig. 7D be revised as it may be difficult to distinguish between pupal formation and prepupal maintenance given the experimental set-up. Fig. 7E may also need to be modified since the development of the pupa may require Kr-h1. It is worth noting that at the prepupal stage, JH and Br-C are co-expressed in many insects. If the authors think that Kr-h1 expression needs to be low at this time, this would imply a novel interaction between Kr-h1 and Br-C, and should be discussed.

    3. Author response:

      We thank the editors and reviewers for their careful evaluation and constructive suggestions. During the review process, we identified and corrected several presentation, terminology, citation, and figure-legend errors, and these corrections have been incorporated into the current version of preprint. Following the reviewers’ comment, we have also revised the title of the manuscript. We are now preparing a substantive revision that will distinguish more clearly between experimentally supported conclusions and hypothetical regulatory relationships. We plan to examine gene expression at an earlier time point after Br-c knockdown, further investigate the relationships among Kr-h1, chinmo, and E93 using additional RNAi experiments, characterize the cuticular phenotypes of precocious prepupae at higher magnification, and determine whether severe E93-knockdown individuals exhibit evidence of a repeated pupal developmental program. After completing these experiments, we will cautiously revise the proposed regulatory model and moderate conclusions that are not directly supported by the current evidence.

    1. eLife Assessment

      What can a neural network trained to imitate animal behavior tell us about biology? This valuable work uses deep reinforcement learning to train an artificial neural network to transform the dynamics of a recurrent neural network based on the C. elegans connectome into an adult Drosophila walking program in a physical model of the fly body, demonstrating that achieving plausible output dynamics does not in and of itself imply biologically meaningful simulation. Evidence for this basic claim is solid, but more extensive analyses, better methodological description, and a discussion of deeper challenges in the undertaking of biological brain modeling would strengthen the study. This result demands the attention of the practitioners of the growing field of connectome simulation for the purpose of gaining mechanistic understanding of nervous system function.

    2. Reviewer #1 (Public review):

      Summary

      The authors build a "digital sphinx" by stitching together two neural network models: (i) a recurrent network with fixed parameters derived from the C. elegans connectome and imputed physiological (e.g. neural input/output) functions, and (ii) a feedforward encoder-decoder model with learnable parameters intended to represent a central brain - to - motor interface, then harnessing the combined model to a Drosophila biomechanical model situated in a physics simulator, and finally using deep reinforcement learning (DRL) training to optimize the parameters of the encoder-decoder model to reproduce a set of spatiotemporal patterns of jointed limb activations that together produce the overall organismal behavior of walking, within the physics simulator.

      The primary intent of this paper is to dispel the recent grandiose claims made in the mainstream press by a private company, Eon Systems, to have achieved a major advance in biologically based brain simulation of the production of a set of ethologically relevant motor behaviors by the fly. Representatives of the company referred to this modeling and training process euphemistically and deceptively as "brain uploading". The authors proceed with a reduction-to-triviality exercise by constructing their own high-parameter dynamical brain-plus-body model situated in a physical simulation that produces, after training by reinforcement learning, satisfying ethological behavioral imitation in the same vein as the private company claim, but based on a clearly absurd and biologically unrealistic set of model assumptions.

      Secondarily, the paper provides two overall admonitions that they assert their computational demonstration illustrates: that training high parameter network models to imitate behavior, even if they possess some biological detail, will deliver little or no biological insight, and that models of behavioral generation must be built from detailed biological data and, crucially, developed in a hypothesis generation/falsification loop with experimental validation, in order to be scientifically useful.

      Appraisal

      The authors are well justified in challenging the non-rigorous claims of "uploading" or even the delivery of a neurobehavioral simulation with potential scientific utility, in unison with the vocal criticisms of many other researchers in the fields of AI and neuroscience, and it is an important message to deliver to the world. However, the authors' own modeling counter-exercise, while clever and vivid in imagery, suffers from its own lack of rigor, both in disclosure of implementation and in scientific case-making. Some sacrifice of clarity and thoroughness in the interest of brevity is inevitable under the brief format of this manuscript; however, we suggest that crucial additions and modifications should be made to avoid falling into a similar trap of non-rigorous sensationalism.

      Because the private company claims were not accompanied by a scientific paper, preprint, code repository, or much methodological disclosure of any kind, the authors have the particular challenge of building a refutation case against an undefined target. As a consequence, the authors chose their own task, model structure, and training paradigm.

      The authors argue that brain models need to be built from biological data to be useful for yielding biological insight. We agree with the overall principle; however, in practice, this procedure is fraught with epistemological difficulty. Biological modeling suffers from a unique challenge within the larger endeavor of scientific/physical modeling, which is that it is generally unclear as to precisely what biological quantities should be measured and at what level of detail they should be measured. Additionally, biological data will by necessity be incomplete and noisy, and thus decisions of coarse-graining must be made at the outset of large-scale data collection projects, and some, possibly a substantial, level of data imputation will have to be performed in order to build testable models in our lifetimes. Despite the astonishing success of scaling (in both parameter count and corpus size) in engineered neural networks for certain human-like tasks, it is not at all clear that simply adding more detail to biological models will produce deeper scientific insight, or whether cataloging parameters from snapshot data will yield functional simulations. The failed Blue Brain mega-project should provide a lesson, as well as Marder's longstanding work on parameter variation in neural systems. The coupled, pernicious questions of choosing measurement detail and modeling detail represent a deep, unsolved challenge area for the field, and this context should be raised in the text.

      The message about overinterpreting models trained with deep reinforcement learning, while valid and important, should be broadened to be a message about overinterpreting trained high-parameter models in general, in their ability to fit data or reproduce simple behavior. Other parameter optimization/learning procedures for building underdetermined and/or high-parameter models risk the same misinterpretation. The prescription of building models in conjunction with experimental prediction and validation is an important point.

      The authors leave out an additional important and underappreciated challenge of brain-model-building, which is that imitating a time segment of behavior is a computational task of unspecified, and possibly low complexity. Successful recapitulation of behavioral time series may simply not be considered cognitively interesting, even if the model is built entirely on biological data. While quantifying task complexity is another open area of computational and neuroscientific research, the authors should, at a minimum, describe their particular task data in explicit mathematical terms and preferentially provide some complexity analysis. In the absence of task complexity analysis, at a minimum, computational controls should be applied to demonstrate the necessity of whatever structure or data is being asserted in the model. This epistemological practice is glaringly absent in much, if not most, of the neurobehavioral modeling literature. This paper would be a good opportunity to set an example of rigor.

      Finally, the authors' description of prior work in the field of whole-organism neurobiological simulation feels incomplete and skewed toward work in Drosophila versus other model organisms. An internet search reveals many published efforts to build neurobehavioral models at varying levels of detail in C. elegans, of which only two are referenced.

      We do feel this work constitutes an illustrative scientific exercise and important counterpoint to the sensationalism building around efforts in neurobiological simulation. It should inspire further work in defining a rigorous and scientifically productive epistemological framework for these kinds of brain modeling efforts.

      Further Comments

      (1) The authors oversell the completeness and quality of connectome datasets and what they lack.

      Language such as "complete wiring diagrams," "nearly comprehensive connectomes" neglects the well-appreciated gaps in biological data that most practitioners believe necessary for useful, detailed models to be built. There is a brief mention that biological parameters "remain unknown" and that interfaces are "incompletely characterized", but beyond that, the authors do not explain which parameters are missing, why these parameters might matter, and what still needs to be addressed in order to make any plausible whole-brain emulation claims. This may also inadvertently bolster the sensationalist claims that the manuscript is trying to deflate by giving the impression that neurobiological and physiological data collection is a near-complete exercise.

      (2) Prior work in C. elegans neurobehavioral modeling should be more acknowledged, if nothing else, for why it has been largely unsatisfying.

      C. elegans is rarely discussed, while Drosophila is primarily focused on. The status of C. elegans connectomics, physiological mapping, biomechanics, and neurobehavioral modeling is worth more treatment.

      (3) Critiques of Eon Systems announcements also, by and large, apply to more detailed and disclosed efforts in neurobehavioral modeling using RL for parameter imputation, and this should be recognized.

      By way of reference to a tweet in the first paragraph, the authors are responding to a recent claim made by a startup that they have fully "uploaded" a fly brain, a significant advance vis-à-vis prior work in neurobehavioral modeling in Drosophila, such as references [3 and 9], which are mentioned as background in the paper but left out of the methodological critique. But one of the central warnings of the paper is around the challenge of interpretability when using reinforcement learning to optimize model parameters. The authors also should acknowledge that the use of RL has been justified by building neurobehavioral model builders as a proxy for the learning and tuning processes thought to occur during animal development.

      (4) Substantiate the reservoir computing explanatory claim with appropriate computational controls.

      The reservoir computing idea is the only piece of hypothesizing a necessary function for the central brain component model in the paper. This claim could be substantiated with some basic computational controls rather than just hypothesized. We suggest the following possibilities as additions to the model: (a) replace the connectome with an RRNN, (b) shuffle the connectome, or (c) use other simple dynamical systems in place of the worm brain model.

      Specific Manuscript Comments

      (1) Abstract

      "New connectome datasets and musculoskeletal models now enable integrated, closed-loop simulations of the neural and biomechanical systems of the fruit fly Drosophila, an ideal model organism to investigate embodied intelligence."<br /> This sentence could mislead non-specialists into thinking all current simulations are novel because the connectome datasets are new. In fact, FlyWire (2024), NeuroMechFly (2022), and other connectomes have already been available for some years now. We believe that this sentence is a chance to make the opposite point that these resources have existed for a while, and that many simulations have been built before.

      "However, many biological parameters of the nervous system and the body, as well as how they interface, remain unknown."<br /> Some examples of specific parameter/physiological data types that are missing and thought to be critical, such as neuronal input/output functions, are warranted. See below for a comment on the confusing construct of "interface" as a distinct entity from the neural network.

      (2) Introduction

      "Among animals that walk, the integration of brain wiring and body models is perhaps closest to fruition in Drosophila, due to the recent completion of multiple complete wiring diagrams (known as connectomes) of the fly nervous system." ...and... "The fly is the only animal with legs for which nearly comprehensive connectomes of its brain and nerve cord exist."<br /> The walking qualifier allows the authors to skirt around the substantial and decades-long work on connectomes in C. elegans, which crawls and does not walk. Yet sinusoidal crawling is a multidimensional, adaptive behavior, so it seems this exclusion was for narrative convenience rather than contextual accuracy.

      "Despite this progress, closed-loop integration of biomechanical and neural models remains far from straightforward."<br /> Work (and shortcomings) in C. elegans neurobehavioral modeling should also be stated here alongside the fly.

      "Where interfaces between brains and body models are missing or only partially characterized, one approach is to train an artificial neural network (ANN) to approximate these interfaces with deep reinforcement learning (DRL)."<br /> The choice of "interface" as a distinct, well-defined neurobiological entity is somewhat confusing and may mislead non-practitioner readers. If neuronal and muscular (and their interactions) physiology are incorporated into a neurobehavioral model, then in principle there is nothing left to call an "interface". It would be clearer to explain that prior neurobehavioral models have often inserted a trainable multilayer feedforward network between sensory and central brain and between the central brain and motor effectors in order to have a substrate for learning, and that this insertion may render the entire biological modeling exercise scientifically pointless, or at a minimum require a set of computational controls.

      "In building virtual animal models, a motor policy is commonly learned by DRL so that the integrated, closed-loop virtual body successfully mimics the detailed kinematics of real animal behavior."<br /> The authors could define "motor policy" in simple terms and give a brief example.

      "Additional realism is added when the motor policy network is constrained by a connectome dataset. However, many biophysical parameters for individual neurons and synapses remain un-measured."<br /> "motor policy network" is confusing; this is referring to the entire network model here, presumably.

      (3) Methods

      "We used the adult hermaphrodite C. elegans nematode connectome dataset [15, 16, 5], including the identities of its 302 neurons and their synapses (Fig. 1A)."<br /> We believe the authors should specify the dataset type, which is a structural, unsigned connectome lacking grounding in physiological function.

      "The policy network was trained in closed loop using PPO as implemented by MIMIC-MJX"<br /> The authors should define "PPO" and "MIMIC-MJX" in simple terms and explain why they were used.

      (4) Discussion

      "Its role in the movement policy could be fulfilled equally well by a randomly connected RNN, akin to reservoir computing [20], since all the learning happens in the black-box ANN motor decoder."<br /> See above - this computational exercise should actually be performed.

      "Looking further ahead, swapping brain and body models of related species may one day yield real insights into how their brains and bodies diverged through evolution. However, far more model development and experimental validation is needed before we can learn anything from such a digital sphinx."<br /> These two sentences about future possible cross-species chimeras feel superfluous and unsubstantiated, and weaken the main argument of the paper about whole-brain emulation.

    3. Reviewer #2 (Public review):

      Summary:

      The authors use DRL to train a C. elegans connectome-based ANN to control stepping in a D. melanogaster body model. The resulting system can walk. This shows that one needs further constraints to derive biologically meaningful results from this approach.

      Strengths:

      The authors perform a very simple experiment with a clear outcome. The interpretation (or lack of interpretation) is a striking cautionary tale.

      Weaknesses:

      There is little analysis of precisely how robust this result is to parameter variation and network wiring. The worm also undulates in an oscillatory fashion. Thus, it is possible that the network is tapping into biologically meaningful motifs to generate oscillations for walking. As well, it would be useful to examine which heuristics one can use to determine whether modeling efforts are sufficiently constrained (i.e., how much biological data will be necessary to start obtaining fruitful, interpretable outcomes from DRL task optimization). For example, their "solution" using the worm connectome is not sparse (i.e., it uses many neurons). Perhaps a signature of a biologically-meaningful, interpretable result is one that is sparse?

    4. Reviewer #3 (Public review):

      Summary:

      The authors construct a computational chimera by attaching a C. elegans connectome to a Drosophila body biomechanical model and use deep reinforcement learning to link neural activity to motor output. The model is able to produce walking, but is considered a priori to be scientifically meaningless, and the work is treated as a cautionary tale in complex interpretation layers unconstrained by experiment or data.

      Strengths:

      In a period of increasing excitement about linking AI and neuroscience, I respect very much that the authors work through a nontrivial example of nonsense results, rather than just making a theoretical case. It offers a clear and memorable existence proof that matching outputs of complex trained networks does not mean the internal dynamics are themselves emulated.

      Weaknesses:

      While I understand that the work was a rapidly produced comment on science-by-press-release, the message seems too important to be treated in quite as pithy a manner as it is. In particular, because the computational experiment is so memorable, it is worth getting the message right to avoid a set of readers who take from it that they should dismiss this category of neuroAI wholesale (which the authors absolutely do not imply!).

      One part of me reads this work and thinks that by intentionally wiring up the sensory feedback in a particularly nonsense way, the authors have just made a bad model, and sometimes bad models can still generate sensible outputs, especially when expressive models are optimized to fit those sensible outputs. But I think this work is trying to say something more specific than this, and I would like it to be a bit clearer about that. The authors do a fairly good job of sharing a view about what should have been done instead, but this message would benefit from having some more concrete suggestions to avoid a simplistic interpretation. A few thoughts:

      (1) It's not entirely obvious to me that the model is "scientifically meaningless." As the authors know extremely well, Drosophila walking is thought to be driven by simple central pattern generators coupled to leg-specific implementations. The C. elegans neural circuit is clearly capable of producing rhythmic activity as well. A version of the model they ran could have identified biologically valid rhythmic activity in the C elegans circuit and mapped it via the DRL to the right locomotor behavior in the fly. While this would not be a good emulation of the fly, it's not a concept devoid of scientific meaning. Similarly, if the ANN is converting a rhythmic signal to coordinated walking, it's not obvious to me that there aren't useful principles to identify in how it achieves this - it's basically the equivalent of that post-CPG circuitry, no?

      (2) Similarly, is this outcome going to be relatively specific to rhythmic behaviors? I suspect that it would be harder to push the C. elegans connectome to produce some behaviors than others - for example, adding in visual navigation and other motor patterns, or a ring attractor. Rhythmic circuits arise in many places, and both biology and dynamical systems tell us they can come from numerous configurations of elements and interactions.

      (3) Aside from the nonsense formulation of the problem, I would have liked to know more about what the authors should have done to know their model was useless. Put another way, if the authors hadn't known that their model was bad from the beginning (e.g., if they had stuck a fly brain in the middle of it, gotten the sensory feedback right), would there have been some way to figure out if it was meaningful or meaningless based on the results of the trained model itself?

    1. eLife Assessment

      This study provides valuable insights into the role of the bile acid receptor TGR5 in regulating bone marrow adipose tissue. This revised manuscript has additional characterization of TGR5 expression in hematopoietic and stromal populations, and phenotypic differences observed between sexes, during aging, and under high-fat diet conditions. Although the scope of the study has expanded, the mechanism of how TGR5 in the microenvironment regulates hematopoietic recovery is still incomplete.

    2. Reviewer #1 (Public review):

      This study by Alonso-Calleja and colleagues aimed to determine whether TGR5 regulates hematopoiesis and the bone marrow microenvironment under steady-state conditions and following transplantation. The revised manuscript substantially improves upon the original submission by providing additional characterization of TGR5 expression in hematopoietic and stromal populations, incorporating analyses in female mice, and expanding the investigation of bone marrow adipose tissue under aging and high-fat diet conditions. These additions more convincingly establish TGR5 as a regulator of bone marrow adipose tissue and stromal composition.

      Major strengths of the study include the comprehensive characterization of the bone marrow adipose tissue phenotype across multiple experimental settings and the demonstration that TGR5 deficiency consistently alters the stromal compartment. The strongest and most convincing aspect of the work is the identification of TGR5 as a regulator of bone marrow adipose tissue and the bone marrow microenvironment. These findings provide useful insights into how metabolic signaling pathways influence the hematopoietic niche.

      However, the evidence supporting a direct role for TGR5 in hematopoietic recovery following transplantation remains limited. Although reciprocal transplantation experiments and peripheral blood recovery analyses strengthen the manuscript, the conclusions regarding hematopoietic regeneration continue to rely largely on correlative observations. The study does not directly demonstrate that expansion of adipocyte progenitors is responsible for the enhanced recovery phenotype, nor does it establish improved regeneration of hematopoietic stem or progenitor cells within the bone marrow. Overall, the revised work addresses many of the concerns raised in the original review and provides useful new insights into the regulation of the bone marrow microenvironment by TGR5. Nevertheless, the conclusions regarding hematopoietic recovery should remain appropriately tempered, as the mechanistic basis linking the stromal phenotype to enhanced regeneration has not been directly demonstrated.

    3. Reviewer #2 (Public review):

      Summary:

      The authors showed the expression of TGR5 in hematopoietic compartments and that loss of TGR5 doesn't impair steady-state hematopoiesis. Notably, TGR5 knockout significantly decreases BMAT, increase the APC population and accelerate the recovery upon bone marrow transplantation.

      Strengths:

      The role of TGR5 is interesting.

      Weaknesses:

      Additional mechanistic studies would further strengthen the work and provide deeper insight into how TGR5 regulates the bone marrow microenvironment.

    4. Author response:

      The following is the authors’ response to the original reviews

      eLife assessment

      This study investigates the role of the bile acid receptor TGR5 in adult hematopoiesis of the mouse model. The findings are potentially useful because the loss of TGR5 leads to dysregulation of bone marrow adipose tissue (BMAT) that has emerging regulatory functions. However, the study is still incomplete because the mechanism of TGR5 is not clear, the stromal cells expressing TGR5 have not been well defined, and there is not strong evidence for the role of TGR5 in recovery from transplant stress.

      We thank the eLife editorial team for handling our manuscript. In our revised version, we took into consideration the suggestions of the reviewers, which we believe have significantly improved the quality of the current study. In summary, our new data provide further evidence that TGR5 is expressed in both hematopoietic cell lineage and stromal cells of the bone marrow (BM), including subpopulation analyses for both. While steady-state hematopoiesis remains intact in TGR5 knockout mice, we demonstrate that loss of TGR5 significantly impacts progenitor reconstitution under stress conditions, alters the BM adipose tissue (BMAT) homeostasis in a sex-specific manner and influences the balance of stromal cell differentiation. In particular, TGR5 deficiency resulted in reduced regulated BMAT and an accumulation of adipocyte progenitors, correlating with improved hematopoietic recovery following BM transplantation. The BMAT decrease is further observed in physiologically and pathophysiologically relevant contexts such as aging and obesity, where TGR5 deficiency is associated with a decrease in the myeloid bias. Collectively, our findings support a previously unrecognized role for TGR5 in maintaining BM niche integrity and highlight its potential as a modulator of hematopoietic support, although the precise molecular mechanisms still need to be elucidated.

      Public Reviews:

      Reviewer #1 (Public Review):

      Summary:

      Alonso-Calleja and colleagues explore the role of TGR5 in adult hematopoiesis at both steady state and post-transplantation. The authors utilize two different mouse models including a TGR5-GFP reporter mouse to analyze the expression of TGR5 in various hematopoietic cell subsets. Using germline Tgr5-/- mice it's reported that loss of Tgr5 has no significant impact on steady-state hematopoiesis, with a small decrease in trabecular bone fraction, associated with a reduction in proximal tibia adipose tissue, and an increase in marrow phenotypic adipocytic precursors. The authors further explored the role of stroma TGR5 expression in the hematopoietic recovery upon bone marrow transplantation of wild-type cells, although the studies supporting this claim are weak. Overall, while most of the hematopoietic phenotypes have negative results or small effects, the role of TGR5 in adipose tissue regulation is interesting to the field.

      Strengths:

      This is the first time the role of TGR5 has been examined in the bone marrow.

      This paper supports further exploration of the role of bile acids in bone marrow transplantation and possible therapeutic strategies.

      We thank the reviewer for pinpointing the strengths of our study.

      Weaknesses:

      (1) The authors fail to describe whether niche stroma cells or adipocyte progenitor cells (APCs) express TGR5.

      Using the TGR5:GFP reporter model, we identified GFP<sup>+</sup> cells in the stroma-enriched CD45-Ter119-CD31- population that contains the adipogenic progenitor cells (APC).

      These data, along with the corresponding gating strategy are outline in Figure 6A and B.

      We found the subpopulation analyses within the stroma gate challenging at the individual mouse level given the limited cell numbers for these progenitor populations. We attempted to circumvent this by concatenating the individual files per genotype (WT and TGR5:GFP) for two independent experiments. We were surprised to find that there is little to no GFP expression in the APC population. Nevertheless, we found that the CD45<sup>-</sup>Ter119<sup>-</sup>CD31<sup>-</sup>CD24<sup>+</sup>Sca1</sup>+</sup>, multi-potent stem cell-like population (Ambrosi et al., 2017), reproducibly showed a TGR5:GFP positivity comparable to that of Ly6<sup>lo</sup> monocytes. We believe this would be compatible with our results showing an increase in CFU-F in Tgr5<sup>-/-</sup> mice, but we remain cautious about its interpretation. Results are shown in Author response image 1 for the concatenated flow plots and numerically in Error! Reference source not found. considering all individual mice analyzed (i.e., non pooled) for completeness. We kindly request the Reviewer’s opinion on whether these subpopulation analyses should be included in the main manuscript or solely here as part of the public review section.

      Author response image 1.

      Flow cytometry gating strategy used to identify stroma subpopulations in the stroma enriched CD45<sup>-</sup>Ter119<sup>-</sup>CD31<sup>-</sup> gate and their GFP signal for BM cells in TGR5:GFP mice. Subpopulations were immunophenotypically defined as CD45<sup>-</sup>Ter119<sup>-</sup>CD31<sup>-</sup>CD24<sup>+</sup>Sca1</sup>+</sup> (adipogenic progenitor cells (APC)) and CD45Ter119<sup>-</sup>CD31<sup>-</sup>CD24<sup>+</sup>Sca1</sup>+</sup> (multi-potent stem cell-like populations) (Ambrosi et al., 2017). Results are shown as the concatenated data of all mice in each phenotype for two independent experiments, as indicated.

      Author response table 1.

      Frequencies in the total live cell gate for adipogenic progenitor cells (APC) and CD45Ter119<sup>-</sup>CD31<sup>-</sup>CD24<sup>+</sup>Sca1</sup>+</sup> (multi-potent stem cell-like populations, MPSC-like) (gating as in Ambrosi et al., 2017) in the experiments presented in Author Response Figure 1, expressed as average +/- 95% confidence interval for the two independent experiments. The number of mice per genotype and per experiment is indicated in parenthesis. **Paired, two-tailed Student’s t-test statistical analysis for the combined experiments (n=8 per group) indicates p < 0.01 for GFP expressing cells in APC versus MPSC-like populations.

      (2) Although the authors note a significant reduction in bone marrow adipose tissue in Tgr5-/- mice, they do not address whether this is white or brown adipose tissue especially since BA-TGR5 signaling has been shown to play a role in beiging.

      The nature of BMAT and how it relates to brown, white or brown/beige adipose tissue has been a persisting question in the field. Our understanding is that BMAT is currently considered as a distinct adipose depot that is neither white nor brown/beige (Sebo et al., 2019; Suchacki et al., 2020). BMAT does not express UCP1 to an appreciable extent, with reports showing that detectable expression possibly stems from contamination by tissues surrounding bone (Craft et al., 2019). Beyond this consideration, as the regulated BMAT in Tgr5<sup>-/-</sup> mice is almost absent, determination of the brown/beige vs white nature of the little regulated BMAT that remains would be technically extremely challenging.

      (3) In Figure 1, the authors explore different progenitor subsets but stop short of describing whether TGR5 is expressed in hematopoietic stem cells (HSCs).

      We have added these data to the manuscript as part of Figure 1C.

      We have further expanded our data in Figure 1D and Figure 1–figure supplement 1A with the expression of TGR5:GFP in megakaryocyte progenitors (Lin<sup>-</sup>cKit<sup>+</sup>Sca1CD150<sup>+</sup>CD41<sup>+</sup>) as shown in Author response image 2.

      Author response image 2.

      A, representative flow cytometry gating strategy used to identify megakaryocyte progenitors (MkProg) and GFP positivity in TGR5:GFP mice and their wild-type controls. B, frequencies of GFP<sup>+</sup> cells in MkProg population in the BM of 8-12-week-old male TGR5:GFP mice and their controls (n=3 for wild-type control mice, n=4 for TGR5:GFP mice). Results represent the mean ± s.e.m., n represents biologically independent replicates. Two-tailed Student’s t-test (B) was used for statistical analysis. p-values (exact value) are indicated.

      Finally, we have completed the characterization of BM progenitor populations to include the erythroid lineage, thus covering the main hematopoietic populations. We have added these data to Figure 1E and Figure 1–figure supplement 1B.

      (4) Are there more CD45+ cells in the BM because hematopoietic cells are proliferating more due to a direct effect of the loss of Tgr5 or is it because there is just more space due to less trabecular bone?

      We observe an average 20% increase in CD45<sup>+</sup> cell counts in baseline Tgr5<sup>-/-</sup> mice. The absolute volume of bone and BMAT lost in these animals does not account for 20% of the medullary cavity volume, so we speculate that the increase in CD45<sup>+</sup> counts is not solely due to increased available volume for these cells.

      (5) In Figure 4 no absolute cell counts are provided to support the increase in immunophenotypic APCs (CD45-Ter119-CD31-Sca1+CD24-) in the stroma of Tgr5-/- mice. Accordingly, the absolute number of total stromal cells and other stroma niche cells such as MSCs, ECs are missing.

      These data are now included in the manuscript and in Author response image 3. Although we detect an increase in the relative frequency of APCs (Figure 7A), on a per-leg basis we did not detect an increase in APC numbers per leg (Author response image 3). We did however observe a decrease in total CD45<sup>-</sup>Ter119sup>-</sup>CD31sup>-</sup> stromal-enriched cells in Tgr5<sup>-/-</sup> mice. Nevertheless, given that only 2–5% of stromal cells are recovered in single-cell suspensions compared with native tissue (Coutu et al., 2017; Gomariz et al., 2018), we consider results for absolute stromal cell numbers prone to over interpretation. Our conclusion, therefore, remains one of relative enrichment of immunophenotypic APCs, supported by in vitro findings of increased adipogenesis and CFU-F formation after plating equal cell numbers.

      Endothelial cells (CD45<sup>-</sup>Ter119<sup>-</sup>CD31<sup>+</sup>) were also quantified, with no differences observed between groups.

      Author response image 3.

      Absolute number of adipocyte progenitor cells (APC), stroma cells (CD45-Ter119CD31<sup>-</sup>) and endothelial cells (CD45<sup>-</sup>Ter119<sup>-</sup>CD31<sup>+</sup>) (n=5 for both Tgr5<sup>+/+</sup> and Tgr5<sup>-/-</sup> mice). Results represent the mean ± s.e.m., n represents biologically independent replicates. Unpaired, two-tailed Student’s t-test was used for statistical analysis. p values (exact values) are shown.

      (6) There are issues with the reciprocal transplantation design in Fig 4. Why did the authors choose such a low dose (250 000) of BM cells to transplant? If the effect is true and relevant, the early recovery would be observed independently of the setup and a more robust engraftment dataset would be observed without having lethality post-transplant. On the same note, it's surprising that the authors report ~70% lethality post-transplant from wild-type control mice (Fig 4E), according to the literature 200 000 BM cells should ensure the survival of the recipient post-TBI. Overall, the results even in such a stringent setup still show minimal differences and the study lacks further in-depth analyses to support the main claim.

      We thank the reviewer for this comment. On the one hand, we respectfully disagree on the relevance of the effect size, as Tgr5<sup>-/-</sup> mice recover from low platelet and leukocyte counts significantly faster than wild-type controls. Tgr5<sup>-/-</sup> recipients recovered their platelet levels faster than the Tgr5<sup>+/+</sup> recipients, showing values consistently above 200.000/µL one week sooner; platelet levels below this threshold are considered a risk factor for bleeding events (Morowski et al., 2013; Vannini et al., 2019). In addition, on day 15 post-irradiation, Tgr5<sup>-/-</sup> recipients had higher neutrophil levels, at over the 500 cells/µL-threshold value for infection risk (Morowski et al., 2013; Vannini et al., 2019), but this difference was no longer statistically significant upon stringent multiple-comparison correction (uncorrected p-value 0.028, multiplecomparison corrected p-value 0.108). Underlining the relevance, in a clinical setting, G-CSF is routinely administered to patients daily to enhance myeloid recovery even if the acceleration of recovery is by 1-2 days (Trivedi et al., 2009).

      From the perspective of mortality, we agree that it is higher than expected and constitutes a limitation of our work. Note that we discovered a mistake in data plotting during the preparation of the revised version of the manuscript which renders the mortality curve no longer statistically significant, with a p value that changed from 0.0254 to 0.0676. We have duly changed our interpretation in the text to that of a trend towards higher survival.

      Regarding the mortality in our experiments being higher than expected, we have unfortunately suffered from cases of “swollen muzzles syndrome” in our facilities that have greatly hampered our ability to perform myeloablation experiments (Garrett et al., 2018), as even sublethal doses have resulted in the appearance of side effects that were reasons for euthanasia under Swiss legislation. For example, a strong reduction in mobility requires immediate euthanasia. All experiments were performed blinded to genotype allocation, so we can reasonably exclude experimenter bias. Finally, it could be argued that mice with more marked symptomatology leading to euthanasia are more likely to have hematopoietic deficits, which in our case was mostly seen for WT animals. We have therefore chosen to report mortality alongside the longitudinal assessment of peripheral blood counts to ensure clarity of the coherent effect and full transparency. We strongly believe that this disclosure, when examined as a limitation in the discussion section, aligns with the open science efforts advocated by eLife. Unfortunately, it is beyond the scope of this manuscript and the authors' timeline to rederive the Tgr5<sup>-/-</sup> colony in our new facility and perform new bone marrow transplantation studies with titrating doses of BM.

      Lastly, the choice of 250,000 BM cells serves as the control dose per recommended standards for competitive repopulation assays (Purton and Scadden, 2007). Quoting from this methodological reference paper: “In our experiments, each recipient mouse receives cell doses … together with 2x10<sup>5</sup> competing congenic bone marrow. We have found these cell doses sufficient to detect both reductions (Purton et al., 2006) and increases (Janzen et al., 2006; Walkley et al., 2005) in HSC numbers in different mutant mice… Furthermore, caution should be used when designing competitive repopulation assays, as it has been shown that the reliability of this assay is critically dependent on the numbers of HSCs present in the populations being assessed: when too few or too many HSCs (recipients of <1 1x10<sup>5</sup> or >2x10<sup>7</sup> bone marrow cells each from donor and competing sources) are present, the data may not be meaningful (Harrison et al., 1993).”

      Moreover, our lethal radiation rescue experiments with 2.5x105 cells are designed to deliver a minimal hematopoietic source known to rescue the great majority of animals in standard conditions, and thus to best mimic the situations when enhancement of hematopoietic recovery would be clinically meaningful (as used by our group members in (Naveiras et al., 2009; Tratwal et al., 2020; Vannini et al., 2019)). In our hands and in the absence of “swollen muzzles syndrome”, this approach leads to 80-95% overall survival, which mirrors the clinical setting of autologous transplantation that our transplants are meant to mimic. In clinical practice strategies, enhancing hematopoietic recovery would be most useful to improve morbimortality in either alternative hematopoietic progenitor allotransplants, known associated with delayed engraftment, or in the case of febrile neutropenia associated to bacteriemia, which affects 11-15% of both hematopoietic autologous or allogeneic transplant patients (Gil et al., 2013; Gooley et al., 2010). In summary, septic neutropenia was unfortunately modelled by the rescue experiments presented in Figure 7 C-J with swollen muzzles syndrome concomitant to the reconstitution, but we strongly believe that this complication enhances the clinical relevance of our data.  

      (7) Mechanistically, how does the loss of Tgr5 impact hematopoietic regeneration following sublethal irradiation?

      As delineated in the previous point, we have been seriously conditioned by cases of “swollen muzzles syndrome” (Garrett et al., 2018), which has stopped us from proceeding with more irradiation experiments for this particular study. Mechanistic studies are unfortunately beyond the scope of this manuscript, but the mechanistic basis for the relationship between BM adipocyte differentiation and hematopoiesis is now a focus for one of the involved laboratories and should produce follow-up manuscripts in the near future.

      (8) Only male mice were used throughout this study. It would be beneficial to know whether female mice show similar results.

      We thank Reviewer #1 for this question, as it led us to perform the characterization of steady-state hematopoiesis and morphological bone and BMAT evaluation of young female mice, yielding new findings that strengthen our manuscript. In summary, we have found that the decrease in BMAT is sexually dimorphic, with young females not showing reduced BMAT levels. Conversely, we did find a trend towards lower trabecular bone content in females. We present these data as part of Figure 3.

      Reviewer #2 (Public Review):

      Summary:

      In this manuscript, the authors examined the role of the bile acid receptor TGR5 in the bone marrow under steady-state and stress hematopoiesis. They initially showed the expression of TGR5 in hematopoietic compartments and that loss of TGR5 doesn't impair steady-state hematopoiesis. They further demonstrated that TGR5 knockout significantly decreases BMAT, increases the APC population, and accelerates the recovery upon bone marrow transplantation.

      Strengths:

      The manuscript is well-structured and well-written.

      We thank Reviewer #2 for this comment.

      Weaknesses:

      The mechanism is not clear, and additional studies need to be performed to support the authors' conclusion.

      We agree with Reviewer #2 that more studies are needed to understand the role of TGR5 in the hematopoietic system. We have been hampered in our studies of stress hematopoiesis because of frequent cases of swollen muzzles syndrome (Garrett et al., 2018), which prevented us from conducting additional experiments involving myelosuppression. Furthermore, the identification of the mechanism that links changes in the adipocyte differentiation axis and hematopoietic support, which is more complex than initially thought, has become a top priority for one of the involved laboratories and should produce follow-up manuscripts in the near future.

      Recommendations For The Authors:

      Reviewer #2 (Recommendations For The Authors):

      (1) Figure 1: the authors showed the presence of TGR5 in hematopoietic cells using a GFP report in mice. What's the expression pattern of TGR5 in the nonhematopoietic cells? For example, adipocytes, stromal cells, endothelial cells, etc. In addition to analyze the percentage of TGR5-GFP+ cells, the authors should also quantify the expression levels of TGR5 in various hematopoietic and niche components.

      We thank Reviewer #2 for this question, which we have addressed as points 1 and 3 from Reviewer #1. The expression of TGR5:GFP signal in stromal cells, HSCs, and the various hematopoietic compartments is now presented respectively as new panels in Figure 6A-B, Figure 1C-D, Figure 1–Figure Supplement 1A-B, and Figure 1figure supplement 2A, as well as Author response image 1 for stromal progenitors. Please note that specifically for the stromal compartment, we kindly requested Reviewer #1’s opinion on whether these subpopulation analyses should be included in the main manuscript or solely here as part of the public review section, as we are concerned about over interpretation for these rare APC and multi-potent stem cell-like subpopulations.

      Consistent with the typical low-abundance expression of G protein-coupled receptors, where ligand-mediated activation is more relevant than absolute receptor abundance, TGR5 is also lowly expressed in most cell types. Although technical limitations prevent us from directly quantifying TGR5 expression in adipocytes (due to the difficulty of isolating these populations from bone marrow), previous studies have reported TGR5 expression in human BMSC-derived adipocytes (Velazquez-Villegas et al., 2018) . For similar reasons, we could not quantify TGR5 expression levels by RT-qPCR in the bone marrow niche, as isolating enough cells from each compartment to reliably detect a low-abundance receptor is technically challenging.

      Regarding endothelial cells, previous studies have shown TGR5 expression in vascular endothelial cells (Kida et al., 2013) as well. In our dataset the number of endothelial cells recovered was limited, largely due to the lack of a dedicated endothelial isolation protocol for the BM. Even so, the data we collected are shown as Author response image 4 and Author response table 2 and suggest that TGR5:GFP level in endothelial cells is lower than in the stromal and hematopoietic compartments. As for Author response image 1 and Author response table 1, we remain cautious about the interpretation of GFP expression in these low frequency stromal and endothelial populations and we kindly request the Reviewer’s opinion on whether these subpopulation analyses should be included in the main manuscript or solely here as part of the public review section.

      Author response image 4.

      Representative flow cytometry gating strategy used to identify endothelial cells in BM, defined as CD45<sup>-</sup>Ter119<sup>-</sup>CD31<sup>+</sup> and their GFP positivity in TGR5:GFP mice.

      Author response table 2.

      Frequency of GFP-expressing cells in the CD45<sup>-</sup>Ter119<sup>-</sup>CD31<sup>+</sup> gate for endothelial cells presented in Author Response Figure 4, expressed as average (standard deviation). The number of mice per genotype and per experiment is indicated in parenthesis.

      (2) Figure 2: Regarding the competitive transplantation, the authors should bleed the mice every 4 weeks and show the dynamics of donor chimerism up to 16 or 20 weeks. The difference at the 3-week time point is very tiny and is this change significant? There is no statistical significance shown in Supplemental Figure 2 Panel C at a 3-week time point.

      The full data for repetitive monthly bleedings in primary, secondary and tertiary transplants is now shown in Figure 2J and Figure 2-figure supplement 1I-K. The statistically significant difference on the first month, which represents a 26% loss in short-time progenitor phenotype, is small but potentially clinically relevant as it is these progenitors that sustain early hematopoietic recovery from severe leucopenia and thrombocytopenia. Indeed, the hematopoietic phenotype described in Figure 7 C-J for accelerated rescue of Tgr5<sup>-/-</sup> recipients with wild-type bone marrow would be coherently associated to this short-term progenitor phenotype.  

      (3) Figure 3: in addition to the irradiation/transplantation model, the authors should also use alternative models for stress hematopoietic, for example, 5-FU, etc.

      This is an excellent suggestion, but unfortunately beyond the scope of this manuscript due to the move of one of the co-senior authors to another institution. Follow-up studies are however planned with ablative chemotherapy models relevant to the hematopoietic transplant setting (non 5-FU).

      (4) To get a better understanding of the role of TGR5 in the bone marrow niche, the authors should get and/or generate the TGR5 floxed mice and this will allow the authors to delete it from specific cell populations using cell-type specific Cre. In addition, it would be very interesting to investigate further the molecular mechanisms downstream of TGR5.

      We much appreciate this comment and the interest it reflects in TGR5 BM biology. Mechanistic studies are unfortunately beyond the scope of this manuscript, but the mechanistic basis for the relationship between BM adipocyte differentiation and hematopoiesis are now a focus for one of the involved laboratories and should produce concrete follow-up manuscripts soon.  

      References

      Ambrosi, T.H., Scialdone, A., Graja, A., Gohlke, S., Jank, A.-M., Bocian, C., Woelk, L., Fan, H., Logan, D.W., Schürmann, A., Saraiva, L.R., Schulz, T.J., 2017. Adipocyte Accumulation in the Bone Marrow during Obesity and Aging Impairs Stem Cell-Based Hematopoietic and Bone Regeneration. Cell Stem Cell 20, 771-784.e6. https://doi.org/10.1016/j.stem.2017.02.009

      Coutu, D.L., Kokkaliaris, K.D., Kunz, L., Schroeder, T., 2017. Three-dimensional map of nonhematopoietic bone and bone-marrow cells and molecules. Nat Biotechnol 35, 1202–1210. https://doi.org/10.1038/nbt.4006

      Craft, C.S., Robles, H., Lorenz, M.R., Hilker, E.D., Magee, K.L., Andersen, T.L., Cawthorn, W.P., MacDougald, O.A., Harris, C.A., Scheller, E.L., 2019. Bone marrow adipose tissue does not express UCP1 during development or adrenergic-induced remodeling. Sci Rep 9, 17427. https://doi.org/10.1038/s41598-019-54036-x

      Garrett, J., Sampson, C.H., Plett, P.A., Crisler, R., Parker, J., Venezia, R., Chua, H.L., Hickman, D.L., Booth, C., MacVittie, T., Orschella, C.M., Dynlachta, J.R., 2018. Characterization and Etiology of Swollen Muzzles in Irradiated Mice. Radiation Research 191, 31. https://doi.org/10.1667/RR14724.1

      Gil, L., Poplawski, D., Mol, A., Nowicki, A., Schneider, A., Komarnicki, M., 2013. Neutropenic enterocolitis after high-dose chemotherapy and autologous stem cell transplantation: incidence, risk factors, and outcome. Transpl Infect Dis 15, 1–7. https://doi.org/10.1111/j.1399-3062.2012.00777.x

      Gomariz, A., Helbling, P.M., Isringhausen, S., Suessbier, U., Becker, A., Boss, A., Nagasawa, T., Paul, G., Goksel, O., Székely, G., Stoma, S., Nørrelykke, S.F., Manz, M.G., Nombela-Arrieta, C., 2018. Quantitative spatial analysis of haematopoiesis-regulating stromal cells in the bone marrow microenvironment by 3D microscopy. Nat Commun 9, 2532. https://doi.org/10.1038/s41467-018-04770-z

      Gooley, T.A., Chien, J.W., Pergam, S.A., Hingorani, S., Sorror, M.L., Boeckh, M., Martin, P.J., Sandmaier, B.M., Marr, K.A., Appelbaum, F.R., Storb, R., McDonald, G.B., 2010. Reduced mortality after allogeneic hematopoietic-cell transplantation. N Engl J Med 363, 2091–2101. https://doi.org/10.1056/NEJMoa1004383

      Harrison, D.E., Jordan, C.T., Zhong, R.K., Astle, C.M., 1993. Primitive hemopoietic stem cells: direct assay of most productive populations by competitive repopulation with simple binomial, correlation and covariance calculations. Exp Hematol 21, 206–219.

      Janzen, V., Forkert, R., Fleming, H.E., Saito, Y., Waring, M.T., Dombkowski, D.M., Cheng, T., DePinho, R.A., Sharpless, N.E., Scadden, D.T., 2006. Stem-cell ageing modified by the cyclin-dependent kinase inhibitor p16INK4a. Nature 443, 421–426. https://doi.org/10.1038/nature05159

      Kida, T., Tsubosaka, Y., Hori, M., Ozaki, H., Murata, T., 2013. Bile acid receptor TGR5 agonism induces NO production and reduces monocyte adhesion in vascular endothelial cells. Arterioscler Thromb Vasc Biol 33, 1663–1669. https://doi.org/10.1161/ATVBAHA.113.301565

      Morowski, M., Vögtle, T., Kraft, P., Kleinschnitz, C., Stoll, G., Nieswandt, B., 2013. Only severe thrombocytopenia results in bleeding and defective thrombus formation in mice. Blood 121, 4938–4947. https://doi.org/10.1182/blood-2012-10-461459

      Naveiras, O., Nardi, V., Wenzel, P.L., Hauschka, P.V., Fahey, F., Daley, G.Q., 2009. Bonemarrow adipocytes as negative regulators of the haematopoietic microenvironment. Nature 460, 259–263. https://doi.org/10.1038/nature08099

      Purton, L.E., Dworkin, S., Olsen, G.H., Walkley, C.R., Fabb, S.A., Collins, S.J., Chambon, P., 2006. RARgamma is critical for maintaining a balance between hematopoietic stem cell self-renewal and differentiation. J Exp Med 203, 1283–1293. https://doi.org/10.1084/jem.20052105

      Purton, L.E., Scadden, D.T., 2007. Limiting factors in murine hematopoietic stem cell assays. Cell Stem Cell 1, 263–270. https://doi.org/10.1016/j.stem.2007.08.016

      Sebo, Z.L., Rendina-Ruedy, E., Ables, G.P., Lindskog, D.M., Rodeheffer, M.S., Fazeli, P.K., Horowitz, M.C., 2019. Bone Marrow Adiposity: Basic and Clinical Implications. Endocr Rev 40, 1187–1206. https://doi.org/10.1210/er.2018-00138

      Suchacki, K.J., Tavares, A.A.S., Mattiucci, D., Scheller, E.L., Papanastasiou, G., Gray, C., Sinton, M.C., Ramage, L.E., McDougald, W.A., Lovdel, A., Sulston, R.J., Thomas, B.J., Nicholson, B.M., Drake, A.J., Alcaide-Corral, C.J., Said, D., Poloni, A., Cinti, S., Macpherson, G.J., Dweck, M.R., Andrews, J.P.M., Williams, M.C., Wallace, R.J., Van Beek, E.J.R., MacDougald, O.A., Morton, N.M., Stimson, R.H., Cawthorn, W.P., 2020. Bone marrow adipose tissue is a unique adipose subtype with distinct roles in glucose homeostasis. Nat Commun 11, 3097. https://doi.org/10.1038/s41467-020-16878-2

      Tratwal, J., Bekri, D., Boussema, C., Sarkis, R., Kunz, N., Koliqi, T., Rojas-Sutterlin, S., Schyrr, F., Tavakol, D.N., Campos, V., Scheller, E.L., Sarro, R., Bárcena, C., Bisig, B., Nardi, V., de Leval, L., Burri, O., Naveiras, O., 2020. MarrowQuant Across Aging and Aplasia: A Digital Pathology Workflow for Quantification of Bone Marrow Compartments in Histological Sections. Front Endocrinol (Lausanne) 11, 480. https://doi.org/10.3389/fendo.2020.00480

      Trivedi, M., Martinez, S., Corringham, S., Medley, K., Ball, E.D., 2009. Optimal use of G-CSF administration after hematopoietic SCT. Bone Marrow Transplant 43, 895–908. https://doi.org/10.1038/bmt.2009.75

      Vannini, N., Campos, V., Girotra, M., Trachsel, V., Rojas-Sutterlin, S., Tratwal, J., Ragusa, S., Stefanidis, E., Ryu, D., Rainer, P.Y., Nikitin, G., Giger, S., Li, T.Y., Semilietof, A., Oggier, A., Yersin, Y., Tauzin, L., Pirinen, E., Cheng, W.-C., Ratajczak, J., Canto, C., Ehrbar, M., Sizzano, F., Petrova, T.V., Vanhecke, D., Zhang, L., Romero, P., Nahimana, A., Cherix, S., Duchosal, M.A., Ho, P.-C., Deplancke, B., Coukos, G., Auwerx, J., Lutolf, M.P., Naveiras, O., 2019. The NAD-Booster Nicotinamide Riboside Potently Stimulates Hematopoiesis through Increased Mitochondrial Clearance. Cell Stem Cell 24, 405-418.e7. https://doi.org/10.1016/j.stem.2019.02.012

      Velazquez-Villegas, L.A., Perino, A., Lemos, V., Zietak, M., Nomura, M., Pols, T.W.H., Schoonjans, K., 2018. TGR5 signalling promotes mitochondrial fission and beige remodelling of white adipose tissue. Nat Commun 9, 245. https://doi.org/10.1038/s41467-017-02068-0

      Walkley, C.R., Fero, M.L., Chien, W.-M., Purton, L.E., McArthur, G.A., 2005. Negative cellcycle regulators cooperatively control self-renewal and differentiation of haematopoietic stem cells. Nat Cell Biol 7, 172–178. https://doi.org/10.1038/ncb1214

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public reviews):

      Weaknesses:

      Reliance on self-reports

      A primary limitation of this study, acknowledged by the authors, is its reliance on self-reports of participants’ emotional states. Although considerable effort was made to minimize expectation effects, further research is needed to confirm that the observed behavioral changes reflect genuine alterations in emotional states. Additionally, the generalizability of the findings to long-term remediation strategies remains an open question.

      We agree with this characterisation and have strengthened the corresponding acknowledgment in the Discussion. We would also note that, while self-report measures are inherently subjective, the regularities governing subjective emotional experience are of primary scientific interest in their own right, and no consensus definition of a ”genuine” emotional state that supersedes self-report currently exists. We have added a sentence to the Discussion making this point explicit: ”While emotional self-reports are inherently subjective, we note that the regularities governing subjective emotional experience are of primary scientific interest in their own right, and no consensus definition of a ‘genuine’ emotional state that supersedes self-report currently exists.”

      Additionally, we agree that what we have described is limited to a short-term intervention and change. Whether these changes bear on longer-term changes remains to be assessed. Furthermore, the mechanisms or processes that would support such a maintenance are of substantial interest, and will be the focus of future work.

      Statistical analysis and interpretation of the dynamics matrix

      Second, the statistical analysis, particularly the computational approach, sometimes lacks sufficient detail and refinement. While I will not elaborate on specific points here, one notable issue is the interpretation of the intrinsic matrix (A). The model-free analysis reveals correlations between emotions at a given time or within an emotional state across time points. However, it does not provide evidence to support lagged interactions across states that would justify non-diagonal elements in A. The other result concerning the dynamics matrix only highlights a trend in the dominant eigenvalue, which is difficult to interpret in isolation. The absence of a statistically significant group x intervention interaction furthermore makes this finding a little compelling. This weakens the study’s conclusions about the importance of intrinsic dynamics, as claimed in the title.

      We thank the reviewer for raising this important methodological point. We address it in three parts.

      (i) Justification for the full dynamics matrix. In response to this comment, we have added a diagonal-constrained variant of A to the model comparison. The full A model was selected by BIC over the diagonal-constrained version, meaning that cross-emotion lagged interactions contributed to model fit over and above what could be explained by individual emotion autocorrelation alone, thereby justifying the inclusion of off-diagonal elements. We have updated the model comparison results section and figure caption accordingly.

      (ii) Interpretation of the dominant eigenvalue. We agree that the dominant eigenvalue alone is difficult to interpret. Our primary claim regarding intrinsic dynamics rests on model comparison (which identified a change in A as necessary to explain post-intervention data in the distancing group) together with the change in the direction of the dominant eigenvector (tested via Hotelling T<sup>2</sup>, p = 0.019). The eigenvalue magnitude is reported as a complementary, interpretable summary of the stability shift.

      (iii) General analysis of the dynamics matrix. We have added a visualization of the full dynamics matrix A to the Supplementary Materials, reporting per-emotion diagonal elements (persistence) and their variability across participants. This shows that the emotion calm exhibited the highest persistence, consistent with our hypothesis, whereas sadness showed lower-than-expected stability, and other emotions exhibited low mean persistence with substantial individual variability. We have added the following sentence to the Results: ”A general analysis of the dynamics matrix A revealed that while calm exhibited the highest persistence, consistent with our hypothesis (0.59 ± 0.85), sadness did not show the expected stability (0.09 ± 0.58), and other emotions demonstrated low persistence (0.13 to 0.24) alongside substantial individual variability.”

      Terminology (controllability, stability, sensitivity, Gramian)

      Finally, to avoid potential misunderstandings of their work, the authors should be more careful about their use of terms pertaining to the control theory and take the time to properly define them. For example, the ”controllability” of emotional states can either denote that those states are more changeable (control theory definition), or, conversely, more tightly regulated (common interpretation, as used in the abstract). This is true for numerous terms (stability, sensitivity, Gramian, etc.) for which no clear definition nor references are provided. Readers unfamiliar with the framework of control theory will likely be at a loss without more guidance.

      This is an excellent and important point. We have made the following changes throughout the manuscript.

      (i) Terminology table. We have added a new Table 1 in the Methods (Conceptual Definitions) providing a side-by-side mapping of the formal control-theory definition and the psychological interpretation for the three key terms: Controllability, Stability, and Sensitivity.

      (ii) Abstract. The abstract previously used ”controllability” in a way that could be read in the psychological sense. We have replaced the sentence referencing the Controllability Gramian with a clarified version describing the measure as ”continuous measures derived from the controllability matrix, capturing relative differences in how emotional states can be driven by external inputs.”

      (iii) Introduction. We have added a paragraph explicitly distinguishing the control-theoretic definition of controllability from the psychological concept of emotion control or emotion regulation, noting that high controllability in the formal sense does not imply tight regulation or reduced variability. Definitions of sensitivity and stability are also provided.

      (iv) Discussion. Occurrences of ”controllability” that could be ambiguous are annotated with brief clarifications of which sense is intended.

      (v) Gramian. We now use the ”controllability matrix” (C) throughout, and its SVD-based analysis is described using ”left singular vectors” rather than ”eigenvectors of the Gramian”.

      Reviewer #2 (Public reviews):

      Online recruitment, selection effects, and remuneration

      Acquiring data online inevitably gives rise to selection and self-selection effects. This needs to be acknowledged clearly. Exacerbating this, participant remuneration seems low at an amount below the minimum or living wage in Western countries (do the authors know where their participants came from?).

      We thank the reviewer for raising this. All participants were recruited from the UK via Prolific Academic and reimbursed at £7.50/h. Remuneration rates were comparable to other experimental settings, in keeping with other online studies, UK living wage recommendations, and ultimately determined according to institutional ethical guidance. We acknowledge that online recruitment via Prolific may introduce self-selection effects and that the sample may not be representative of the general population. We have added a corresponding statement to the Limitations section: ”Online recruitment via Prolific Academic may introduce self-selection effects, and the sample may not be representative of the general population.”

      Intervention ongoing during the second block

      Another concern is that the intervention does not simply take place before the second block begins but is ongoing during the whole of the second block in that it is integrated into the phrasing of the task on each trial. It is therefore somewhat misleading to speak of a period ’after the intervention’, and it would have been interesting to assess the effect of this by including a third group where the phrasing does not change, but the floating leaves intervention takes place.

      This is a valid and important observation. In the distancing group, the trial-by-trial question phrasing during the second block included a reminder, meaning the intervention was reinforced on every trial. We have acknowledged this as a procedural difference and potential confound in the Limitations section, noting that this reminder may have encouraged a form of retrospective reappraisal rather than in-the-moment distancing. The design choice was intentional (the reminder was included to encourage continued application of the strategy, mirroring how such techniques are deployed in practice) but we agree it is an imperfect feature of the current design and have discussed it openly. We have added to Limitations: ”Relatedly, a procedural difference existed in that only the distancing group received a strategy reminder during the second video block. While intended to facilitate real-time regulation, this prompt may have inadvertently encouraged retrospective reappraisal to align with the distancing narrative.”

      Observation noise

      As mentioned in the Limitations section, observation noise was assumed and not estimated. While this is understandable in this case, the effect of this assumption could have been assessed by simulation with varying levels of observation (and process) noise.

      We would like to clarify that both observation noise (Γ) and process noise (Σ) were in fact estimated from the data, constrained to be diagonal. We have extended the parameter recovery analysis (Supplementary Section: Parameter Recovery) in which, for each of 104 subjects and both time periods (N = 208 observations), we generated 100 surrogate trajectories from the fitted model and re-estimated all parameters. Recovery quality is quantified via Pearson correlations between true and recovered parameters (A, C, Σ, and Γ), alignment of dominant eigenvectors and left singular vectors, and bias analysis for dominant eigenvalue and singular value magnitudes. The analysis confirms that A and C matrix parameters and the critical eigenvector-based metrics all exceed our target threshold of r = 0.7. Noise covariances are poorly recovered, consistent with typical Kalman filter behaviour on short time series. We have expanded the Limitations to note that the Gaussian observation noise assumption does not fully capture the bounded nature of 0–100 rating scales, and suggest that future work could use a truncated or censored observation model.

      Reliance on formal model comparison

      Relatedly, the reliance on formal model comparison is unfortunate since the outcome of such comparisons is easily influenced by slight changes to assumptions such as noise levels. An alternative approach would have been to develop a favoured model based on its suitability to address the research question and its ability, established by simulation, to distill relevant changes of behaviour into reliable parameter estimates.

      We appreciate this methodological concern but would argue that formal model comparison is wellsuited to our research question. Our central aim is not simply to fit emotion trajectories, but to determine which components of the dynamical system (intrinsic dynamics (A), input weights (C), or both) are altered by the distancing intervention. This is inherently a model comparison question: without comparing models that do and do not allow each component to change, no principled inference about the mechanism of action is possible. A single favoured model, however well-motivated, would presuppose the answer.

      We also note that the reviewers concern, that outcomes are sensitive to underlying assumptions, applies equally to the favoured-model approach: simulations rely on predefined structures and noise specifications that shape parameter recovery, and a misspecified favoured model risks confounding the parameters intended to capture the intervention effect with artefacts of model structure. By contrast, our approach evaluates a principled, nested set of models of increasing complexity, using BIC to penalise unnecessary parameters, which guards against overfitting. The models are intentionally simple (linear, Gaussian, time-invariant) to limit the degrees of freedom available to absorb noise.

      Critically, the approach is not model comparison alone. We followed established best-practice procedures for computational modelling, including posterior predictive checks (simulated trajectories closely matched observed data; Fig. 4C), and parameter recovery establishing that the key metrics are reliably recovered (see Supplementary Materials F Parameter Recovery). We have also added a Random-Effects Bayesian Model Selection (RFX-BMS) to characterise individual heterogeneity in model preferences. Together, these provide converging evidence for the validity of our inference. We have clarified this reasoning in the revised manuscript.

      Statistical limitations; Bayesian inference

      The statistical analyses clearly show the limitations of classical statistical testing with highly complex models of the kind the authors (commendably) use. Hunting for statistically significant interactions in a multivariate repeated-measures design relying on inputs from time series- derived point estimates is a difficult proposition. While the authors make the best of the bad 3 situation they create by using null-hypothesis significance testing, a more promising approach would have been to estimate parameters using a sampler like Stan or PyMC and then draw conclusions based on posterior predictive simulations.

      We agree that fully Bayesian parameter estimation via Stan or PyMC would be a valuable methodological advance. Implementing this for 104 subjects across two time periods with a 5-dimensional Kalman filter is, however, a substantial undertaking beyond the scope of the current revision. In the interim, the RFX-BMS analysis added in response to Comment 4 above provides a group-level Bayesian perspective on model uncertainty that partially addresses this concern.

      Reviewer #3 (Public reviews):

      Dual meanings of controllability

      An interesting but perhaps at present slightly confusing aspect of their described results relates to the ’controllability’ of emotions, which they define as their susceptibility to external inputs. Readers should note this definition is (as I understand it) quite distinct from, and sometimes even orthogonal to, concepts of emotional control in the emotion literature, which refer to intentional control of emotions (by emotion regulation strategies such as distancing). The authors also use this second meaning in the discussion. Because of the centrality of control/controllability (in both meanings) to this paper, at present it is key for readers to bear these dual meanings in mind for juxtaposed results that distancing ”reduces controllability” while causing ”enhanced emotional control”

      We are grateful for this observation, which echoes Reviewer 1’s concern about terminology. We have addressed this comprehensively; please see our response to Reviewer 1 Comment 3 above.

      Strategy reminder and possible reappraisal

      As above the authors use an active control – a relaxation intervention – which is extremely closely matched with their active intervention (and a major strength). However, there was an additional difference between the groups (as I currently understand it): ”in the group allocated to the distancing intervention, the phrasing of the question about their feelings in the second video block reminded participants about the intervention, stating: ”You observed your emotions and let them pass like the leaves floating by on the stream.” I do wonder if the effects of distancing also have been partially driven by some degree of reappraisal (considered a separate emotion regulation strategy) since this reminder might have evoked retrospective changes in ratings.

      This is a well-founded concern. As noted in our response to Reviewer 2 Comment 2, we have added an explicit acknowledgment of this procedural difference and the potential for retrospective reappraisal to the Limitations section. We note, as the reviewer themselves observe in the Strengths section, that demand effects are unlikely to account for the specific pattern of dynamic changes observed: uniform demand effects would be expected to produce flat reductions across emotions, whereas our findings show emotion-specific changes in eigenmode structure and controllability direction. Nevertheless, a partial contribution of reappraisal cannot be ruled out from the current design.

      Mechanism of distancing effects (eye movement and oculomotor avoidance)

      Not necessarily a weakness, but an unanswered question is exactly how distancing is producing these effects. As the authors point out, there is a possibility that eye-movement avoidance of the more emotionally salient aspects of scenes could be changing participants’ exposure to the emotions somewhat. Not discussed by the authors, but possibly relevant, is the literature on differences between emotion types on oculomotor avoidance, which could have contributed to differential effects on different emotions.

      We thank the reviewer for raising the oculomotor avoidance hypothesis. Research suggests that different emotions elicit distinct patterns of gaze behaviour: disgust is associated with visual avoidance, whereas anxiety and other negative emotions show increased attentional bias following fear conditioning. These emotion-specific oculomotor patterns could have contributed to the differential effects we observe on the input weight matrix C. What would be particularly interesting to examine in future work is whether a distancing intervention induces multiple, emotionally-specific gaze behaviours, or a single undifferentiated avoidance response. We have expanded the Limitations to: ”[...] The literature on emotion-specific oculomotor avoidance suggests that gaze patterns differ across emotion categories, which could contribute to differential effects on the input weight matrix for specific emotions. [...]”

      Recommendations for the Authors:

      Reviewer #1 (Recommendations for the Authors):

      (1) In the procedure description, the authors suggest that some emotions (e.g. disgust) would be more volatile and stimulus driven, while others (eg. sad) would be more stable. Is this hypothesis reflected in the model-based state dynamics, typically in the diagonal elements of A?

      Yes, we added a supplementary figure showing the dynamics matrix A, which confirms that calm exhibited the highest persistence, consistent with our hypothesis. Disgust showed lower persistence, also in line with this expectation. However, sadness did not display the expected stability and instead showed relatively low persistence, contrary to our hypothesis.

      (2) Could the authors elaborate on the emotional space covered by the chosen ratings? If the axes are positive-negative and slow-fast, why 5 and not 4?

      We added a sentence in the Methods clarifying that five emotions were selected based on the specific affective qualities of the video stimuli provided in the validated databaseto and to better capture the high-dimensional, nuanced states elicited by the stimuli rather than to fit a traditional four-axis model.

      (3) The whole methods section crucially lacks references. As an example, the whole derivation of the most important metrics (eigenvalues of the Gramian, energy ellipse, etc) leaves the reader completely on its own.

      References have been added throughout the Methods, including for the controllability matrix, SVDbased analysis, and eigendecomposition.

      (4) Before equation 1, when introducing x and u: a) time appears twice (typo), and b) the 1Tˆ notation is not standard (especially without bold) and unclear until way below when the one-hot encoding is mentioned.

      The typo has been corrected.

      (5) Why use one hot-encoding rather than the original ratings from the video database? Videos must vary if not in spread (as suggested in Figure B.1) at least in intensity. Ignoring this variance surely diminishes the accuracy of the modelling.

      We used one-hot encoding so that the input weight matrix C can directly estimate participantspecific intensity and sensitivity, rather than fixing input magnitudes to database averages. Using database ratings would assume that the emotional intensity of each video clip generalizes perfectly to our sample; any mismatch would be absorbed as error in C, potentially biasing parameter estimates and obscuring individual differences in emotional reactivity, which are central to our analysis.

      (6) Could the authors develop the rationale behind the bias in Equation 1?

      We added a sentence clarifying that the bias term h captures the steady-state baseline of the emotional system, i.e. the mean rating toward which emotions converge in the absence of external inputs.

      (7) While I can understand why the authors included a set of models with a diagonal C matrix, I do not see why they did not do the same with A. While the diagonal elements are necessary to persist emotional states and induce some autocorrelation in the ratings, as observed empirically, the influence of the non-diagonal elements is not justified (and Figure 5G seems to confirm that). This is critical as, in the end, the controllability metrics will highly depend on those non-diagonal elements which remain very obscure throughout the manuscript.

      We now included both diagonal and full variants of A in the model comparison; still the full A was selected by BIC.

      (8) Concerning the model comparison, I am not sure what the authors mean by using the BIC at the group level. Did they just sum them across participants? This approach is known to be highly susceptible to outliers and cannot be relied upon in general. So-called ”random effect analysis” tends to be regarded as the gold standard and can be easily implemented by taking - 0.5 * BIC as an approximation to the model evidence. Such an approach would also allow to properly test for group differences (cf. Rigoux et al. 2014).

      We have added a Random-Effects Bayesian Model Selection analysis as a new Supplementary Section; see Comment 4 of the public review response above.

      (9) The sentence ”proportion of the total amount of predictive power provided by the full set of models contained in the model being assessed” does not make any sense to me. Please rephrase.

      The sentence has been rephrased: ”Cumulative model weights (w<sub>j</sub>) normalize raw BIC scores so they can be interpreted as the relative probability that a specific model is the best one among the set being compared:”

      (10) ”the largest eigenvalue of the dynamics matrix A identifies the most stable combination of emotions” is only true if the eigenvalues are below 1, which is not granted.

      We have added the qualifier that this holds provided the dominant eigenvalue lies within the unit circle (|λ| < 1), indicating a system that converges to a steady state.

      (11) Equation 3 does not define the Gramian but the controllability matrix, a confusion that goes through the manuscript. The Gramian is formally defined as W = P(AkBB′A′k). Luckily, for discrete systems, it can be approximated by CC′ and therefore the singular values of C can be used to approximate the eigenvalues of W, which are the usual metrics used to define the energy ellipse and so on. While the results reported in the manuscript are correct (up the the approximation), the general description is wrong or misleading.

      We thank the reviewer for this important correction; we now consistently refer to C as the ”controllability matrix” throughout, and its SVD-based analysis is described in terms of left singular vectors rather than eigenvectors of the Gramian.

      (12) Figure 2: what is the matrix V? If it’s from the singular value decomposition W = USV, then (if I am not mistaken) the direction of the ellipsoid is defined by U. Again, a reference would help.

      We have corrected the figure and caption to refer to ”left singular vectors” throughout. We retain the variable name V rather than adopting the standard SVD convention of U to avoid confusion with the input vector u, which appears throughout the model equations.

      <(13) Correction for multiple comparisons is mentioned as a way to correct for the number of conducted tests. However, later on, some post-hoc analyses are reported with the mention that the correction is done across emotions (so p/5), while multiple tests are run for each emotion. This is critical when all the pairs across the cells of an ANOVA are tested and no correction seems to be applied, which is inducing a huge risk of false positives.

      Along the same line: the correct way to demonstrate the effect of the intervention is to first do an ANOVA to reveal an interaction between group and time, and then only to do post-hoc tests to pinpoint where the interaction is coming from, and not the other way around as reported in the manuscript. Further, a difference in significance is not equivalent to a significant difference, and showing that a time effect is significant in one group but not in the other does not imply that the intervention differs between groups, only testing the interaction can confirm this.

      We have added reporting of the significant group × time interaction effects (F(5,208) = 2.6, p = 0.026 for mean ratings; F(5,200) = 2.5, p = 0.03 for the most controllable direction) and flagged these in the figure captions. In the interest of transparency we have left the structure of the results section intact rather than retrospectively reframing it.

      (14) The notation DV = b0 + b1IV*b2G is confusing as a full model (interaction + main effects) should have 3 parameters in addition to the intercept. Also, why use different models for testing the main effect and the interaction?

      The regression equation has been corrected to DV = β<sub>0</sub> + β<sub>1</sub>IV + β<sub>2</sub>G + β<sub>3</sub>(IV × G) + ϵ, making the interaction term explicit.

      (15) Figure 4: the control subject in panel C seems to rate close to 0 in all emotions except for ”calm”. How was this subject fitted? Does model selection (at the subject level) correctly identify a change of dynamics in this case? I don’t see how the behaviour after the intervention could be realistically fitted with a 65-parameters dynamical system. What type of checks were operated to ensure the quality of the fit beyond the recovery analysis (see below)?

      The top participant in Figure 4C was from the distancing group and the participant rating close to zero on most emotions except calm shown at the bottom was from the control group. For the distancing participant, model selection correctly identified a change in dynamics and input weights (BIC = 4806) over the same-parameters model (BIC = 4821). For the control participant, model selection similarly favoured a change in dynamics and input weights (BIC = 3723) over the sameparameters model (BIC = 4011), suggesting that the relaxation intervention produced a comparable effect on emotional dynamics to the distancing intervention. Note that these two participants were selected randomly to illustrate the visual quality of model fit (i.e. that simulated trajectories closely resemble the empirical rating curves) and are not intended to be representative examples of group differences.

      Regarding the data-to-parameter ratio: the 65 parameters are estimated from 55 observations per emotion per block, giving a more favourable ratio than it might appear.

      Beyond the visual trajectory overlays in Figure 4C, we have now added R<sup>2</sup> and peak cross-correlation as a quantitative measure of individual fit quality. Mean R<sup>2</sup> across all subjects and emotions was 0.6 and mean temporal correlation was r = 0.74−0.80, confirming that both the timing and magnitude of emotional responses were well reproduced. Notably, for the specific control participant shown in Figure 4C, R<sup>2</sup> values were 0.74, 0.83, 0.81, 0.83, and 0.83 for disgusted, amused, calm, anxious, and sad respectively, confirming that even for this visually striking participant the model fit was adequate across all five emotion dimensions.

      (16) What do the authors mean by ”eigenmodes”? In the following sentence, what does ”This component” refer to?

      Eigenmodes is defined in the Stability section as the independently evolving combinations of emotions obtained by projecting the state vector onto the eigenvectors of A, and ”This component” has now an explicit referent.

      (17) Figure 5: see above the comment about the necessity for testing the interactions, which should also be reported in the figures.

      Interaction effects are now included in the relevant figure caption (Figure 5).

      (18) When looking for the relationships between questionnaires and controllability, looking only at the most controllable direction seems rather inefficient due to the multiple comparisons correction. Why not compute the angle (or other measure of similarity) with an ideal ”calm” unit vector?

      Along the same line, it’s unlikely that the most controllable direction will smoothly rotate as a function of symptoms. More likely, the winning (highest eigenvalue) direction will switch from one to another, creating some discontinuity in the summary statistic used for the correlation with clinical scores. How could one work around this issue?

      This is an interesting suggestion; we have not implemented it in the current revision, but we acknowledge it as a promising analysis for future work.

      (19) More generally, it would be interesting to see if there are some regularities in the dynamics across participants. If this is the case, one could construct a canonical emotional dynamics and project all participants on this eigenspace. Emotional trajectories would then differ only in their controllability (eigenvalues) in this common space, making a comparison across participants more straightforward. Could the authors comment on this?

      This is a valuable suggestion that we have not pursued in the current revision, as constructing a common eigenspace across participants requires additional methodological choices.

      (20) I am a bit puzzled by the hypothesis that questionnaires should mediate the intervention effect. Shouldn’t questionnaires be related to before-intervention controllability only? Similarly, could one use initial controllability to predict the intervention response (irrespective or not of the clinical score)?

      Our hypothesis was that participants with greater difficulties in emotion regulation (high DERS-18) would show a smaller intervention effect, as their emotional system might be less amenable to brief distancing. However, this was not the case. We also note that DERS-18 scores were not significantly related to the overall magnitude of pre-intervention controllability (norm of the controllability matrix), but were related to its direction: participants with higher DERS-18 scores showed a most controllable direction pointing toward disgust and away from amusement and calmness, suggesting that trait-level regulation difficulties are linked to the specific emotional configuration of the system rather than its overall controllability.

      Regarding using initial controllability to predict the intervention response: we agree this is a mechanistically appealing question, but it is unfortunately not straightforward to address here. The intervention effect would be quantified as the change between pre- and post-intervention controllability, and since pre-intervention controllability is a constituent of that change score, any correlation between the two would be partially circular by construction.

      (21) The recovery procedure should be way more detailed. How many surrogates were run, etc.? Which kind of quality checks were used to ensure the recoverability was sufficient at the subject level, especially as parameter recovery seems relatively low for some subjects?

      The parameter recovery section now reports that 100 surrogate trajectories were generated per subject (N> = 104) per time period (before and after intervention: N = 208 observations total), with per-subject recovery quality reported across simulations together with across-subject variability; see Supplementary Materials F Parameter Recovery.

      (22) As the result of the eigen decomposition is the endpoint of the analysis, it would be a nice addition to test the recoverability of those measures (eg. correlation between simulated and inferred eigenvalues).

      Recovery of the dominant eigenvalue/vector and dominant singular value/left singular vector is now reported explicitly, including bias analysis and scatter plots of true versus recovered values. Beyond subject-level recovery, we also assessed whether the observed group differences in emotional dynamics and controllability could be reliably recovered at the group level. See Supplementary Materials F Parameter Recovery.

      (23) Although this comment comes close to last, this is a major concern of mine. I am not convinced that the recoverability procedure is sufficient to prove that the inference is working. The model assumes that the observation noise is Gaussian, which is clearly not the case in the data. By simulating surrogate time series with normally distributed noise, the authors do not account for any saturating effects that could destroy a large part of the behavioural information necessary for a successful inversion (eg Figure 4.C showing that simulated data contains a lot of negative ratings). A workaround would be to bind the surrogate time series to mimic the saturation caused by the rating scale, and then run the model estimation on those capped time series.

      We acknowledge this important limitation: the Gaussian assumption does not capture the bounded 0–100 scale, and we have added this to the Limitations with a suggestion that future work use a truncated or censored observation model.

      (24) Supplementary tables with placeholders (v1, v2) that can vary in meaning depending on the line are extremely hard to decipher.

      The supplementary tables have been restructured to a hierarchical format.

      (25) Table I9 is not referenced in the manuscript.

      A reference to this table has been added in the appropriate Results section.

      (26) It’s a shame that neither data nor analysis code has been made available.

      Fully anonymised data and analysis code are now publicly available on GitHub (https://github.com/huyslab/emotioncon public).

      Reviewer #2 (Recommendations for the Authors):

      (1) Abstract: By some definitions, controllability is binary, present or absent, according to whether the controllability Gramian is positive definite. Mention that you use a continuous definition, otherwise ’quantified’ leads to confusion.

      The abstract now describes the measure as: ”Controllability was assessed using continuous measures derived from the controllability matrix, capturing relative differences in how emotional states can be driven by external inputs.”, making the non-binary usage explicit.

      (2) p 5: ’on [not in] the recruiting platform’.

      Corrected.

      (3) p 8: Clarify notation of x<sub>t</sub>t = 1<sup>T</sup> (and same for u). What does this mean?

      The notation has been corrected.

      (4) p 8: Define h in the paragraph following Eq 1, don’t wait until the next section.

      The definition of the bias h (steady-state baseline) has been moved to immediately after Equation 1.

      (5) p 9: Give clear references for your methods here. There are different definitions of controllability etc. than the ones you use.

      References have been added at each key definition.

      (6) p 28: ’20 videos per emotion category were chosen resulting in 50 videos per sequence’ - doesn’t make sense.

      This has been clarified: 20 videos per category across 5 categories yield 100 videos in total. These were split into two matched sequences of 50 videos each. Including 2 videos that were repeated twice resulted in 54 videos per sequence and 108 videos in total.

      (7) Figure B2: 54 videos are listed, and categorized into five categories. How does this relate to the remark right above?

      A clarifying note has been added to the supplementary explaining that each sequence of 54 clips includes repeated videos and is drawn from the pool of 100, with emotion-category sequences matched between blocks.

      (8) p 29: What were the process noise Σ and observation noise Γ assumed in the parameter recovery exercise? What were the consequences of that assumption as assessed by simulation?

      Both Σ and Γ were estimated from the data constrained to be diagonal; the parameter recovery section now explicitly reports their recovery quality.

      (9) p 30: How are results affected by including the excluded participants? The level required to pass attention checks seems arbitrary. How was it chosen?

      We have added a sensitivity analysis including the four excluded outliers, showing results remained qualitatively and statistically similar. With 10 binary attention checks, chance performance is 50%, meaning a participant scoring below 70% is performing only marginally above chance and likely not attending consistently. At the same time, 70% is permissive enough to retain participants who may have missed one or two checks due to momentary distraction.

      (10) Figure F4: Typographically distinguish capital letters referring to panels in the figure from those referring to matrices.

      Panel labels in Figure F4 are formatted in bold to distinguish them from italicised matrix notation.

      (11) Table G1: Showing that differences between groups were non-significant before the intervention but significant after is not enough, you need to show that there was a significant interaction between time point and intervention. [I wrote this after reading the supplementary but before reading the main text. It turns out you know what I’m telling you here. You should mention it more prominently though, including in the abstract and the discussion, because in your chosen null-hypothesis significance testing framework, this is the crucial test of your study. I don’t think there’s any harm at all in being up-front about this - certainly much better than making excuses like the one about randomization at the top of page 11, which I recommend removing].

      This is very right. The significant group × time interaction effects are now reported prominently in the main text Results and figure captions. We also wish to be transparent: the interaction tests were conducted post-hoc rather than as the primary analysis, which is the reverse of the correct order. In the interest of transparency we have left the structure of the results section intact rather than retrospectively reframing it.

      Reviewer #3 (Recommendations for the Authors):

      (1) I would encourage the authors to re-add some basic details regarding their power analyses from the supplement to the main text so the reader can immediately reference the intended effect size, which analysis/analyses were considered primary for the power analysis, etc.

      More information about the power analysis have been added to the Participants section of the main text.

      (2) Similarly, I wondered if there could be a little extra information on how the test-retest reliability was calculated (page 13) - on the first and last views of a video pre-intervention? I wasn’t sure - why do the authors only present ICCs for amusement/joy and disgust/horror? Seems useful to present all (particularly because there may be individual differences in habituation to some emotions).

      We have expanded the test-retest section to clarify that reliability was assessed using duplicate videos shown three times pre-intervention, and now report ICCs with confidence intervals, Cronbach’s α, and habituation/sensitization tests for both disgust and amusement; the selection of these two categories reflects which videos were repeated in the design, they were chosen at random during experimental design.

      (3) Regarding my comments about controllability/emotional control, I think the authors probably have two choices - address this head-on (e.g. with a note describing the relationship/distinction between these two concepts of controllability), or else avoid using it in one of the senses (I would suggest the mathematical sense since overriding the concept of cognitive control seems harder - the authors could use phrases like ’impact of emotional inputs’ instead of ’controllable’). In particular, the abstract could be clearer about the nature of controllability as implemented by the authors - this seems critical for communicability. Because of the high relevance of both of these ’control’ concepts to the paper, if the authors agree with my concern, I would also suggest changes throughout, such as in the results section phrasing: ”In those participants with high DERS-18 scores, the most controllable direction pointed towards disgust ( = 0.26, p = 0.006), and away from amusement ( = 0.26, p = 0.005) and calmness ( = 0.24, p = 0.011; though this did not survive Bonferroni correction).”

      Thank you for this comment. Please see our response to the public review comments above, which we hope address this.

      (4) Regarding the control intervention, which is great, is it possible the follow-up question/reminder affected the results – e.g., is there reason to believe that prospective regulation was the primary difference in the distancing group and not a retrospective effect via this question?

      See our response to Reviewer 2 public review Comment 2 and the corresponding Limitations addition.

      (5) Lastly, purely for interest, the authors could consider elaborating on their brief interpretation as to why difficulties in regulating emotions were specifically linked to the controllability of disgust, amusement, and calmness, but not other emotions (anxiety/sadness) (page 20). I wonder if there is a brief space to discuss the emotional specificity of these results further given the relevance to the wider literature on specific emotion types, e.g. fear vs disgust.

      We have expanded the Discussion to elaborate on the emotional specificity of these findings. Amusement and disgust are strongly influenced by external events, suggesting that stimulus-driven controllability is particularly relevant for these emotions. By contrast, anxiety and sadness are maintained through internally generated processes such as rumination and anticipatory cognition, exhibiting greater emotional inertia over time, and their regulation may therefore be less sensitive to momentary stimulus controllability. This provides a mechanistic account of why controllability effects emerged selectively for disgust, amusement, and calmness, and aligns with growing evidence that emotion regulation is emotion-specific rather than domain-general.

    2. eLife Assessment

      This important manuscript proposes a dual behavioral/computational approach to assess emotional regulation in humans. The authors present convincing evidence for the idea that emotional distancing (as routinely used in clinical interventions for e.g. mood and anxiety disorders) enhances emotional control.

    3. Reviewer #1 (Public review):

      Summary:

      Using sequences of short videos to elicit emotional changes in participants, Malamud and Huys demonstrate how a brief, controlled emotion regulation intervention (distancing) can effectively alter subsequent emotion ratings. A novel computational approach based on state-space models captures the trajectories of emotion ratings and leverages tools from control theory to quantify the intervention's impact on emotion dynamics.

      Strengths:

      The experiment is well designed and tailored to the computational modeling approach advanced in the paper. It also relies on a selection of previously validated stimuli. Within the constraints of a controlled experiment, the intervention successfully implements a relatively common tool used in psychotherapeutic treatment, supporting its clinical relevance.

      The computational modeling is grounded in the well-established framework of dynamical systems and control theory. This foundation offers a conceptually clear formalization, along with powerful quantification tools that go beyond previous, more data-driven approaches.

      Overall, this timely study presents a coherent approach that bridges concepts from clinical psychology and computational theory, providing a stepping stone toward more quantified, evidence-based psychological interventions targeting emotion control.

      Weaknesses:

      A limitation of this study is that the data were acquired online, resulting in some heterogeneity in the measured effects and reduced statistical power when testing for complex interactions. While the current data are sufficiently solid to demonstrate the validity of the general concept and computational approach, future work should aim to replicate these results in a more controlled laboratory setting.

      Additionally, the repeated reminders of the distancing instruction during the second phase of the experiment raise questions about the generalizability of the findings to longer-term remediation strategies, as typically implemented in clinical settings.

    4. Reviewer #3 (Public review):

      Summary:

      The manuscript takes a dynamical systems perspective on emotion regulation, meaning that rather than a simplistic model conceptualising regulation as applying to a single emotion (e.g. regulation of sadness), emotion regulation could cause a shift in the dynamics of a whole system of emotions (which are linked mathematically to one another). This builds on the idea that there are 'attractor states' of emotions between which people transition, governed by both the system's intrinsic characteristics (e.g. temporal autocorrelation of a particular emotion/person) and external driving forces (having a stressful week). Conceptually this is a very useful advance because it is very unlikely that emotions are elicited (or reduced) singly, without affecting other emotions. This paper is a timely implementation of these ideas in the context of a psychotherapeutic intervention, distancing, which participants were trained (randomised) to perform while watching emotion-inducing videos.

      The authors' main conclusion is that distancing both stabilises specific emotional patterns and reduces the impact of external video clips. I would consider these results strong and believable, and to have the potential to impact models of emotion regulation as well as the field's broader views on the mechanisms of psychological therapies.

      Strengths:

      This paper has very many strengths: I would especially note the authors' very-well-matched active control condition and the robustness of their model comparison approach. One feature of the authors' approach in is that they explicitly add noise - not what you typically see in an emotion time-series analysis - which allows for participants to make errors in their own subjective ratings (a reasonable thing to assume); this noise can then be smoothed during filtering. In their model comparison approach, they explicitly test whether a true dynamical system explains emotion change/emotion regulation effect on emotions - demonstrating that both intrinsic dynamics and external inputs were needed to explain subjective emotion. Powerfully, they also used this approach to test the differential effects of the treatment groups (see below).

      The main result seems quite robust statistically. Verifying the effects of the distancing intervention on emotion, the authors found an interaction between time (pre- to post-intervention) and intervention group (distancing vs. relaxation) suggesting that distancing (but not relaxation) reduced ratings of almost all emotions. Participants allocated to the distancing intervention also showed decreased variability of emotion ratings compared to those in the relaxation intervention (though note this interaction was not significant).

      Using a model comparison approach, the authors then demonstrated that whilst the control group was best-explained by a model that did not change its dynamics of emotions, the active intervention (distancing) group was best-explained by a model that captured both changing emotion dynamics and a changing input weights (influence of the videos) - results confirmed in follow-up analyses. This is convincing evidence that emotion regulation strategies may specifically affect the dynamics of emotions - both their relationships to one another and their susceptibility to changes evoked by external influences.

      The authors also perform analyses that suggest their result is not attributable to a demand effect (finding that participants were quicker during the control intervention, which one would expect if they had already decided how to respond in advance of the emotion question). I personally also think a demand effect is unlikely given the robustness of their control intervention (which participants would be just as likely to interpret as a mental health-enhancing training as distancing) and am convinced by the notion that demand effects would be unlikely to elicit their more specific effects on the dynamic quality of emotions.

      Weaknesses:

      The authors use an active control - a relaxation intervention - which is extremely closely matched with their active intervention (and a major strength). However, there was an additional difference between the groups: "in the group allocated to the distancing intervention, the phrasing of the question about their feelings in the second video block reminded participants about the intervention, stating: "You observed your emotions and let them pass like the leaves floating by on the stream." Therefore, some of the effects of distancing may have also been driven by different emotion regulation strategies, i.e. reappraisal, since this reminder might have evoked retrospective changes in ratings.

      An unanswered question is exactly how distancing is producing these effects. As the authors point out, there is a possibility that eye-movement avoidance of the more emotionally-salient aspects of scenes could be changing participants' exposure to the emotions somewhat, which could vary by emotion, as the authors now discuss in their limitations.

      Comments on revised version.

      The authors have addressed my concerns.

    1. eLife Assessment

      This important study partially fills the gap in the knowledge of olfaction at the level of the Anterior Olfactory Nucleus (AON) and Piriform Cortex with functional magnetic resonance imaging, electrophysiology, and modeling. The methods used are convincing. Some of the findings confirm ongoing hypotheses, such as the behavioral importance of AON for odor source discrimination, while other results shed light on the dynamics of the connection between the olfactory system and the rest of the brain. However, the lack of access to raw data hampers the assessment of data quality.

    2. Reviewer #1 (Public review):

      Summary:

      This manuscript combined rat fMRI, optogenetics and electrophysiology to examine the large-scale functional network of the olfactory system as well as its alteration in an aged rat model.

      Strengths:

      Overall methodology is very solid and the results provided an interesting perspective on large-scale functional network perturbation of the olfactory system.

      Weaknesses:

      The biological relevance and validation of the current results can be improved.

      Comment on revised version.

      Authors made satisfactory revision and I have no further comments.

    3. Reviewer #2 (Public review):

      Summary:

      Ma and colleagues presented a study on the characterization of brain-wide spatio-temporal impact of olfactory cortical outputs. They take advantage of multi-modal techniques on rats: fMRI, optogenetics and electrophysiology. In addition, they used cutting-edge analytical techniques and modeling to support and interpret their data. The main findings of the study are:

      (1) Neurons in Olfactory Bulb (OB) predominantly activate primary olfactory network regions, while stimulation of OB afferents in Anterior Olfactory Nucleus (AON) and Piriform Cortex (Pir) primarily orthodromically activates hippocampal/striatal and limbic networks, respectively.<br /> (2) Non-specified adaptation or habituation mechanisms may play a significant role in modulating olfactory outputs over subsequent fMRI sessions.

      (3) Artificially induced aging in rats induces profound modification in the functional interaction between olfactory cortices and multiple brain regions.

      The results on AON are of particular interest because of the lack of functional information on this region, despite its recognized importance in shaping OB output and behavior (odor localization tasks).

      Strengths:

      The manuscript is very accurate. The figures are well-crafted, clear and provide much information with the most appropriate plots and graphics. The study's amount and data quality are remarkable, and the experimental size adequately addresses the scientific questions. I particularly appreciated the details in the description of the methods regarding the missing data and the size of the different animal groups. The supplementary data complete the leading figures and provide information at a single animal level.

      Weaknesses:

      (1) One of the main reasons the Piriform Cx is understudied in rodents is because of the proximity to air, which creates artifacts in fMRI images. This issue becomes more critical at ultra-high magnetic fields, but I would expect it also at 7T. One main achievement of this study is, indeed, the acquisition of fMRI data from Piriform, and this point should be highlighted by showing raw functional data from a rat. The best would be if an fMRI data sample for a rat, no matter which stimulation, is shared on a public repository, like Zenodo or similar. I am curious to check the quality of the BOLD data from such an 'enormous' field of view, particularly in the OB, with a single-shot sequence. Also, the visual inspection of raw data is essential to appreciate how many 0.5 x 0.5 x 1 mm voxels fit into AON, and others analyzed small brain structures, like the amygdala, etc. Was the amygdala entirely visible in BOLD, or did the air in the ear channel make an artifact partially shadowing it?

      (2) Surprisingly, the only information missing in the methods is the post-surgery period and the time between two consecutive fMRI sessions. How much time was accorded to rats to recover from the surgeries, and what time interval between two scans? This information is crucial for interpreting the decrease in most BOLD responses in subsequent recordings. The supposed adaptation should fit into the known time frames for odor adaptation. Usually, fast adaptation does not last for days (and it should be measured within a single experiment: is it the case?), while for long-lasting adaptation the stimulus (odor or opto) should be maintained constantly ON. This does not seem to be the case in this study. The hypothesis, alternative to adaptation, of a less efficient light activation, for example, due to gliosis around the fiber tips, should be discarded with more evidence than the preservation of OB > Pir responses or acknowledged in the manuscript.

      (3) The D-galactose experiments were conducted only after administering the aging molecule, with no baseline/reference data on the same animals. Then, comparisons were made with healthy rats, but the two groups not only can be discriminated with respect to D-galactose administration but also with age (10 VS 18 weeks). A control group for 18-weeks-old rats with no D-galactose treatment would better compare the D-galactose effect and avoid any potential bias from group comparisons of rats at different ages. Do you confirm that D-galactose was injected into each rat 56 times/days in a raw, or am I mistaken?

      The updated version of the manuscript partially addresses the flaws of the original submission. Here are my general concerns:

      (1) Overall, the revised version comes with a few modifications/additions and no new data. Apart from a new correlation analysis, the improvements are mainly discursive, often non-convincing, justifications of the authors' choices. This may reflect a lack of interest in a publication that, in the meantime, lost its original peer-review value. However, it should be acknowledged that the Authors made an effort to partially address the concerns raised by the reviewers.

      (2) My main concern was the quality of fMRI recordings. In the revised version, the Authors provided a new figure with an example single-mouse fMRI data. However, the depicted regions of interest (ROIs) mostly cover the brain spots that I expected to be the most impacted by the BOLD artifacts caused by the proximity of the air and the big field-of-view. In addition, these ROIs do not appear to match the mouse anatomy shown above the functional data. As an example, the EPI images in the OB are almost entirely covered by the colored mask. The feeling is that the fMRI data was indeed poor, as I worried, and the lack of any public repository of raw data reinforces that feeling. To make this point clear: I do not think the findings are not true, but poor fMRI data quality might have hidden more insightful results and does not foster the use of fMRI to monitor the olfactory pathway, which lowers the impact of this article.

    4. Author response:

      The following is the authors’ response to the original reviews

      eLife Assessment

      This important study partially fills the gap in the knowledge of olfaction at the level of the Anterior Olfactory Nucleus (AON) and Piriform Cortex (Pir) with functional magnetic resonance imaging, electrophysiology, and modeling. The methods used are convincing. Some of the findings confirm ongoing hypotheses, such as the behavioral importance of AON for odor source discrimination. Other results shed light on the dynamics of the connection between the olfactory system and the rest of the brain.

      We sincerely thank the editors and reviewers for the thorough review of our manuscript. We appreciate the insightful comments, which have significantly contributed to improving our work. In this revision, we addressed all the concerns posed by the reviewers, including conducting additional analyses and providing the data generated and codes used in this study.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This manuscript combined rat fMRI, optogenetics, and electrophysiology to examine the large-scale functional network of the olfactory system as well as its alteration in an aged rat model.

      Strengths:

      Overall methodology is very solid and the results provided an interesting perspective on large-scale functional network perturbation of the olfactory system.

      Weaknesses:

      The biological relevance and validation of the current results can be improved.

      We thank the reviewer for the comments and suggestions regarding our manuscript. They have been invaluable and instrumental in enhancing our work. Please see R1-1 to R1-8 below for our corresponding responses and revisions.

      Comments:

      (R1-1) Figure 1A, on the top of the figure, ChR2 may be replaced by ChR2-mCherry, as only mCherry is fluorescent. And also, it’s somewhat surprising that in AON and Pir regions (where only axon fibers should be labelled as red), most fluorescence appeared dot-like and looked more similar to cell body instead of typical fiber. The authors may want to double-check this.

      (a) Thank you for pointing out the missing fluorescent marker for ChR2 expression in the label of Figure 1A. We distinguished the transfection of ChR2 in olfactory bulb (OB) neurons from axonal terminals in AON and Pir by identifying the expression patterns in these three regions (see Figure 1). In OB, the cellular nucleus (i.e., labelled in blue by DAPI) was surrounded evenly by ChR2-mCherry expression (red fluorescent, indicated by green arrows in Figure 1), which clearly traced the neuronal cell body shape. Meanwhile, the expression of mCherry in AON and Pir consisted of concentrated red fluorescent dots that were sparsely assembled near the cell nucleus, indicating the presence of synaptic boutons at the axonal terminals. Further, the patterns of expression here did not trace the shape of the neuronal cell body.

      (b) In this revision, we amended the label of Fig. 1A to “ChR2-mCherry”.

      In this revision, changes were made to the confocal images of AON and Pir, based on Figure 1, for clarity and the figure caption was also edited accordingly.

      (R1-2) The authors primarily presented 1 Hz stimulation results. What is the most biologically relevant frequency (e.g., perhaps firing frequency under natural odor stimulation) among all frequencies that were used?

      (a) There are various firing rate and oscillatory frequency of neural activities within the olfactory system (i.e., OB, Pir) in rodents during olfaction. For example in OB, neural oscillations ranging from 1 to 12 Hz (i.e., slow to theta) are driven by sensory stimulation and are closely linked to respiration [1]. Specifically, odor-evoked responsive excitatory bursts of mitral and tufted cells in OB are coupled with the low-frequency respiration rhythm (1-4 Hz) [2,3]. Such coupling has also been documented during light anaesthesia [4]. Meanwhile, higher frequencies such as beta oscillations (15 ~ 30 Hz) have been associated with odor learning and sensitization [5,6], and gamma oscillations (40 ~ 80 Hz) evoked by sensory stimulation are associated with fine olfactory discrimination and odor learning [6-8].

      The piriform cortex (Pir) also exhibits natural oscillatory activity across various frequency bands, including slow, theta, beta, and gamma oscillations [9-11]. As in OB mitral and tufted cells, slow oscillations (< 1.5 Hz) in Pir are also correlated with the respiratory rhythm, indicating the intrinsic respiration-related oscillatory properties within the olfactory system [12]. However, while it has been shown that the primary burst of firing in the anterior Pir is locked to respiration, its coding strategies differed from OB during olfaction [13].

      Despite not being able to entirely mimic the evoked neural activities under natural odor stimulation due to the synchronizations induced by our optogenetic stimulations, our frequencies were chosen to be within the range of firing rates of neurons in the olfactory system under natural circumstances. Our selection of 1 Hz frequency for optogenetic stimulation was based on the balance of experimental simplicity of inducing neural excitation and the biological significance of low-frequency neural oscillations in OB, Pir and AON. The robustness of evoked brain-wide long-range activations is one of our criteria for the selection of 1 Hz stimulation frequency at OB, considering that the goal of this study is to examine long-range olfactory networks. Furthermore, we also examined other frequencies of stimulation covering theta, beta and gamma frequency bands, which provided insights into the response characteristics within the olfactory networks.

      (b) In this revision, we added a brief statement in the Results section that 1 Hz was the most biologically relevant frequency based on discussions in R1-2a above.

      (R1-3) In Figure 2, the statistical thresholding is confusing: in the figure legend, it was stated that “t > 3.1 corresponding to P < 0.001” but later “further corrected for multiple comparisons with thresholdfree cluster enhancement with family-wise error rate (TFCE-FWE) at P < 0.05”? Regardless of the statistical thresholding, such BOLD activation seemed to be widespread (almost whole-brain activation). Does such activation remain specific to the optogenetic stimulation, or something more general (e.g., arousal level change)? Furthermore, how those results (I assume they are group-level results) were obtained was not described very clearly. Is it just a simple average of individual-level results, or (more conventionally) second-level analysis?

      (a) We thank the reviewer for drawing our attention to clarity issues regarding the generation of group-level BOLD activation maps and statistical thresholding methods used.

      In brief, we first generate the BOLD activation map of individual animal through conventional general linear model (GLM) analysis. These individual activation maps then underwent a two-step statistical thresholding method to generate the group-level results. Firstly, uncorrected one-sample t-tests were conducted with a threshold of P < 0.001. Secondly, these thresholded activation maps were further corrected using nonparametric inference with threshold-free cluster enhancement multiple comparison correction of family-wise error rate (TFCE-FWE, P < 0.05) before they were averaged. Hence, the BOLD activation maps presented in our manuscript were the result of group-level analysis instead of individual-level analysis.

      (b) We note the reviewer’s concern about whether the observed widespread BOLD activations were caused by the animal’s general brain state (e.g., arousal levels) rather than the specificity of optogenetic stimulation. First, such widespread activations were only specific to 1 Hz stimulation of OB neurons and OB afferents at AON (Figs. 2A, B and 3A, B) whereas activations were localized to regions in the primary olfactory network with increasing stimulation frequencies. Furthermore, varied neural activity adaptation properties were observed (Fig. 3A, B) following repeated 1 Hz optogenetic excitation of OB neurons and OB afferents at AON indicating that such widespread propagation of evoked neural activity was not driven by the animal’s general brain state, which would be relatively random across animals as we interleaved the presentation of each stimulation frequencies. While we cannot discount that certain characteristics of brain states differ under anaesthetized and awake conditions, we showed that light (1.0% isoflurane) anaesthesia minimally affects the BOLD fMRI activations and/or the propagation of optogenetically-evoked neural activity across thalamo-cortical, hippocampal-cortical and vestibulo-cortical networks [14-20].

      Second, the neural representations in OB have been shown to be brain state-independent to ensure the high fidelity of olfactory inputs to higher olfactory cortices, despite differences in spontaneous baseline activity across states [21-24]. In fact, OB neural activity (i.e., slow to delta oscillations) remains highly coupled to respiration rhythms under low anaesthesia as in awake conditions [4]. Further, low-dose isoflurane (i.e., 1% as in the present study) does not affect cortical gamma oscillations (40 Hz) [25-28]. Although odorant-evoked responses in olfactory cortices (e.g., Pir and olfactory tubercle, Tu) can be modulated by overall brain state, various interactions between such cortical circuits to process odor inputs persist under anaesthesia [29-31].

      (c) In this revision, we clarified the two-step statistical thresholding method that was applied to generate the group-level BOLD activation maps, as discussed in R1-3a above, in the captions of Fig. 2 and the Methods section of the manuscript.

      In this revision, we also included a statement in the Discussion section that the widespread BOLD fMRI activations were specific to the 1 Hz optogenetic stimulation and not the animal’s general brain state, as discussed in R1-3b above.

      (R1-4) In Figure 2, why use AUC to quantify the activation, not the more conventional beta value in the GLM analysis?

      (a) We thank the reviewer for raising this concern. We are aware of the more conventional beta value, b, in GLM analysis and have used them for quantification and comparison of the amplitudes of fMRI activations in our previous rodent fMRI studies [18,32,33]. However, in this study, we chose to utilize area under curve (AUC) as it offers a more comprehensive measure of BOLD signal change over time, including shape, duration, and magnitude, thereby capturing the bulk of neural activities and their dynamics throughout the stimulation period. b primarily represents the peak amplitude of BOLD responses (i.e., the % BOLD signal change) [34] and can be constrained by the assumptions and limitations of the GLM analysis, such as the shape of the canonical hemodynamic response function (HRF). Therefore, AUC provides greater accuracy in capturing different aspects of neural responses across various brain regions, such as transient peaks and/or sustained responses.

      (b) In this revision, we have included the justifications for using AUC over the more conventional b value to quantify BOLD fMRI activations, as in R1-4a above, in the Methods section.

      (R1-5) For Figure 2D, the way that it was quantified can be better described as “relative” activation within one condition, and I don’t know how to interpret the comparison among the relative fraction of activated regions. Perhaps comparison using percentage change (i.e., beta values) is more straightforward.

      (a) We thank the reviewer’s comment and suggestion here. “BOLD activation strength” is the AUCs normalized to the respective sum of BOLD activations of their respective stimulation target. We are of the opinion that “strength” can accurately reflect our purpose for conducting this comparison, which is to distinguish the brain networks that were predominantly recruited by AON- or Pir-driven neural activities compared to OB stimulation. 

      (b) We conducted a comparison using relative fraction for fair comparison considering the differences in absolute BOLD signal amplitudes upon different stimulation targets. As we found the overall amplitudes of activations evoked by Pir stimulation were appreciably weaker, our normalization step ensures that the contribution of Pir is not underrepresented. Without this normalization process, the Pir stimulation seems not to activate most of the downstream targets based on the relatively weak activations. As discussed earlier in R1-4a above, comparison using percentage change/b value would be skewed towards the absolute amplitude of BOLD responses and would be erroneous for subsequent interpretation of brain networks that are primarily recruited by OB vs. AON vs. Pir.

      (b) In this revision, we chose not to include additional statements as we have already described in the Results section the reasons for normalizing AUC to reflect and subsequently compare the BOLD activation strength across various networks that were recruited by optogenetic stimulation of three distinct olfactory regions (i.e., OB, AON and Pir). 

      (R1-6) For Figure 3, it may be more convenient for readers to include the results of 1st activation for direct comparison. The current layout makes it difficult to make direct, visual comparisons among all 3 activations. Again, I think using beta values (instead of AUC) may be more conventional.

      (a) We agree that it will be more convenient for readers if we include the 1st activation maps in Figure 3 for direct comparison. Please refer to R1-4 for the usage of AUC rather than beta values.

      (b) In this revision, we modified Fig. 3 to include activation maps from the 1st fMRI session for easy comparison with the 2nd and 3rd sessions.

      (R1-7) Can the DCM results (at least part of it) be verified using the current electrophysiological data? For example, the long-range inhibitory effective connectivity of AON is rather intriguing. If that can be verified using the electrophysiology data, it would be really great. In the current form, the DCM and electrophysiology results seem to be totally unrelated.

      (a) We thank the reviewer for raising this concern, and it’s a great suggestion to causally link the outcomes from dynamic causal modeling (DCM) analysis and electrophysiology findings.

      (b) In principle, the recorded local field potentials (LFPs) can be treated as the neural dynamic component in DCM analysis under specific conditions. They are as follows:

      (1) LFPs are recorded at the identical brain state as the optogenetic fMRI experiments with identical stimulation paradigms.

      (2) The electrode locations of LFP recordings must match the nodes of the a priori matrix defined for the existing DCM analysis (Fig. 4A).

      However, in the present study, we did not have LFP recordings at the entorhinal cortex (Ent), which will likely influence the modeling of effective connectivity strength as one of the nodes defined is dropped from the analysis. Previous studies have showed that changes in the a priori matrix will result in different connectivity estimations [35-38]. We chose to model the connectivity involving Ent due to its documented interactions with the primary olfactory cortices and hippocampal regions [39,40]. Hence, it would be incomplete without modelling the contributions from Ent.

      (c) In this revision, we decided not to redefine the a priori matrix by removing Ent as a node in the DCM analysis to estimate new effective connectivities. Instead, we acknowledge the absence of Ent recordings.

      (R1-8) In Figure 6, it would be great if the adaptation of BOLD and electrophysiology signals can be correlated at the brain region level. The current figure only demonstrated there is adaptation in the electrophysiology recording, but did not show if such adaptation is related to the BOLD adaptation.

      (a) We thank the reviewer for pointing out that we should associate the electrophysiology and fMRI data to validate the olfactory adaptation. As such, we conducted additional correlation analysis between LFP power and BOLD signal profiles. The method for computing the correlation coefficient is described below:

      (1) For each 1 Hz optogenetic stimulation target (i.e., OB, AON, and Pir), regions of interest (ROIs) that have both LFP recordings and BOLD activations were selected. These ROIs were AON, Pir, ventral caudate putamen (vCPu), visual cortex (V1), and ventral hippocampus (vHP) for OB stimulation; OB, Pir, vCPu, V1, amygdala (Amg), and vHP for AON stimulation; and OB, AON, vCPu, V1, Amg, and vHP for Pir stimulation. Note that we excluded the respective stimulated region as a ROI because of BOLD fMRI signal dropout caused by the implanted optical fibre.

      (2) We first calculate the individual animal AUC differences of 2nd stimulation session vs. 1st session and 3rd session vs. 1st session in LFP power and BOLD signal profiles, respectively.

      Subsequently, the mean of the differences for each ROI across animals was then calculated.

      (3) Correlation coefficients and P values of AUC differences between LFP power and BOLD signal profiles were computed using Spearman nonparametric correlation.

      (b) The relationship between the differences of LFP power and BOLD signal profiles showing neural adaptation was described by the correlation coefficient, r, and the corresponding P value. The r and P values upon the three distinct stimulations were OB: r = 0.77 (P = 0.01), AON: r = 0.46 (P = 0.13), and Pir: r = - 0.06 (P = 0.85), respectively. This indicates that the decrease of LFP power is highly correlated with the decrease of BOLD activations upon OB and AON stimulation compared to those upon Pir stimulation.

      (c) In this revision, we added the description of the correlation analysis conducted, as in R1-8a above, in the Methods section.

      In this revision, we also added several statements in the Results section to indicate that neural adaptation observed is significantly related to the decreased BOLD activations when stimulating OB excitatory neurons or OB afferents at AON, as discussed in R1-8b above. Additionally, we also included the outcomes of the correlation analysis as a new Supplementary Fig. 7.

      Reviewer #2 (Public review):

      Summary:

      Ma and colleagues presented a study on the characterization of brain-wide spatio-temporal impact of olfactory cortical outputs. They take advantage of multi-modal techniques on rats: fMRI, optogenetics, and electrophysiology. In addition, they used cutting-edge analytical techniques and modeling to support and interpret their data. The main findings of the study are:

      (1) The neurons in the Olfactory Bulb (OB) predominantly activate primary olfactory network regions, while stimulation of OB afferents in Anterior Olfactory Nucleus (AON) and Piriform Cortex (Pir) primarily orthodromically activates hippocampal/striatal and limbic networks, respectively.

      (2) Non-specified adaptation or habituation mechanisms may play a significant role in modulating olfactory outputs over subsequent fMRI sessions.

      (3) Artificially induced aging in rats induces profound modification in the functional interaction between olfactory cortices and multiple brain regions.

      The results on AON are of particular interest because of the lack of functional information on this region, despite its recognized importance in shaping OB output and behavior (odor localization tasks).

      Strengths:

      The manuscript is very accurate. The figures are well-crafted, and clear and provide much information with the most appropriate plots and graphics. The study’s amount and data quality are remarkable, and the experimental size adequately addresses the scientific questions. I particularly appreciated the details in the description of the methods regarding the missing data and the size of the different animal groups. The supplementary data complete the leading figures and provide information at a single animal level.

      We are grateful for the reviewer’s appreciation of our study’s breadth and depth, and the quality and quantity of our data. The recognition of the strengths in our methods and the inclusion of single-animal-level data in the supplementary information is highly encouraging. We are also appreciative for the reviewer’s constructive comments below, which has help us to address specific weaknesses of the study and thereby improving the manuscript text and flow. We have read through the comments carefully and have made the necessary corrections. We hope that our responses to R2-1 to R2-11 below has sufficiently addressed all the reviewer’s concerns.

      Weaknesses:

      (R2-1) One of the main reasons the Piriform Cx is understudied in rodents is because of the proximity to air, which creates artifacts in fMRI images. This issue becomes more critical at ultra-high magnetic fields, but I would expect it also at 7T. One main achievement of this study is, indeed, the acquisition of fMRI data from Piriform, and this point should be highlighted by showing raw functional data from a rat. The best would be if an fMRI data sample for a rat, no matter which stimulation, is shared on a public repository, like Zenodo or similar. I am curious to check the quality of the BOLD data from such an ‘enormous’ field of view, particularly in the OB, with a single-shot sequence. Also, the visual inspection of raw data is essential to appreciate how many 0.5 x 0.5 x 1 mm voxels fit into AON, and others analyzed small brain structures, like the amygdala, etc. Was the amygdala entirely visible in BOLD, or did the air in the ear channel make an artifact partially shadowing it?

      (a) As per the reviewer’s suggestion, we displayed the representative EPI images after preprocessing at a matrix size of 128 x 128 with a pixel size of 0.25 x 0.25 mm in Supplementary Figure 2. Up-sampling was not applied out-of-plane. Note that the atlas-based region-of-interest (ROIs) used to extract BOLD signal profiles (i.e., indicated by colored overlays in Supplementary Figure 2) were drawn based on the up-sampled EPI images (from the acquired 64 x 64 to 128 x 128). Notably, the piriform cortex (Pir) and amygdala (Amg) are visible with no appreciable signal dropout at these regions. To clarify, the Pir in our study represented the anterior Pir.

      Signal dropout is pronounced in the posterior ventral regions of the brain (e.g., Bregma- 4 mm, below the Ent) and as expected in a localized region where the implanted optical fiber region that targeted the AON.

      (b) In this revision, we included as a new Supplementary Figure 8 and added a statement in the Results section to describe the absence of appreciable signal dropouts at OB, Pir and Amg, which could affect subsequent quantitative analyses.

      (R2-2) Surprisingly, the only information missing in the methods is the post-surgery period and the time between two consecutive fMRI sessions. How much time was accorded to rats to recover from the surgeries, and what time interval between two scans? This information is crucial for interpreting the decrease in most BOLD responses in subsequent recordings. The supposed adaptation should fit into the known time frames for odor adaptation. Usually, fast adaptation does not last for days (and it should be measured within a single experiment: is it the case?), while for long-lasting adaptation the stimulus (odor or opto) should be maintained constantly ON. This does not seem to be the case in this study. The hypothesis, alternative to adaptation, of a less efficient light activation, for example, due to gliosis around the fiber tips, should be discarded with more evidence than the preservation of OB > Pir responses or acknowledged in the manuscript.

      (a) We thank the reviewer for drawing our attention to the missing information regarding the timeline of each optogenetic fMRI experiment. We conducted fMRI experiments immediately after the surgical procedure for implanting the optical fiber cannula. Such an implantation procedure typically last for 45mins under 1.2-2.0% isoflurane. The animal is then moved from the surgical table to the magnet to begin the optogenetic fMRI experiment. Intervals between two fMRI scans were within one minute. As the stimulation frequencies (i.e., 1 Hz, 5 Hz, 10 Hz, 20 Hz and 40 Hz) were pseudorandomized, the interval between any two identical stimulation frequency was on average 30 minutes. In this case, we are measuring fast adaptation (on the order of tens of minutes) rather than long-lasting adaptation. Further, as the fMRI experiment was conducted immediately following fiber implantation, the risk of gliosis around the fiber tips is expected to be low.

      (b) The main representation of olfactory adaptation is the attenuation of neural responses upon continuous or repeated stimuli [41,42]. Several fMRI studies in rodents have reported that olfactory adaptation was detected in primary olfactory cortices (i.e., OB, AON, and Pir) upon repeated odor stimulation [43,44]. Specifically, these studies showed robust decreased responses in primary olfactory cortices under repeated odor stimulation with an odor interval of 30 minutes, which were comparable to those observed in our study upon repeated 1 Hz optogenetic stimulation. As such, the pronounced neural adaptation that we demonstrated when stimulating OB excitatory neurons or OB afferents at AON is unlikely to be caused by less efficient light activation.

      (c) In this revision, we added statements in the Methods section, as discussed in R2-2a above, to clarify the timeline of our optogenetic fMRI experiment.

      (R2-3) The D-galactose experiments were conducted only after administering the aging molecule, with no baseline/reference data on the same animals. Then, comparisons were made with healthy rats, but the two groups not only can be discriminated with respect to D-galactose administration but also with age (10 VS 18 weeks). A control group for 18-weeks-old rats with no D-galactose treatment would better compare the D-galactose effect and avoid any potential bias from group comparisons of rats at different ages. Do you confirm that D-galactose was injected into each rat 56 times/day in a row, or am I mistaken?

      (a) We appreciate the reviewer’s comment here and understand the concerns regarding the absence of an age-matched control group with saline administration instead of D-galactose. We acknowledge that the inclusion of a control group would have provided a more direct comparison to evaluate the differences between healthy and aged animals. However, we were unable to include this group in the present study due to logistical constraints.

      To clarify, our experiments were conducted on healthy rats at the median age of 14 weeks to ensure the animals had reached adulthood. The median age of the D-galactose injected animal was 18 weeks. In terms of the lifespan of rats, they reach sexual maturity at approximately 6 weeks of age [45], at which point we conducted the optogenetic viral vector injection. The age of rats at 14 and 18 weeks can both be regarded as teenage adult [45,46]. The primary purpose of the experiments on the aged animal model and corresponding analyses was to provide some potential clues into the dysfunction of olfactory networks at the system level as animals aged. Despite the age difference between the two groups, we believe that the comparison between healthy and aged rats still offers valuable insights into the ageing-related dysfunctions within the olfactory system. 

      (b) Note that D-galactose was injected into each rat 56 times in total over the whole experiment (i.e., one injection per day over the course of 8 weeks).

      (c) In this revision, we acknowledge that the comparisons made were not with age-matched healthy controls in the Discussion section.

      In this revision, we also clarified that D-galactose was administered once daily for 8 weeks (i.e., a total of 56 injections) in the Methods section.

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      Minor:

      (R2-4) The look of the activation maps will greatly improve if the displayed brain slices are coronal (as in the supplementary figures) with no angle.

      We thank the reviewer for this suggestion. The purpose of such a display was to maintain the consistency of displaying the overlay of BOLD fMRI activation maps on the 3D renders of rat brain anatomical images. In addition, such 3D display can better impress readers that the optogenetically evoked BOLD activations were indeed long-range and brain-wide. 

      As such, in this revision, we decided to maintain the existing displays of BOLD activation maps in the main figures.

      (R2-5) The choice of testing different stimulation frequencies deserves more justification. Why was it performed in the OB, which is naturally activated on the breathing rhythm?

      (a) We thank the reviewer’s comment here to seek further clarification. OB is widely recognized as the first stage of olfactory information processing in the brain. By activating the OB excitatory neurons, we can ensure that our optogenetically-evoked activations are olfactory-related.

      The choice of various frequencies ranging from low (i.e., 1 Hz) to high (e.g., 40 Hz) was made to cover a range of firing rate and oscillatory frequency of neural activities neural oscillations that exist within the rodent olfactory system during olfaction. For example in OB, neural oscillations ranging from 1 to 12 Hz (i.e., slow to theta) are driven by sensory stimulation and are closely linked to respiration [1]. Specifically, odor-evoked responsive excitatory bursts of mitral and tufted cells in OB are coupled with the low-frequency respiration rhythm (1-4 Hz) [2,3]. Such coupling has also been documented during light anaesthesia [4]. Meanwhile, higher frequencies such as beta oscillations (15 ~ 30 Hz) have been associated with odor learning and sensitization [5,6], and gamma oscillations (40 ~ 80 Hz) evoked by sensory stimulation are associated with fine olfactory discrimination and odor learning [6-8].

      (b) In this revision, we added a statement in the Results section to further describe our justifications for testing various stimulation frequencies.

      (R2-6) The absence of natural (odor) stimulation should be acknowledged as a limitation for the findings of this study in the Discussion.

      We thank the reviewer for the suggestion. In this revision, we made the acknowledgement in the Discussion section. 

      (R2-7) Line 46: overstatement, the role of the Piriform cortex at the system level is well known and was not discovered by this study. (see Gordon Shepherd’s book: The Synaptic Organization of the Brain, Figure 10.6)

      As per the reviewer’s suggestion, we decided to use “distinguishes” rather than “uncovers” for an accurate description of our study’s findings, which demonstrated the differences between the piriform cortex and anterior olfactory nucleus in driving the downstream activations of olfactory and non-olfactory targets.

      (R2-8) Line 75: please rephrase for an easier understanding.

      As per the reviewer’s suggestion, we edited this statement to “Further, the need to present a multitude of odor combinations also makes it challenging to efficiently and reliably interrogate long-range olfactory networks and their properties”.

      (R2-9) Line 313: the intermediary region between the Piriform and Entorhinal cortices is the Perirhinal Cortex. No doubt about that.

      As per the reviewer’s comment, we modified the corresponding statement by adding “such as the perirhinal cortex” and cited the appropriate references [47,48].

      (R2-10) Line 391: there is no evidence from this study to exclude that the downstream targets would be less activated because of a decreased OB activation in favor of broad inhibition.

      As per the reviewer’s suggestion, we modified the statement by replacing “rather than” with “in addition to” for a more precise statement.

      (R2-11) Line 585: which was/were the regressor/s used in GLM? Any convolution with common HRFs?

      We regarded the block-designed optogenetic stimulation as a task. The regressors are then obtained by convolving the expected task-evoked neural activity (i.e., boxcar function for block design) with the canonical hemodynamic response function, HRF (SPM12, Wellcome Department of Imaging Neuroscience, University College London, UK).

      In this revision, we added a statement in the Methods section to provide clarification on the regressors used in GLM and the convolution step with canonical HRF.

      References

      (1) Kay, L. M. & Stopfer, M. Information processing in the olfactory systems of insects and vertebrates. Semin Cell Dev Biol 17, 433-442 (2006). https://doi.org/10.1016/j.semcdb.2006.04.012

      (2) Cang, J. & Isaacson, J. S. In vivo whole-cell recording of odor-evoked synaptic transmission in the rat olfactory bulb. J Neurosci 23, 4108-4116 (2003). 

      (3) Margrie, T. W. & Schaefer, A. T. Theta oscillation coupled spike latencies yield computational vigour in a mammalian sensory system. J Physiol 546, 363-374 (2003). https://doi.org/10.1113/jphysiol.2002.031245

      (4) Fontanini, A. & Bower, J. M. Variable coupling between olfactory system activity and respiration in ketamine/xylazine anesthetized rats. Journal of Neurophysiology 93, 3573-3581 (2005). https://doi.org/10.1152/jn.01320.2004

      (5) Kay, L. M. et al. Olfactory oscillations: the what, how and what for. Trends in Neurosciences 32, 207-214 (2009). https://doi.org/10.1016/j.tins.2008.11.008

      (6) Gervais, R., Buonviso, N., Martin, C. & Ravel, N. What do electrophysiological studies tell us about processing at the olfactory bulb level? Journal of physiology, Paris 101, 40-45 (2007). https://doi.org/10.1016/j.jphysparis.2007.10.006

      (7) Beshel, J., Kopell, N. & Kay, L. M. Olfactory bulb gamma oscillations are enhanced with task demands. J Neurosci 27, 8358-8365 (2007). https://doi.org/10.1523/JNEUROSCI.119907.2007

      (8) Martin, C., Beshel, J. & Kay, L. M. An olfacto-hippocampal network is dynamically involved in odor-discrimination learning. J Neurophysiol 98, 2196-2205 (2007). https://doi.org/10.1152/jn.00524.2007

      (9) Kay, L. M. Theta oscillations and sensorimotor performance. Proc Natl Acad Sci U S A 102, 3863-3868 (2005). https://doi.org/10.1073/pnas.0407920102

      (10) Lowry, C. A. & Kay, L. M. Chemical factors determine olfactory system beta oscillations in waking rats. J Neurophysiol 98, 394-404 (2007). https://doi.org/10.1152/jn.00124.2007

      (11) Vanderwolf, C. H. & Zibrowski, E. M. Pyriform cortex beta-waves: odor-specific sensitization following repeated olfactory stimulation. Brain Res 892, 301-308 (2001). https://doi.org/10.1016/s0006-8993(00)03263-7

      (12) Fontanini, A., Spano, P. & Bower, J. M. Ketamine-xylazine-induced slow (< 1.5 Hz) oscillations in the rat piriform (olfactory) cortex are functionally correlated with respiration. J Neurosci 23, 7993-8001 (2003). https://doi.org/10.1523/JNEUROSCI.23-22-07993.2003

      (13) Miura, K., Mainen, Z. F. & Uchida, N. Odor representations in olfactory cortex: distributed rate coding and decorrelated population activity. Neuron 74, 1087-1098 (2012). https://doi.org/10.1016/j.neuron.2012.04.021

      (14) Leong, A. T. et al. Long-range projections coordinate distributed brain-wide neural activity with a specific spatiotemporal profile. Proc Natl Acad Sci U S A 113, E8306-E8315 (2016). https://doi.org/10.1073/pnas.1616361113

      (15) Chan, R. W. et al. Low-frequency hippocampal–cortical activity drives brain-wide resting-state functional MRI connectivity. Proc Natl Acad Sci U S A 114, E6972-E6981 (2017). https://doi.org/10.1073/pnas.1703309114

      (16) Leong, A. T. L. et al. Optogenetic fMRI interrogation of brain-wide central vestibular pathways. Proceedings of the National Academy of Sciences 116, 10122-10129 (2019). https://doi.org/10.1073/pnas.1812453116

      (17) Leong, A. T. L., Wang, X., Wong, E. C., Dong, C. M. & Wu, E. X. Neural activity temporal pattern dictates long-range propagation targets. Neuroimage 235, 118032 (2021). https://doi.org/10.1016/j.neuroimage.2021.118032

      (18) Leong, A. T. L., Wong, E. C., Wang, X. & Wu, E. X. Hippocampus Modulates Vocalizations Responses at Early Auditory Centers. Neuroimage 270, 119943 (2023). https://doi.org/10.1016/j.neuroimage.2023.119943

      (19) Wang, X. et al. Functional MRI reveals brain-wide actions of thalamically-initiated oscillatory activities on associative memory consolidation. Nat Commun 14, 2195 (2023). https://doi.org/10.1038/s41467-023-37682-8

      (20) Xie, L. et al. Brain-wide resting-state fMRI network dynamics elicited by activation of single thalamic input. Nat Commun 16, 11247 (2025). https://doi.org/10.1038/s41467-025-66104-0

      (21) Li, A., Gong, L. & Xu, F. Brain-state-independent neural representation of peripheral stimulation in rat olfactory bulb. Proc Natl Acad Sci U S A 108, 5087-5092 (2011). https://doi.org/10.1073/pnas.1013814108

      (22) Lang, J. et al. Odor representation in the olfactory bulb under different brain states revealed by intrinsic optical signals imaging. Neuroscience 243, 54-63 (2013). https://doi.org/10.1016/j.neuroscience.2013.03.057

      (23) Chery, R., Gurden, H. & Martin, C. Anesthetic regimes modulate the temporal dynamics of local field potential in the mouse olfactory bulb. J Neurophysiol 111, 908-917 (2014). https://doi.org/10.1152/jn.00261.2013

      (24) Wachowiak, M. et al. Optical dissection of odor information processing in vivo using GCaMPs expressed in specified cell types of the olfactory bulb. J Neurosci 33, 5285-5300 (2013). https://doi.org/10.1523/JNEUROSCI.4824-12.2013

      (25) Hudetz, A. G., Vizuete, J. A. & Pillay, S. Differential effects of isoflurane on high-frequency and low-frequency gamma oscillations in the cerebral cortex and hippocampus in freely moving rats. Anesthesiology 114, 588-595 (2011). https://doi.org/10.1097/ALN.0b013e31820ad3f9

      (26) Joliot, M., Ribary, U. & Llinas, R. Human oscillatory brain activity near 40 Hz coexists with cognitive temporal binding. Proc Natl Acad Sci U S A 91, 11748-11751 (1994). https://doi.org/10.1073/pnas.91.24.11748

      (27) Murthy, V. N. & Fetz, E. E. Oscillatory activity in sensorimotor cortex of awake monkeys: synchronization of local field potentials and relation to behavior. J Neurophysiol 76, 3949-3967 (1996). https://doi.org/10.1152/jn.1996.76.6.3949

      (28) Tallon-Baudry, C., Bertrand, O., Delpuech, C. & Pernier, J. Stimulus specificity of phaselocked and non-phase-locked 40 Hz visual responses in human. J Neurosci 16, 4240-4249 (1996). https://doi.org/10.1523/JNEUROSCI.16-13-04240.1996

      (29) Murakami, M., Kashiwadani, H., Kirino, Y. & Mori, K. State-dependent sensory gating in olfactory cortex. Neuron 46, 285-296 (2005). https://doi.org/10.1016/j.neuron.2005.02.025

      (30) Wilson, D. A. & Yan, X. Sleep-like states modulate functional connectivity in the rat olfactory system. J Neurophysiol 104, 3231-3239 (2010). https://doi.org/10.1152/jn.00711.2010

      (31) Schreck, M. R. et al. State-dependent olfactory processing in freely behaving mice. Cell Rep 38, 110450 (2022). https://doi.org/10.1016/j.celrep.2022.110450

      (32) Gao, P. P., Zhang, J. W., Chan, R. W., Leong, A. T. L. & Wu, E. X. BOLD fMRI study of ultrahigh frequency encoding in the inferior colliculus. Neuroimage 114, 427-437 (2015). https://doi.org/10.1016/j.neuroimage.2015.04.007

      (33) Gao, P. P., Zhang, J. W., Fan, S. J., Sanes, D. H. & Wu, E. X. Auditory midbrain processing is differentially modulated by auditory and visual cortices: An auditory fMRI study. Neuroimage 123, 22-32 (2015). https://doi.org/10.1016/j.neuroimage.2015.08.040

      (34) Goddard, E. & Mullen, K. T. fMRI representational similarity analysis reveals graded preferences for chromatic and achromatic stimulus contrast across human visual cortex. Neuroimage 215, 116780 (2020). https://doi.org/10.1016/j.neuroimage.2020.116780

      (35) Friston, K. J., Harrison, L. & Penny, W. Dynamic causal modelling. Neuroimage 19, 12731302 (2003). https://doi.org/10.1016/s1053-8119(03)00202-7

      (36) Friston, K. J., Kahan, J., Biswal, B. & Razi, A. A DCM for resting state fMRI. Neuroimage 94, 396-407 (2014). https://doi.org/10.1016/j.neuroimage.2013.12.009

      (37) Bernal-Casas, D., Lee, H. J., Weitz, A. J. & Lee, J. H. Studying Brain Circuit Function with Dynamic Causal Modeling for Optogenetic fMRI. Neuron 93, 522-532 e525 (2017). https://doi.org/10.1016/j.neuron.2016.12.035

      (38) Zeidman, P. et al. A guide to group effective connectivity analysis, part 1: First level analysis with DCM for fMRI. Neuroimage 200, 174-190 (2019). https://doi.org/10.1016/j.neuroimage.2019.06.031

      (39) Salimi, M. et al. Disrupted connectivity in the olfactory bulb-entorhinal cortex-dorsal hippocampus circuit is associated with recognition memory deficit in Alzheimer's disease model. Sci Rep 12, 4394 (2022). https://doi.org/10.1038/s41598-022-08528-y

      (40) Chen, Y. N., Kostka, J. K., Bitzenhofer, S. H. & Hanganu-Opatz, I. L. Olfactory bulb activity shapes the development of entorhinal-hippocampal coupling and associated cognitive abilities. Curr Biol 33, 4353-4366 e4355 (2023). https://doi.org/10.1016/j.cub.2023.08.072

      (41) Pellegrino, R., Sinding, C., de Wijk, R. A. & Hummel, T. Habituation and adaptation to odors in humans. Physiol Behav 177, 13-19 (2017). https://doi.org/10.1016/j.physbeh.2017.04.006

      (42) Sinding, C. et al. New determinants of olfactory habituation. Sci Rep 7, 41047 (2017). https://doi.org/10.1038/srep41047

      (43) Zhao, F. et al. fMRI study of olfaction in the olfactory bulb and high olfactory structures of rats: Insight into their roles in habituation. Neuroimage 127, 445-455 (2016). https://doi.org/10.1016/j.neuroimage.2015.10.080

      (44) Zhao, F. et al. fMRI study of the role of glutamate NMDA receptor in the olfactory adaptation in rats: Insights into cellular and molecular mechanisms of olfactory adaptation. Neuroimage 149, 348-360 (2017). https://doi.org/10.1016/j.neuroimage.2017.01.068

      (45) Sengupta, P. The Laboratory Rat: Relating Its Age With Human's. Int J Prev Med 4, 624-630 (2013). 

      (46) Quinn, R. Comparing rat’s to human’s age: How old is my rat in people years? Nutrition 21, 775-777 (2005). https://doi.org/10.1016/j.nut.2005.04.002

      (47) Burwell, R. D. & Amaral, D. G. Perirhinal and postrhinal cortices of the rat: interconnectivity and connections with the entorhinal cortex. J Comp Neurol 391, 293-321 (1998).

      (48) Kajiwara, R., Takashima, I., Mimura, Y., Witter, M. P. & Iijima, T. Amygdala input promotes spread of excitatory neural activity from perirhinal cortex to the entorhinal-hippocampal circuit. J Neurophysiol 89, 2176-2184 (2003). https://doi.org/10.1152/jn.01033.2002

    1. eLife Assessment

      This valuable study addresses two debates about how confidence is computed: whether it continues the accumulation process that produced the initial choice or reflects a separate one, and whether that accumulation stops at a time limit or an evidence boundary. The evidence is solid, combining systematic model comparison with an independent neural marker of evidence accumulation to adjudicate between models that behaviour alone cannot separate. Support for the boundary-based stopping rule is stronger than for the single-process account, as the neural comparison is qualitative rather than quantified and the alternatives tested cover only a narrow range of plausible two-stage architectures. The work will be of interest to researchers studying decision-making and metacognition.

    2. Reviewer #2 (Public review):

      Summary:

      Overall, the authors aimed to provide evidence that clarifies two debates within metacognition research concerning subjective confidence reports:

      (1) Does the post-decision confidence report arise from the same process that drives the initial decision, or does a separate, independent process support confidence computation?

      (2) How do we stop accumulating evidence for the post-decision confidence report? Is it based on a self-imposed time limit, or on accumulated evidence crossing a boundary?

      For the investigation, the authors constructed four models (2 × 2 factorial) to compare each combination of processes to account for random-dot motion tasks data with speed/accuracy manipulations. The models are generally embedded in the drift diffusion model framework, retaining basic parameters such as drift rate, boundary separation, starting point, and non-decision time, while adding linearly collapsing boundaries to model the initial choice. For the single vs. distinct process dimension, the difference lies in whether post-decision evidence accumulation is referenced to the endpoint of the initial decision process or restarts from a new, freely estimated starting point. For the time- vs. boundary-based stopping rule dimension, the key difference is that post-decision evidence accumulation stops either at a deadline sampled from a normal distribution or when the accumulated evidence hits a collapsing boundary.

      Based on model comparison, the boundary-based stopping rule clearly outperformed the time-based stopping rule. However, models with the boundary-based stopping rule performed similarly regardless of whether a single or distinct process was used. Here, the authors drew additional insights from EEG recordings during the task, focusing on the centro-parietal positivity (CPP), which has been proposed as a neural correlate of the evidence accumulation process. By simulating evidence accumulation trajectories (with additional assumptions) and comparing the patterns of those trajectories with observed ERP waveforms, the authors argued that the single-process model provided a better match to the CPP findings and was therefore preferred. This was specifically demonstrated by the model's superior ability to match the pre-response CPP amplitude differences conditioned on the post-decision confidence-related variables.

      Strengths:

      (1) The authors translated existing theories into computational models of decision-making and systematically compared different cognitive processes by assessing model fits to the data. This provides strong evidence supporting the idea that post-decision confidence reports could be better explained by boundary crossing rather than a self-imposed deadline to respond.

      (2) Beyond model evidence, an important result is that CPP amplitude predicted confidence before the initial choice was reported, which is a unique prediction of the single-process model. The use of EEG as an independent validation measure provided additional evidence in favour of this model.

      (3) Combining points 1 and 2, this study successfully addressed the two key debates with solid evidence to favour one theory over another.

      (4) Another strength of this study is the data quality. The high number of trials provided a strong foundation for model inference as well as ERP analysis. The experiment also contained a speed-accuracy manipulation to evaluate model performance across diverse situations.

      Weaknesses:

      I have two main concerns around the modelling work and neural analyses, which in my opinion could have limited the interpretation of the findings. My responses here will be lengthier, but this reflects the nature of the modelling work rather than implying stronger criticisms than those suggested by the strengths discussed above.

      (1) There are a few assumptions in the models that lack psychologically meaningful interpretations, and this study placed more effort into model comparison while lacking discussion of the cognitive processes inferred from parameter estimates.

      To start, I think some of the parameterisations were not properly justified. For the boundary models, it is not very clear why the upper and lower boundaries were different and collapsed at different rates for confidence decisions, given that a single boundary parameter and collapse rate were used for the initial decision. This allows more flexible shifts in the model's predictions of confidence ratings without strong justification. Specifically, it is unclear why the boundary-single model has such an implementation while the boundary-distinct model was only equipped with one boundary parameter (a2, compared to a2up and a2down).

      Similarly, the inclusion of metacognitive noise creates another layer of flexibility in the predictions of confidence ratings. In most existing evidence accumulation models with a diffusion process, noise comes from two sources: within-trial noisy evidence accumulation and across-trial variability (e.g., drift rate variability). Beyond these, such models almost always assume that the decision is made deterministically once the evidence reaches a specific boundary. The inclusion of metacognitive noise here sounds more like a noisy decision-to-action mapping.

      I also have similar doubts about allowing the non-decision time parameter for confidence accumulation in the distinct model to be negative. The authors argued that confidence accumulation may begin during initial evidence accumulation. However, this is a flawed implementation, as the non-decision time was simply added to the evidence accumulation time rather than being incorporated within it. Allowing negative non-decision times may achieve similar predictions, but it is ad hoc.

      The inclusion of a collapsing boundary mechanism in the post-decision confidence accumulator helped the model reach more diverse levels of accumulated evidence and ultimately improved predictions of confidence ratings. However, no strong argument is presented for this implementation beyond the observation that the model performs worse without it. The collapsing boundary mechanism has traditionally been interpreted as reflecting a sense of urgency. For the boundary models, I noticed that the collapse rate of the upper boundary differed significantly between speed and accuracy conditions, which is consistent with the urgency interpretation. Overall, I would like to see more discussion of the specific model mechanisms included by the authors, interpreted in light of parameter estimates.

      (2) While the ERP findings provided external evidence and validation of the modelling results, I find the simulation practices not particularly useful and potentially misleading for naive readers. Specifically, the authors attempted to draw a parallel between patterns of simulated evidence accumulation traces and observed CPP waveforms. While the CPP has received support as a correlate of the evidence accumulation process, the DDM is by no means a neural model capable of generating predictions of neural observations. To my understanding, the superior fit of the boundary-single model was primarily due to the fact that pre-response CPP amplitude predicts post-decision confidence ratings. Therefore, as the boundary-distinct model did not connect the two phases of evidence accumulation, it would fail to account for this observation. I think this point could be clearly demonstrated without the need to introduce additional assumptions into the model simulations in order to directly compare averaged trajectories with averaged ERP waveforms. While the authors did not explicitly claim otherwise, this approach creates an illusion that the model can mechanistically account for ERP data. I would like the authors to provide explicit clarification on this point.

      Appraisal:

      Overall, the authors have provided solid evidence in support of their research aims. The findings contribute to longstanding debates with insights from model mechanisms and neural findings that should not be overlooked by future studies on this topic. This study also offers a good starting point for future model development and refinement in broader contexts of confidence reporting, such as paradigms involving simultaneous initial decisions and confidence judgements. The high quality of the behavioural and EEG data will make a valuable contribution to future research.

    3. Reviewer #1 (Public review):

      Summary:

      A central question in decision-making is whether confidence arises from the same evidence-accumulation process that led to the choice or from a separate second process. The manuscript addresses this question using a random-dot motion (RDM) task, with an initial choice followed by a confidence report, and a time-pressure manipulation on the confidence report. The authors fit a family of four models differing on two dimensions: the source of confidence (a continuation of the choice accumulator vs. a distinct accumulation process) and the stopping rule (time-based vs. boundary-based). The models are fitted to behavior, and then their simulated dynamics are compared with the CPP signal associated with evidence accumulation (not included in the fit). The main methodological contribution is the use of a neural signal to decide between the two boundary-based models, which are nearly indistinguishable behaviorally. The authors conclude that boundary-based stopping rules outperform time-based rules, and that the Boundary-Single model reproduces the certainty-related CPP dynamics better than the Boundary-Distinct model, supporting a single accumulation process for both choice and confidence.

      Strengths:

      The main strength is methodological: using a neural signal (CPP) as an out-of-sample arbiter between the two boundary-based models (which are nearly equivalent behaviorally). This addresses the model-identifiability problem: when behavior does not distinguish between competing models, a neural signal not included in the fit can provide external evidence.

      The work is also thorough and empirically rigorous. The finding that CPP amplitude predicts subsequent certainty ratings several hundred milliseconds before the initial choice response is statistically supported and interesting in its own right, independent of its interpretation. The small-N/many-trials design (2,160 trials per participant) is well suited to capturing change-of-mind trials and is accompanied by thorough model and parameter recovery, with the generating model recovered in most simulations. The study also includes preregistration of the design and planned behavioral and neural analyses, and it replicates a previously reported behavioral pattern.

      Weaknesses:

      The central conclusion may well be correct, but in my view the current evidence does not fully support it. The behavioral comparison between the competing models did not resolve the issue and in fact showed a slight preference for the distinct-process model, so the weight of the decision falls mainly on the neural comparison.

      (1) The neural evidence supports access to pre-choice variation, but does not necessarily establish a single continuous process. The critical difference between the single and distinct models is the presence of trial-to-trial variation in the evidence at choice commitment, which is inherited by the confidence process: the single model preserves it (the confidence process begins from the trial-specific endpoint of the choice DV), while the distinct model does not inherit it (z2 fixed). Thus, the neural test examines whether confidence has access to the state of evidence accumulation before the choice response but does not establish that the same accumulation process must continue seamlessly to determine confidence. The authors also test a model in which a distinct post-choice process is initialized using information from the endpoint of the choice process. However, its failure rules out one specific implementation of information transfer, rather than the broader class of two-stage models in which a distinct confidence process receives a readout of the decision state and may additionally integrate other metacognitive cues.

      (2) A broader class of two-stage metacognitive models that combine a decision-state readout with additional cues is not tested. The distinct-process model implemented here captures only a limited subset of possible metacognitive architectures. Its starting point is independent of the choice-DV endpoint, and it remains driven by the available sensory evidence, potentially within a different reference frame. It does not capture second-order accounts in which confidence combines a readout of the decision state with additional cues not explicitly represented in the choice accumulator, such as response time, motor conflict, subjective stimulus clarity or attention. Moreover, the paradigm used in the manuscript provides few independently manipulated information sources that would allow such a process to be identified separately from the choice accumulator.

      (3) The neural distinction between models is not quantitatively evaluated. The claim that the Boundary-Single model better reproduces the CPP rests primarily on a visual/qualitative comparison, without a numerical measure of the discrepancy between each model and the neural data. Although the plotted β coefficients (Figure 4D) provide estimates of the neural effects, no scalar summary of model-to-CPP fit is reported. Because the neural comparison carries much of the inferential weight, a quantitative comparison would strengthen the conclusion substantially.

      (4) The fitted parameters raise a question about the post-choice process. Confidence responses were very fast (average median confidence RT = 243 ms), with no minimum RT threshold. The Boundary-Single model estimated a post-choice drift rate more than twice the pre-choice rate (1.69 vs. 0.75). In the model, confidence accumulation begins at commitment, before the initial response is executed. The measured confidence RT therefore does not capture the full accumulation window, which also includes the motor delay. Still, the sharp rise in drift rate at commitment requires an explanation. It may reflect stronger weighting of the still-available evidence, as the authors suggest. But it is also consistent with a fast readout of a decision state that was largely set before the initial response. The authors could compare the current model against a readout model, or against an intermediate variant that permits only a brief, bounded period of post-choice accumulation.

      (5) The participant-level distribution of model preferences would clarify the comparison. Models were fit separately per participant and condition, but the comparison is summarized as mean BIC (a small average preference for Boundary-Distinct). A mean cannot distinguish two different situations: a consistent, weak preference for one model across all participants, versus a mixture in which some participants clearly favor one architecture and others the opposite. Reporting the distribution of per-participant ΔBIC and the number of participants favoring each model would clarify the result.

    1. eLife Assessment

      The work presented represents important new evidence linking the cascade of neural processes triggered by memory-based prediction errors. The study uses an impressive collection of approaches and methods to characterize and measure cognitive control, arousal, and memory changes as a function of memory-based violations. Analyses are technically sophisticated and rigorous and, taken together, provide solid evidence that there are multiple processes accompanying prediction errors, and that they differentially relate to successful encoding.

    2. Reviewer #1 (Public review):

      This manuscript describes a multi-modal study of associative learning and memory in humans, that combines scalp EEG, pupillometry and behavioral analysis to explore the construct of mnemonic prediction errors (MPEs), in terms of their relationship to attention and cognitive control. Across two pooled studies, participants performed associative memory tasks in which they learned the relationship between a cue word (action verb) and subsequent picture (animate or inanimate) with a strong vs. weak (4 or 1 repetitions) encoding manipulation. At test, participants were encouraged to generate a prediction following the cue word to determine whether the subsequently presented picture was a match or mismatch. The timecourse of pupillary responses during match decisions were decomposed using temporal principal components analysis, which identified 6 distinct and overlapping processes. Some of the components (PC3/PC4) exhibited sensitivity to both the strength and mismatch conditions, as well as behavior (both RT and accuracy) and retrieval success on the subsequent trial. Furthermore, relationships were also observed between pupillary responses (specifically for PC4) and both frontal theta and posterior alpha power measures obtained from scalp EEG in Experiment 2, as well as for frontal theta and subsequent learning from mismatch stimuli (assessed using subsequent memory findings from a surprise recognition test). The authors suggest the findings indicate that MPEs elicit changes in attention, arousal and cognitive control which impact subsequent learning.

      Strengths:

      This manuscript has many strengths, including a clever study design, thoughtful integration of multiple neurocognitive measures, and a set of rigorous and technically sophisticated analyses, which reveal a large set of relationships among the measures and behavior. The findings demonstrating brain/physiology-behavior relationships are particularly important, in that they point to potential functional consequences of MPEs.

      Weaknesses:

      The technical proficiency and complexity of the study and analysis also presents a clear limitation and challenge for interpretation. It is likely that readers, even those that are quite knowledgeable about the methods, constructs, and questions being addressed will often struggle (as this reviewer did) to keep the large set of findings in mind and gain understanding of how they all fit together.

      Indeed, it seems like there many threads running together in the paper which make it challenging to find the through-line of the key findings. The authors do address some of the key questions motivating the paper in the Introduction, but the results are somewhat ambiguous with regard to the primary question of the study as to whether the detection of MPEs leads to interaction among cognitive control, attention, and arousal. To their credit, the authors tackle this question through both cross-correlation and formal mediation analyses, and summarize these in diagrammatic figures (Figure 3, Figure 6). Yet it is not resolved whether the results represent a clear answer pointing to independence, or rather a lack of statistical power, or ill-resolved formulation of the mediational relationship. In particular, the cross-correlation suggests that posterior alpha suppression in response to MPEs does precede frontal theta, yet this indirect relationship does not explain the variation in trial-by-trial RTs on mismatches. This suggests a potential model misspecification.

      In addition to the primary interaction issue mentioned above (between cognitive control, attention & arousal), the Introduction lays out a number of claims: 1) that pupil size will be more sensitive to strong than weak MPEs; 2) that MPE-linked increases in attention (indexed with posterior alpha suppression) and arousal (indexed with pupil size) will be linked to learning; and 3) MPE learning will vary as a function of prediction strength. Given the focus on learning, it is somewhat surprising that learning is not included in the mediation models. As the authors indicate in the Discussion, the use of trial-by-trial RT variation to drive the mediation model might be problematic, given that the RTs are sensitive to a range of factors beyond mnemonic prediction strength and also are under competing pressures (longer for mismatches than matches, due to surprise-linked slowing, but also faster following stronger rather than weaker mnemonic predictions). Thus, an alternative possibility might be to use trial-by-trial recognition of mismatches as the outcome variable in mediation models rather than trial-by-trial RT as the independent variable.

      A large component of the results (Sections 2 and 3) is devoted to analyses of cue-linked pupil and EEG processes that putatively reflect mnemonic predictions (i.e., occurring before picture probes are presented and match/mismatch detection, i.e., MPEs occur). Yet these Results and the subsequent pupillary PCA components (PC1 and PC5) that are elicited are not well-integrated with the primary themes of the paper or the causal hypotheses. One finding that does seem to figure prominently (in that it is mentioned in Abstract, Introduction & Discussion) relates to the amount of attention allocated to the mnemonic prediction generation. Yet this finding is not well emphasized in the Results themselves. Possibly it refers to the negative relationship between posterior alpha during memory retrieval and the magnitude of pupillary PC3 component, described in Section 3. But it was quite challenging to identify amongst the wealth of results described in this Section as well as the others. More generally, the large amount of findings described across all four lengthy Results sections makes it challenging for readers to discern what are the key ones that the authors would like to highlight.

      It is recommended that the authors do another pass through the paper to better highlight the most critical findings that they want to emphasize or which are most interpretable from a mechanistic and causal flow perspective and then de-emphasize or move other findings to the Supplemental Materials. Although the authors are to be commended for such a rigorous and comprehensive set of analyses, there are so many of them and findings, that the key points get buried and the reader needs to struggle potentially unnecessarily to identify the key take-away points.

    3. Reviewer #2 (Public review):

      Summary:

      The authors studied cognitive control and attention in response to mnemonic prediction errors (MPEs): situations in which the external reality violates internal memory-based predictions. The behavioral task first established strong versus weak predictions, and then either confirmed or violated these predictions. The authors examined markers of cognitive control (frontal theta) and attention (posterior alpha suppression, pupil response) while strong and weak predictions were confirmed or violated. They found increased cognitive control (frontal theta) for strong MPEs, which correlated with subsequent memory. Markers of attention (alpha suppression, pupil response) also accompanied strong MPEs but did not correlate with subsequent memory. Pupil response was investigated using an interesting approach that decomposes the response into different components, finding that different components respond earlier or later and show different correlations with MPEs and their strength. The authors also investigated how EEG, reaction time, and pupil responses correlated with one another, providing further insight into the mechanism underlying the response to MPEs. Together, the study points toward multiple control and attention mechanisms involved in MPE response and memory.

      Strengths:

      The study has a clear behavioral paradigm with multiple measures - behavioral, EEG, and pupillometry that offer an investigation into different aspects of MPE response and memory.

      The study is also very comprehensive in looking at multiple phases in processing MPEs: the prediction phase (prior to the violation), the response to MPEs, and subsequent memory of MPEs, all within one study. Specifically, the link between neural mechanisms and subsequent memory is a major advancement, as most prior studies did not include this component. Mechanisms underlying subsequent memory of MPEs are theoretically important, as a primary function of MPEs is to promote learning and memory. As the authors mention, the different neural and pupillary signals are not robustly correlated, suggesting multiple mechanisms underlying MPE detections, which is interesting, offers avenues for future research, and can facilitate a better theory of how MPEs are processed in the brain. Finally, the decomposition of pupil response into different components and their correlation with behavior (RT during match/MPE detection) is interesting.

      Weaknesses:

      The methods are rigorous, and the data support the claims. The weaknesses are minor and are offered here as avenues for future research.

      (1) The relationships the authors find between brain measures and pupil components were largely not specific to mismatches/matches. Thus, the specificity of this relationship is untested.

      (2) The results with subsequent memory are important and address a major gap in the field that largely did not relate neural effects of MPE to subsequent memory. However, one major limitation of the study is that the authors did not test memory for matches. I understand the logic of avoiding testing matches. Because matches were repeated more times in the study, it's not a fair comparison and could change participants' overall criterion for old/new decisions. Future research could address this, e.g., by testing weak matches or potentially using a between-subject design.

      Comments on revised version.

      The authors addressed all my concerns. I appreciate the authors' thoughtful and detailed response.

    4. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Strengths:

      This manuscript has many strengths, including a clever study design, thoughtful integration of multiple neurocognitive measures, and a set of rigorous and technically sophisticated analyses, which reveal a large set of relationships among the measures and behavior. The findings demonstrating brain/physiology-behavior relationships are particularly important, in that they point to potential functional consequences of MPES.

      We thank the reviewer for noting these strengths of the work along with the below encouragement to revise the manuscript to better highlight the key findings and their implications.

      Weaknesses:

      The technical proficiency and complexity of the study and analysis also present a clear limitation and challenge for interpretation. As a reader, even those who are quite knowledgeable about the methods, constructs, and questions being addressed will often struggle (as this reviewer did) to keep the large set of findings in mind and gain an understanding of how they all fit together.

      Indeed, it seems like there are many threads running together in the paper, which makes it challenging to find the through-line of the key findings, or to understand how they might relate to some pre-existing hypotheses, rather than merely interesting patterns detected in the data. In the Introduction and Discussion, it seems as if the key question is to understand the pathways by which MPEs impact cognition, but this is a rather broad topic, so it is not clear exactly what the authors are aiming at with this question and study design.

      As an example, authors operationalize frontal theta power as an index of cognitive control demand, and one of the pathways by which MPEs impact cognition. But this point becomes somewhat circular, since it is not clear how or why the Mismatch x Strength interaction in frontal theta reflects that demand. It would have been better to set this pattern up in the Introduction as a theoretically driven hypothesis, since it currently appears more like a post-hoc interpretation. This is mirrored by how the issue is first brought up in the Introduction, where it states somewhat vaguely: "whether MPEs are followed by an increase in frontal theta... warrants closer examination".

      Again, we appreciate the reviewer’s thoughtful feedback on where the manuscript can be clearer, especially given the rich set of results it reports. Following the reviewer’s guidance, we restructured and revised the Introduction to further motivate the hypotheses that (a) MPEs increase both attention/arousal (grounded in studies of Event Segmentation Theory) and cognitive control (given findings on reward prediction errors and other types of prediction errors), and (b) there are greater increases in these processes triggered by strong compared to weak MPEs. To better link these hypotheses to resulting statistical tests, we note that hypothesis (a) was tested in our trial-level regression models in Fig. 1 by examining main effects of Strength, whereas hypothesis (b) was tested in our models in the Mismatch x Strength interactions. On point (b) and potentially circularity, we note that previous work indicates that frontal theta scales with negative reward prediction errors; as such, we hypothesized that stronger MPEs would elicit more frontal theta, as evidenced by a robust Mismatch x Strength interaction during the probe period.

      Later in the results, there are findings relating frontal theta to pupil dilation, posterior alpha suppression and then subsequent memory. It was hard to understand how all the findings might be linked together functionally or conceptually. Are the authors potentially postulating a mediating or mechanistic pathway, in which the MPE leads to increased cognitive control (frontal theta), which then leads to enhanced subsequent memory of those events? If this is the case, then maybe a formal path analysis would be the best way to test or state this hypothesis. It would also be useful to specify more clearly how the pupil components and alpha suppression factor into this mediating path, since it was not clear.

      Relatedly, the authors suggest that internal attention and arousal also play relevant roles in this pathway, but these are also not clear. In some cases, it is stated as if this is a distinct pathway from the cognitive control one, since there is a focus in the results on the independence of frontal theta and posterior alpha, but elsewhere they seem to be treated as two aspects, or distinct steps, within a single pathway. Again, these different threads of the findings were quite challenging for the reader to follow. Pathway analyses, such as with multiple mediation or moderated mediation, could be a useful way to address this question. For example, it seems as if readiness-to-remember is another behavioral outcome (like subsequent memory) that could be used in the search for mediators.

      We thank the reviewer for highlighting these ambiguities in the original submission and for the thoughtful encouragement to leverage mediation models to more formally test the hypothesized relationships between MPE magnitude and constructs of control, attention, and arousal. In the revision, we now more clearly hypothesize that the effects of strong MPE-driven increases in attention and arousal might be explained, in part, by cognitive control (as indexed by frontal theta) upregulating attention and arousal. To more explicitly test this model of the relationships between our measures, as recommended by Reviewer #1, we now include multivariate mediation analyses to assess whether, at a trial-level, changes in posterior alpha and the immediate pupil effect PC3 are explained in part by increases in frontal theta. Because changes in posterior alpha following MPEs and PC3 scores did not predict subsequent memory, mediation analyses addressing the hypothesis that our attention/arousal measures mediate the effect of frontal theta on subsequent memory were not conducted. Following insights from a cross-correlation analysis, as recommended by Reviewer #2, we tested an additional model to examine whether the effects of MPE magnitude on frontal theta were explained in part by changes in posterior alpha. Examination of the posterior distributions of the indirect effects did not favor our hypothesized model, nor the alternative model that attention upregulates control. Altogether, these outcomes suggest that increases in control, attention, and arousal following strong MPEs may be elicited independently. Yet, we also note that the current set of experiments may be underpowered for these mediation analyses; future work can further investigate the directionality between these effects. Finally, with respect to the readiness-to-remember findings, they were removed in the interest of space, as recommended by Reviewer #2.

      At the minimum, it would be quite helpful to have diagrammatic figures that specify the hypothesized and observed relationships between independent variables (Strength, Mismatch), physiological indices (pupil dilation components, frontal theta, posterior alpha) and key outcome measures (accuracy, RT, next-trial retrieval success, subsequent memory), so that the reader can refer back to them as each component of the analyses is conducted.

      To further increase conceptual clarity, we also followed this helpful suggestion, adding diagrammatic figures to illustrate our hypotheses regarding interactions between the effects (Fig. 3a) and to summarize the observed relationships between MPEs, control, attention, and arousal (Fig. 6).

      Minor Points:

      Many figures had x-axes showing a pupil component or EEG power metric broken down by quartile or quintile. Yet nowhere is it ever explained why this graphical (or analytic?) approach is used and what it reflects, or how it is decided which break down to use (quartile/quintile). If the data are analyzed as a correlation, why is a scatterplot not shown instead?

      In the linear mixed effects models, continuous values were used to assess relationships between variables. In the figures, the continuous variables were binned into quartiles or quintiles for ease of visualization. We opted to visualize the data using this approach, rather than with a scatterplot, given the large number of trials. We updated the figure captions to clarify the approach.

      It was surprising that, unlike readiness-to-remember, which was analyzed via logistic regression and odds-ratio, subsequent memory was not analyzed in the same fashion (i.e., as a binary outcome variable predicted by frontal theta), rather than in a reverse chronological one (subsequent memory predicting frontal theta). Historically, it was the case that subsequent memory was analyzed in this manner, but that was before the era in which trial-level linear mixed-effect models were in wide usage, as they are implemented in this study. Thus, the choice seems like a wasted opportunity or a step backwards analytically.

      We thank the reviewer for this encouragement and agree with the point. In the revision, we note that the readiness-to-remember results were removed in the interest of space and clarity, as was recommended by Reviewer #2. With respect to the subsequent memory analyses, they are now analyzed via logistic regression.

      Reviewer #2 (Public review):

      Strengths:

      The study has a clear behavioral paradigm with multiple measures - behavioral, EEG, and pupillometry that offer an investigation into different aspects of MPE response and memory.

      The study is also very comprehensive in looking at multiple phases in processing MPEs: the prediction phase (prior to the violation), the response to MPEs, and subsequent memory of MPEs, all within one study. Specifically, the link between neural mechanisms and subsequent memory is a major advancement, as most prior studies did not include this component. Mechanisms underlying subsequent memory of MPEs are theoretically important, as a primary function of MPEs is to promote learning and memory. As the authors mention, the different neural and pupillary signals are not robustly correlated, suggesting multiple mechanisms underlying MPE detections, which is interesting, offers avenues for future research, and can facilitate a better theory of how MPEs are processed in the brain. Finally, the decomposition of pupil response into different components and their correlation with behavior (RT during match/MPE detection) is interesting.

      We thank the reviewer for noting these strengths of the work along with the below encouragement to revise the manuscript to better highlight the key findings and their implications.

      Weaknesses:

      The methods are rigorous, and the claims are mostly supported by the data, but there are a few weaknesses or places that could be improved:

      (1) The authors conduct PCA analysis to identify different components of the pupillary response to MPE and relate them to behavior. Specifically, the authors identify components PC3 and PC4, which they interpret as related to MPE. However, some parts of the interpretation could be clearer or better justified:

      (a) The authors refer to PC4 as "post-decision cognitive processing". But, given that RT was between .5-.7s, and PC3 peaked after more than 1s, wouldn't it be cautious to interpret PC3 as postdecision as well?

      Thank you for raising this point. Given that pupil is a relatively sluggish response, it is possible that both components reflect post-decision cognitive processing, even if PC3 peaks before PC4. Following the reviewer’s guidance to adopt more cautious language, we replaced “post-decision” with “post-MPE”.

      (b) MPEs overall elicit longer RTs in this study, suggesting that long RT is a behavioral marker of MPE. Nonetheless, the authors argue on p. 12: "Altogether, these findings indicate that when stronger mnemonic predictions (as indexed by shorter RTs) were violated." And, PC3 is correlated with shorter RTs for mismatches, meaning that behaviorally, these trials were more similar to matches. Thus, how do the authors interpret shorter versus longer RTs for MPEs, and what processes do these RT reflect?

      We thank the reviewer for stressing the need for greater clarity regarding the relationships between RT and the constructs of interest. With respect to RT, we interpret the condition-level difference in RTs between mismatch and match trials as a behavioral marker of an MPE. However, when comparing mismatch trials within a given strength condition to each other (i.e., an analysis at the trial-level), shorter RTs may reflect a stronger prediction, greater certainty that the probe is a mismatch, and therefore the experience of a stronger MPE. Note that while larger PC3 scores were associated with shorter RTs for mismatches (Fig. 2b, right), the mismatch RTs in the largest PC3 quartile were still longer than those on match trials in the corresponding quartile (in other words, there was still a condition-level difference in RTs in the largest PC3 quartile, suggesting that the mismatch trials in this bin are behaviorally still likely to be different from match trials in the corresponding bin).

      To clarify these relationships and our interpretation, we modified the referred to text: “(as indexed by shorter strong mismatch RTs).” Moreover, we added text to the Discussion, further delineating our reasoning here and the implications of our findings for understanding the mechanisms giving rise to and triggering by MPEs of varying strengths. This includes adding an explicit summary of the logic and findings that notes that, at the trial-level, shorter RTs may reflect stronger predictions; at the match/mismatch condition-level, longer RTs may reflect the experience of a MPE. For PC3, shorter RTs (trials with a stronger prediction) in the Strong Mismatch condition were associated with a larger pupillary response. The added text notes that “while longer mean RTs for mismatches compared to matches are a behavioral marker of a MPE, within-condition differences in RTs (i.e., between mismatch trials) may reflect more subtle differences in MPE magnitude, with shorter RTs reflecting stronger predictions and thus stronger MPEs. Strong mismatch RTs were used in the mediation models as a proxy measure of MPE magnitude; this estimate may be noisy because RTs in this experiment are likely sensitive to factors independent of mnemonic prediction strength (e.g., preparatory attention (Supplementary Fig. 6) or memory strength of the mismatch probe). This limitation may have additionally reduced sensitivity to detecting indirect effects.”

      (2) The brain to pupil relationship (p. 13-14): If I understand correctly, this was done on a trial-by-trial basis, but the high temporal resolution allows doing the analysis in a time-resolved manner - does brain activity at a certain time point preceding/following the pupil response correlate with the pupil response? It might be that cognitive control influences attention mechanisms or vice versa (because there is some overlap in the response). Although not testing causality, this temporally resolved correlation would be an interesting way to start probing how signals might influence each other.

      Thank you for this suggestion. We now report a cross-correlation analysis (Fig. 3d) that suggests that cognitive control increases precede attention decreases at retrieval, whereas in response to a strong MPE, cognitive control increases follow attention increases. There were no significant clusters for the temporal relationships between frontal theta and pupil, nor for posterior alpha and pupil.

      (3) The relationships the authors find between brain measures and pupil components were largely not specific to mismatches/matches. However, are they specific to this task? I think it would benefit the paper to show that these relationships are potentially specific to making match/mismatch memory decisions, versus, e.g., any stimulus processing. For example, the authors could run the same analyses locked to stimuli in the study phase, anticipating a different pattern, if indeed these findings are specific to the associative memory task.

      Many of the associations between our measures indeed did not show an interaction with Mismatch. Due to jitter in the ISI in the study phase, some of the analyses in the retrieval phase cannot be performed in the exact same way for the study phase. We will leave these questions to be addressed in future research. We agree that an important question for future research is to address whether these responses depend on making match/mismatch decisions and now include consideration of this point in the Discussion: “Finally, the magnitude of observed increases in control, attention, and arousal following strong MPEs may be influenced by the decision-making process engaged when making match/mismatch judgments. Not all MPEs necessitate behavioral responses. Whether similar magnitudes in neurocognitive responses and consequences for learning are observed upon detection of an MPE, but in the absence of a decision remains unclear.”

      (4) During memory retrieval (i.e., before the probe), the authors find that frontal theta, a marker of cognitive control, was associated on a trial-by-trial basis with more posterior alpha (i.e., less alpha suppression, potentially reflecting less attention), and that this association was stronger for weaker predictions. The authors interpreted this as weaker predictions necessitating more cognitive control, and that more cognitive control was recruited specifically in trials where retrieval included less content (memory reinstatement) to attend to. Generally, cognitive control is recruited to facilitate memory retrieval. If so, one possible interpretation is that this correlation reflects cognitive control effort that has failed to produce enough memory reinstatement. The other possibility is that this correlation reflects more specific retrieval of the correct probe, without retrieval of interfering items (i.e., overall less content). I believe that the former explanation predicts that this correlation would be associated with longer RTs (more difficult decisions), while the latter predicts shorter RTs (easier decisions due to successful retrieval), at least for matches.

      Thank you for these insightful comments. Because this analysis is not key to the main questions about MPEs and given both reviewers’ concerns that the manuscript can be overwhelming for the reader given the sheer number of findings reported, we opted to move this point from the main text to the Supplement. However, following the reviewer’s guidance here, we conducted the proposed analyses and tested these alternative accounting by modeling RTs. The results favour the former interpretation:

      “Greater control being associated with less attention could reflect failure in controlled retrieval efforts to reinstate sufficient memory evidence of the probe and thus fewer retrieval products to which attention is allocated. Alternatively, greater cognitive control could increase the likelihood of retrieval success, eliciting selective retrieval of the correct probe and inhibition of interfering items. To address these alternatives, we examined how the association between frontal theta and posterior alpha related to the difficulty of a trial, as assayed by RTs. The former failure-of-control account would predict that a stronger positive frontal theta- posterior alpha association would relate to longer RTs, whereas the greater-retrieval-specificity account would predict that a stronger positive association would relate to shorter RTs. In a model predicting RTs as a function of frontal theta, posterior alpha, Strength, Mismatch, and their interactions, we found a two-way frontal theta × posterior alpha interaction (β=0.016, CI=[0.001, 0.031], p=0.033), such that a stronger positive association between frontal theta and posterior alpha predicted longer RTs. This relationship did not differ as a function of Strength (no frontal theta × posterior alpha × Strength interaction: β=-0.015, CI=[-0.033, 0.004], p=0.124), Mismatch (no frontal theta × posterior alpha × Mismatch interaction: β=-0.016, CI=[-0.050, 0.018], p=0.342), or interact with Strength and Mismatch (no frontal theta × posterior alpha × Strength × Mismatch interaction: β=0.024, CI=[-0.014, 0.062], p=0.218). Together, these outcomes support the idea that during memory retrieval, positive coupling between frontal theta and posterior alpha may reflect failure or inefficiency of cognitive control efforts to rapidly accumulate mnemonic evidence to which to attend in support of a memory decision.”

      (5) In section 3, the authors found a positive relationship between alpha during memory retrieval and PC3 during MPE. If I understood correctly, this means that less attention during retrieval (less suppression) is correlated with a stronger PC3 response. How do the authors interpret this? Maybe along the same lines as in (5), specifically retrieving the correct information (i.e., less retrieved content to attend to) means a stronger prediction, leading to a stronger MPE, and a stronger MPE response, as reflected by PC3?

      We appreciate this comment, as it highlights a need for greater clarity here. The observed relationship was actually negative (Fig. 4f), meaning that more attention during retrieval was associated with a stronger PC3 response. We suspect that the lack of clarity here may be due to the original statement that “there was a positive relationship between posterior alpha suppression during memory retrieval [and PC3 scores]”. To increase clarity, we have modified this statement to “there was a negative relationship between posterior alpha during memory retrieval [and PC3 scores]”. We interpret the negative relationship with posterior alpha (i.e., positive relationship with posterior alpha suppression) to indicate “that greater attentional allocation during memory retrieval, which occurs when memories are stronger and more retrieval products can be reinstated and attended to (Fig. 1e; Fig. 4c), predicts the magnitude of immediate pupil responses to MPEs.”

      (6) The results with subsequent memory are important and address a major gap in the field that largely did not relate neural effects of MPE to subsequent memory. However, one major limitation of the study is that the authors did not test memory for matches. I understand the logic of avoiding testing matches. Because matches were repeated more times in the study, it's not a fair comparison, and could change participants' overall criterion for old/new decisions. However, one possibility would have been to test only the weak prediction; this could have given some specificity to the neural subsequent memory findings.

      We were indeed concerned about the change in decision criterion and did not include match items for this reason. Nonetheless, this is a useful suggestion that future work could include the weak match items for a better comparison of subsequent recognition memory. We now comment on this in the Discussion: “Whether control-associated enhancements in learning are specific to learning from MPEs can be further tested in future work by including a test of subsequent memory for weak match probes as an additional control condition for comparison.”

      (7) The authors nicely characterized the different PC of pupillary MPE response. But, with respect to subsequent memory, they only present pupil size. Unless there is some methodological reason that prevents testing subsequent memory on the PC, I think this will be very informative about the potential mechanisms underlying memory of MPE.

      The pupil PCs were not associated with subsequent memory, though there are some interesting trends in the 48-delay condition which could be explored in future work. These findings are now reported in the Supplement (Supplementary Fig. 16).

      (8) This paper includes many interesting findings, and I am not sure how they all come together into a cohesive mechanistic understanding of MPE response and subsequent memory. I think the paper would benefit from either a conceptual mechanism figure or, in the Discussion, have a summary of a proposed mechanism integrating the findings together.

      We thank the reviewer for stressing this point, which also was raised by Reviewer #1. To better emphasize the novel contributions of the work and to assist the reader’s understanding of the key findings, we: (a) revised the Introduction to more explicitly describe our hypotheses about the relationships among our measures as they relate to MPE responses and subsequent memory; (b) now report mediation models to more directly test these relationships and include diagrams of the hypothesized relationships; and (c) include a schematic figure at the end of the Results to highlight the key mechanistic relationships supported by the data.

      (9) Relatedly, the section "Immediate, strength-sensitive neurocognitive impacts of MPEs" does not link the arguments to specific data points, so it's hard to follow which data specifically the authors are interpreting.

      The discussion in this paragraph rests on the Strength × Mismatch interactions observed in Figures 1d-f, summarized in the first sentence of the section. To increase clarity, we changed the title of this section to “Neurocognitive impacts of MPEs are strength-sensitive”.

      (10) If I understand correctly, the authors did not find improved memory for strong compared to weak MPE. First, I think this behavioral result should be incorporated in the main paper and in the interpretation of the results. Second, given that the neural effects the authors tested either correlated with memory for strong MPE or did not show a relationship with memory, what neural/pupil response could explain memory for weak MPE?

      Thank you for raising these points. The behavioral result and the frontal theta trial-level regression model for subsequent memory are now described in the main results and included in the Discussion. As noted in the Discussion, the trial-level regression model indicates pre-probe and post-MPE frontal theta effects may explain memory for weak (and strong) mismatch probes. While the magnitude of probe-period MPE frontal theta is additionally predictive of memory for strong mismatch probes (as indicated by the Subsequent Memory × Strength interaction in the trial-level regression model in Fig. 5c and the logistic regression in Fig. 5b), frontal theta during this period does not additionally enhance memory for strong mismatch probes above that of weak mismatch probes (Fig. 5a). We added more discussion on why memory for strong and weak mismatch probes in the current experiment did not differ. We also discuss potential directions for future research to further probe mechanisms underlying MPE-driven learning.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      It is recommended that the authors determine whether formal path analyses, testing for mediation and moderation, would provide a useful approach from which to better integrate the disparate set of findings and make clear their causal/functional implications.

      At a minimum, adding a diagrammatic figure is recommended to visually depict the key components of the study (independent variables, physiological indices, outcome measures) and how they relate to each other both conceptually (ideally in a theoretically hypothesized manner) and in terms of the observed findings. Such a figure will help the reader keep track of the many types of findings and results threads, and with the goal of better organizing the results into a clearer narrative through-line.

      Thank you for the suggestions to add mediation analyses and diagrammatic figures. To more formally address our hypothesis that increases in attention/arousal following MPEs are explained in part by increases in cognitive control, we tested two mediation models: one with posterior alpha as the measure of attention and the other with pupil PC3 scores. We now diagram this hypothesis in Fig. 3a. Given the outcomes of the cross-correlation analysis suggested by Reviewer #2, we also tested an additional model where posterior alpha might explain the impact of strong MPEs on frontal theta (now diagrammed in Fig. 3e). We did not find credible evidence for an indirect effect in any of the models, suggesting that strong MPE-driven increases in control, attention, and arousal may be elicited independently. Finally, in a newly added summary diagram (Fig. 6), we highlight the observed effects of strong vs. weak MPEs on RTs, control, attention and arousal; differences in control, attention, and arousal for strong vs. weak predictions preceding the MPE; trial-level mismatch-specific or strong-specific effects on subsequent memory; and mismatch-specific interactions between processes during retrieval and responses to MPEs. We hope these revisions address Reviewer #1’s concerns and that these figures help the reader keep track of the key hypothesized relationships and main findings.

      Reviewer #2 (Recommendations for the authors):

      (1) The relationship between event segmentation and prediction errors has been reviewed recently in two papers (Nolden et al., 2024, Neuroscience & Biobehavioral Reviews; Rouhani et al., 2024, JOCN for a potentially relevant computational model). I wonder if insights from these papers can inform the Introduction/Discussion of the current manuscript.

      Thank you for these suggestions. These papers are now incorporated into the Introduction and the Discussion.

      (2) Brod et al. (2022, Psych. Bull. Rev.) have previously reported increased pupil dilations for MPE correlating with subsequent memory, specifically for strong prediction errors. I think it's worth including this paper in the Introduction as the finding is highly relevant. The Brod paper might provide more direct evidence of "MPE-related increases in pupil size" than the evidence the authors provide (p. 3).

      Thank you for this suggestion. This paper is now incorporated into the Introduction and Discussion.

      (3) I'm confused about the temporal analysis: "Trial-level regression analyses were conducted on the frontal theta, posterior alpha, and pupil time series from the associative retrieval test to identify temporal clusters that were sensitive to the factors of Mismatch (i.e., mismatch vs. match probes) and/or Strength (i.e., strong vs. weak associative pairs). For each participant and each time point, a linear regression model testing main effects of Mismatch and Strength, and a Mismatch × Strength interaction was run using R to compute beta weights for each regressor." What regression exactly was run? A separate model for each participant and time point? Across trials, then? Later, the authors mention that beta weights were averaged across participants and t-tests and permutation tests were conducted. However, if the data were averaged, what t-test was conducted? And how was the permutation test conducted? It's also unclear what the authors mean by "the sign of each participant's beta weights" - what sign?

      Thank you for raising this point. A regression was run separately for each participant and each time point, across trials: neurocognitive measure ~ Mismatch + Strength + Mismatch: Strength. Beta weights were averaged only for visualization; t-tests were conducted on each set of beta weights (across participants), separately for each time point. We modified the text of the Methods to describe our procedure more clearly.

      (4) The associative memory and recognition accuracy data are presented as d'. In addition, the authors should provide hits and false alarms to facilitate a better interpretation of the results.

      We now report these outcomes in Table 1 and refer to them in the main text.

      (5) The authors argue regarding the frontal theta that "Qualitatively, the main effect of Strength emerged later than the main effect of Mismatch, suggesting that the increase in frontal theta evoked by the probe was more sustained for weak compared to strong trials (or, as a corollary, that the greater control elicited by strong MPEs enabled more rapid resolution of conflict and ultimate choice selection)." It was unclear to me how that stems from the data.

      This statement has been removed altogether.

      (6) Especially in Figure 1, I think clarity can be improved if the authors would indicate the specific subsection they are referring to in the text, because even within, e.g., 1d, there are different graphs, so mentioning which graph is relevant for which statement would be helpful to the reader.

      Thank you for this suggestion for improving the clarity of the manuscript. Fig. 1d-f includes multiple parts because we wanted to show the raw data (the mean time series on the left) as well as the model coefficients (time series on the right). We are hopeful that, given that each subplot is titled and that the main text refers to both the subplots on the left and on the right, this will be clear to the reader as is. We welcome further guidance if this remains a concern.

      (7) This seems highly speculative to me: "Elevated PC4 scores on weak hits, where no error was made, may instead reflect retrieval practice and thus internally oriented attention. On these trials, when cueelicited retrieval may have been weaker, the probe may have provided additional support for pattern completion of the learning episode for that association, and this engagement in memory retrieval may have elicited pupil dilation (c.f., Strength effect in Figure 1f). Overall, PC4 may therefore reflect an attentional orienting response that, depending on the relative success of memory retrieval and probe identity, may direct attention internally or externally. " (p.13). In my opinion, the authors make a lot of assumptions about underlying processes. I'd consider removing.

      We removed this text.

      (8) In Figure 3c, should the x-axis be PC3 quintile? And (c) is not in the figure caption.

      Thank you for this note. We addressed these issues.

      (9) The carryover effects the authors report are interesting, but they seem detached, and it is unclear how they fit with the additional findings. In a paper that already includes many findings, I'd recommend either integrating better or removing.

      Following this guidance, we removed the readiness-to-remember findings in the interest of space and clarity.

    1. eLife Assessment

      This valuable study uses the Cntnap2 mouse model of autism to investigate how reduced activity in the dorsal CA1 region of the hippocampus contributes to impaired temporal binding and memory flexibility. The authors combined trace conditioning and radial maze assays with fibre photometry, optogenetic rescue, and brain-wide c-Fos mapping to provide solid evidence that dorsal CA1 hypoactivity limits the retention of temporally separated associations and biases learning towards less flexible strategies. The work advances understanding of hippocampal contributions to cognitive alterations associated with autism and will be of broad interest to researchers studying memory, hippocampal function, and neurodevelopmental disorders.

    2. Reviewer #1 (Public review):

      Summary:

      The uniqueness of this paper is the study of the formation of temporal binding-dependent memories in the cntnap2 mouse, a long-standing mouse model of autism that has been used to test therapeutic modalities.

      Strengths:

      I liked the combination of optical recordings and interventions and the backup of primary observations with control experiments.

      Weaknesses:

      (1) Fiber photometry recordings are too coarse to give salient clues to the underlying mechanism.

      (2) Are perturbed pyramidal cells causally responsible for the altered trace? What can be concluded about the possible role of inhibitory interneurons as potential drivers? The observations focus on abnormal regional activity as observed with fiber photometry and manipulated by optogenetics. The authors should state clearly the limits of their conclusions.

      (3) I found the "trace" nomenclature confusing. "....in which mice are required to memorize the association between a tone (Conditioned Stimulus) and a mild electric foot-shock (Unconditioned Stimulus), separated by a time interval called Trace (Sellami et al., 2017)." It seems that the conceptual model invokes the creation of an [eligibility] trace, characterized by its progressive disappearance over time. It may be a convention in the field or a matter of language, but it seems perverse to use "trace" to label the time interval rather than the entity that is decaying. If this is an accepted convention going back to Howard Eichenbaum, the authors should cite the paper that first introduced the convention.

      (4) I would advocate for the addition of some discussion points for the authors to consider.

      a) Is the retention of activity in CA1 related to phenomena at the cellular or subcellular level in CA1 pyramidal cells? I'm thinking of dendritic, delayed, and stochastic CaMKII activation (DDSC) as defined by Yasuda's group or short-term and associative plasticity of calcium dynamics (STAPCD) as delineated by Caya-Bissonette and Beique.

      b) Was the optogenetic intervention ever administered in a delayed fashion, capitalizing on the temporal advantages of optogenetics to probe dynamics?

      c) Is the newfound reliance on corticostriatal pathways something more than compensation at the behavioral level? Could it be driven in part by the ASD-related genetic changes?

    3. Reviewer #2 (Public review):

      The authors investigate the contribution of dorsal CA1 hippocampal dysfunction to cognitive impairments in the Cntnap2 knockout mouse model of autism spectrum disorder. Building on previous evidence implicating the hippocampus in episodic and relational memory processes, they combine trace fear conditioning, fiber photometry, optogenetic manipulation, a relational/declarative memory radial maze task, and cFos mapping to test whether altered CA1 function contributes to deficits in temporal binding and memory flexibility.

      The study has several important strengths. First, the work addresses a relatively understudied aspect of autism-related cognition, namely hippocampal-dependent memory processes, whereas much of the literature has focused on social behavior, cortical circuits, or striatal dysfunction. Second, the authors employ multiple complementary approaches that converge on a coherent mechanistic hypothesis. The behavioral data demonstrate a reduced ability of Cntnap2 knockout mice to retain associations across long temporal gaps. Fiber photometry recordings reveal reduced dorsal CA1 activity during conditions that challenge temporal binding, and optogenetic activation of dorsal CA1 neurons during the trace interval is sufficient to rescue memory performance. Together, these findings provide strong support for a causal contribution of dorsal CA1 activity to temporal binding deficits in this model.

      The second major strength of the manuscript is the extension of these findings to a more complex hippocampus-dependent memory task. The radial maze experiments indicate that Cntnap2 knockout mice show impaired memory flexibility and a greater reliance on egocentric learning strategies. The accompanying cFos analyses suggest altered recruitment of hippocampal and striatal networks during learning, providing a systems-level framework that may explain the observed behavioral phenotype.

      Overall, the main conclusions regarding impaired temporal binding and reduced dorsal CA1 engagement are well supported by the data. The optogenetic rescue experiments are particularly compelling because they move beyond correlation and directly test causality. The manuscript therefore makes a meaningful contribution to our understanding of how hippocampal dysfunction may contribute to cognitive abnormalities associated with autism.

      Weaknesses:

      Some conclusions are necessarily more inferential than others. In particular, the interpretation that the observed behavioral phenotype reflects a broader shift from hippocampal-dependent declarative memory toward striatum-dependent procedural learning is supported primarily by cFos activity patterns and behavioral strategy measures. While the data are consistent with this interpretation, they do not directly demonstrate a causal reorganization of memory systems. Similarly, although the findings identify a mechanism in the Cntnap2 model, caution is warranted when extrapolating these conclusions to autism spectrum disorder more broadly; but I believe this caution is addressed in the discussion.

      Despite these limitations, the study presents a coherent and well-executed body of work that provides novel mechanistic insight into hippocampal contributions to cognitive dysfunction in a widely used autism model. The findings should be of considerable interest to researchers studying hippocampal function, memory systems, and neurodevelopmental disorders.

    4. Reviewer #3 (Public review):

      Summary:

      The manuscript evaluated behavioral phenotypes in the Cntnap2 knockout mouse using two behavioral paradigms: trace fear conditioning and a radial maze task. The trace fear conditioning training is normal, but memory generalization is impaired. The inflexibility is suggested to be related to low activity in dCA1 neurons, which can be rescued by ChR2. The radial maze task data suggested a similar conclusion. Brain-wide cFos mapping indicated impairments in the Cntnap2 knockout mouse. The brain-wide cFos mapping does not show direct correlations with Cntnap2, limiting the interpretation of these data in the context of this paper.

      Strengths:

      The behavior data are solid.

      Weaknesses:

      The underlying mechanism is not fully investigated.

      Major points:

      (1) The authors should thoroughly check their manuscript as there are many typos in the current version that affect the readability.

      (2) In trace fear conditioning, the tone test impairment can be rescued by ChR2. Have the authors tried rescue experiments with Cntnap2? Rescue experiments in the radial maze task are also essential, either with ChR2 or Cntnap2.

      (3) The quality of the cFos example image in Figure 3 is too low. The authors should also provide example images for the other brain regions in the supplementary data, if possible.

      (4) The causal link between the brain-wide cFos mapping and the Cntnap2 knockout is weak. How to explain the increase of cFos cell densities in some brain regions, but the decrease in others?

    1. eLife Assessment

      This fundamental study reveals a new downstream mechanism that mediates mTOR's effect on lifespan in C. elegans. Using a combination of genetic, genomic, and functional analyses, the authors uncovered that a bile acid-like hormone, dafachronic acid (DA), acts downstream of mTOR to modulate lifespan. The reviewers found the evidence provided to be compelling.

    2. Reviewer #1 (Public review):

      This manuscript describes a novel downstream mechanism of mTORC1 deficiency-mediated lifespan extension in C. elegans. The authors demonstrated that the biosynthesis and the nuclear hormone receptor daf-12 binding of a bile acid-like hormone, dafachronic acid (DA), are essential for TORC1 mutant raga-1 to extend lifespan. Through RNA-seq and RNAi lifespan screen, they also discovered that a dehydrogenase, dhs-26, which is expressed in the canal-associated neurons, is regulated by DA/daf-12 and downstream of the mTORC1-DA signaling for lifespan extension. The authors also explored the conservation of mTOR/DA/daf-12/dhs-26 signaling in the mouse model. This work demonstrates significant findings that will advance the aging field and will be of interest to many researchers in this field. The conclusions are mostly well supported by data with proper controls.

      Some suggestions to strengthen the manuscript include:

      (1) Other mTOR activity perturbation or mutants should be used to support some of the core lifespan experiments. It will strengthen the conclusions made from raga-1 mutant only, although there is evidence from TOR RNAi in Figure 1g to support the daf-12 data in Figure 1d.

      (2) The authors showed in Figure 1h and 1i that DA supplementation rescued the shortened lifespan of raga-1;daf-9 but not raga-1;daf-12; and also rescued the shortened lifespan of raga-1; dnh-26 in Fig. 5e. Does DA supplementation itself extend lifespan? If its level is increased by mTORC1 inhibition and it is downstream of mTORC1 inhibition, it should theoretically extend lifespan. But from the reported publications, it seems that the DA supplementation lifespan modulation is highly dependent on genetic backgrounds. It will strengthen the conclusions if the authors provide the wild-type condition DA supplementation lifespan data and also related discussions about it.

    3. Reviewer #2 (Public review):

      Summary

      This manuscript by Schilling et al. presents an important advancement in our understanding of how mTOR signaling regulates organismal aging. While the longevity-promoting effects of reduced mTOR activity have been extensively documented across species, the mechanisms by which mTOR communicates systemic metabolic information to regulate lifespan remain unclear. In this study, the authors provide strong evidence that longevity induced by reduced TORC1 signaling requires the bile acid-like steroid hormone dafachronic acid (DA) and its cognate nuclear receptor DAF-12. Furthermore, through a combination of transcriptomics and functional genomics, they identify the conserved short-chain dehydrogenase DHS-26/DHRS1 as a previously unrecognized downstream effector of this pathway. The work integrates genetics, lifespan analyses, sterol measurements, transcriptomics, proteomics, endogenous genome engineering, and comparative mammalian datasets. The resulting model, in which mTOR influences lifespan through regulation of endocrine steroid signaling, represents a conceptual advance that links nutrient sensing, metabolism, and organismal aging. Although several mechanistic questions remain unresolved, the study is comprehensive, technically rigorous, and likely to be of broad interest to investigators studying aging, metabolism, endocrine signaling, and cellular stress responses.

      Strengths:

      One of the major strengths of this manuscript is its conceptual novelty. Rather than reinforcing the well-established role of mTOR as a longevity regulator, the study proposes a specific endocrine mechanism that links reduced mTOR activity to increased lifespan through steroid hormone signaling. This advances the field beyond descriptive observations of mTOR-dependent longevity and introduces a model in which bile acid-like hormones function as systemic mediators of nutrient-sensing pathways. The idea that endocrine steroid signaling may serve as a downstream effector of mTOR provides a new perspective on how longevity signals are coordinated at the organismal level.

      The genetic evidence supporting this model is particularly strong. In Figure 1, the authors use a series of epistasis experiments to demonstrate that mutations in daf-36, daf-9, and daf-12 suppress lifespan extension in raga-1 mutants. The DA supplementation experiments further strengthen the pathway ordering by rescuing longevity in hormone-deficient backgrounds while failing to restore lifespan in receptor-deficient animals. Importantly, the direct quantification of endogenous DA levels elevates the study by providing biochemical support for the proposed model.

      The transcriptomic analyses presented in Figure 2 provide a valuable systems-level perspective on the interaction between mTOR and steroid signaling pathways. The observation that DAF-12 profoundly reshapes the RAGA-1 transcriptional program highlights the importance of steroid signaling in mediating the physiological consequences of reduced mTOR activity. The enrichment of metabolic, lysosomal, and peroxisomal pathways is consistent with established longevity-associated programs and generates a valuable resource for future mechanistic studies.

      Figure 3 effectively integrates discovery-driven and hypothesis-driven biology. The authors use transcriptomic information to prioritize candidate genes and then perform a functional genomic screen to identify factors required for RAGA-1-mediated lifespan extension. This approach converges on DHS-26, which subsequently emerges as a central mechanistic component of the study. The progression from transcriptomics to functional validation is well executed.

      In Figure 4, the generation of CRISPR-engineered dhs-26 deletion mutants and endogenous tagged reporter strains provides strong validation for DHS-26 function. The demonstration that dhs-26 deletion selectively abolishes RAGA-1-dependent longevity without substantially affecting wild-type lifespan strongly supports its role as a context-dependent mediator of mTOR signaling. Furthermore, the conservation analyses linking DHS-26 to mammalian DHRS1 provide biological context and enhance the broader significance of the findings.

      In Figure 5, multiple independent experimental approaches converge on the conclusion that DHS-26 participates in DA-dependent lifespan regulation. The rescue of lifespan by DA supplementation, reductions in DA levels in raga-1;dhs-26 mutants, reporter-based analyses of DAF-12 activity, and proteomic profiling collectively support a mechanistic model. The proposed positive feedback relationship between DA/DAF-12 signaling and DHS-26 is intriguing and offers a plausible explanation for how endocrine signaling may amplify longevity-promoting responses. Finally, the incorporation of mammalian datasets showing regulation of DHRS1 by rapamycin and FXR signaling provides a promising avenue for future studies investigating conservation of this pathway.

      Weaknesses:

      Despite the many strengths of the study, important mechanistic questions remain unresolved. The most significant limitation is that the precise molecular connection between reduced mTOR activity and increased DA production remains unclear. While the genetic and biochemical data convincingly place DA/DAF-12 signaling downstream of mTOR, the study does not establish whether mTOR regulates DA biosynthesis, degradation, intracellular trafficking, sterol uptake, or hormone availability. The observed increase in endogenous DA levels is statistically significant but relatively modest, and the mechanistic basis for this increase remains speculative. Additional experiments examining sterol flux, enzyme activity, or intracellular sterol trafficking would substantially strengthen the proposed model.

      The transcriptomic analyses in Figure 2 are informative but correlative. Because the RNA-sequencing was performed at a single adult time point, it remains difficult to distinguish primary transcriptional responses from secondary adaptive changes. Similarly, while pathway enrichment analyses identify plausible processes, they do not establish direct regulatory relationships. Additional temporal analyses or direct assessment of DAF-12 occupancy at candidate loci would strengthen mechanistic interpretations and help distinguish direct from indirect targets.

      A major unresolved question concerns the biochemical function of DHS-26 itself. While the genetic evidence clearly establishes DHS-26 as an important regulator of RAGA-1-mediated longevity, its endogenous substrate and enzymatic activity remain unknown. The manuscript presents evidence linking DHS-26 to sterol metabolism, but direct biochemical characterization is lacking. Thus, the mechanistic model remains somewhat incomplete. Defining the substrates and products of DHS-26 activity would greatly strengthen the study and provide important insight into how this enzyme influences DA availability.

      Another area requiring additional clarification is the proposed neuroendocrine role of DHS-26. The expression of DHS-26 in canal-associated neurons is interesting and raises the possibility that these cells participate in systemic longevity regulation. However, the current data do not establish whether DHS-26 functions autonomously within these neurons or whether expression in other cell types contributes to the observed phenotypes. Tissue-specific rescue or depletion experiments would strengthen the neuroendocrine model and help establish physiological sites of action.

      Finally, the mammalian data presented in Figure 5 are supportive and suggestive of evolutionary conservation, but they remain correlative. While regulation of DHRS1 expression by rapamycin and FXR signaling is interesting, these observations do not yet demonstrate functional conservation of the longevity mechanism itself. Additional studies directly testing DHRS1 function in mammalian systems will be required before stronger conclusions regarding conservation can be drawn.

      In summary, this manuscript provides a significant contribution to the aging field and introduces a model linking mTOR signaling, endocrine steroid hormones, and longevity. The study is comprehensive, technically sophisticated, and supported by multiple complementary approaches. Although some mechanistic questions remain open regarding the precise regulation of DA production, the biochemical function of DHS-26, and the extent of conservation, these limitations represent opportunities for future investigation. Overall, the work substantially advances our understanding of how nutrient-sensing pathways regulate aging and is likely to stimulate considerable interest within the fields of aging biology, metabolism, and endocrine signaling.

    4. Reviewer #3 (Public review):

      Summary:

      This interesting manuscript provides evidence that the well-established consequences of (reduced) mTOR activity on longevity are, at least in part, mediated by regulation of dafachronic acid (DA) availability and its signalling via its nuclear receptor DAF-12 in C.elegans, with some supporting evidence derived from mouse studies that similar processes may be functional in mammalian systems, i.e., be evolutionarily conserved. Earlier studies by the group have established that DA/DAF-12 signaling promotes adult longevity in several contexts. DA is a bile acid look-alike, and DAF-12 is a homolog of mammalian bile acid-activated nuclear receptors FXR and VDR: recent experimental studies and human cohort studies have indicated a role of (specific) bile acids in mammalian longevity.

      The hypothesis that mTOR and DA/DAF-12 signaling interact to modulate longevity in C.elegans is novel and of great potential interest. The hypothesis has rigorously been tested in a series of well-performed experiments employing mutant strains, functional genomic screens, and DA exposures, etc.. It is convincingly demonstrated that DA/DAF-12 does not directly impact mTOR (assayed on AMPK phosphorylation) and acts downstream of the pathway. The short-chain hydrogenase DHS-26 (mammalian homologue DHRS1) was identified as a downstream target and modulator of this mTOR-DA-DAF12 axis by modulating the lifespan of the mTOR regulator raga-1. As the components of this axis are expressed in different cell types of the worms, this finding indicates a neuroendocrine mode of action. Mode of action of DHS-26 appears to be based on modulation of cholesterol and lathosterol, i.e., substrate availability for DA production.

      Strengths:

      Overall, the manuscript is well-written and builds up the story in a clear fashion. The conclusions are based on solid data and of relevance for ageing research, also because the mechanism identified appears to be evolutionary conserved.

      Weaknesses:

      No overt weaknesses were identified by this reviewer.

    1. eLife Assessment

      This important study combines crystallographic fragment-based screening with fluorescence-polarization competition assays to demonstrate that the SARS-CoV-2 accessory protein ORF9b, and its interface with the mitochondrial import receptor TOM70, is accessible to small-molecule binding. This is a result with practical implications beyond coronavirus biology for efforts to restore interferon signalling and the antiviral response. The structural evidence is convincing, resting on an unusually large body of ligand-bound crystal structures whose fragment hotspots are independently corroborated by orthogonal biophysical approaches including fluorescence polarization and SPR. This work would benefit from additional information regarding the use of CHAI-1 and other attempts to determine the interactions between hit compounds and Tom70.

    2. Reviewer #1 (Public review):

      Summary:

      The authors used FBDD screening to identify numerous compounds interacting with the ORF9b dimer. They expanded the original fragment hit, soaked the derivatives into the crystals and confirmed their binding poses, and showed that the derivatives bind the target with higher affinity. The authors further targeted the ORF9b binding site on TOM70, and used a fluorescence polarization-based (FP) assay to screen a compound library and obtained several hits. Structure-activity relationship (SAR) optimization yielded hit analogs that have higher binding affinity to TOM70.

      Strengths:

      (1) The study adopted novel drug design strategies, including stabilizing ORF9b homodimer to prevent it from binding TOM70, and blocking ORF9b from binding TOM70 by screening compounds that compete with ORF9b for binding TOM70.

      (2) The work established a feasible high-throughput screening assay. This FP-based assay screened ~50,000 compounds, from which two hit compounds were further optimized to yield analogs with higher binding affinity.

      Weaknesses:

      (1) The study lacks functional assays to evaluate whether the ORF9b-stabilized compounds or TOM70 binding compounds could affect IFN inhibition caused by ORF9b or virus infection.

      (2) There is a lack of experimental evidence to reveal the binding mode of lipidated-compounds with ORF9b homodimer.

      (3) There is a lack of experimental evidence to reveal the binding mode of HTS hits or analogs for TOM70.

      (4) Overall, none of the compounds shown in the paper have promising potency warranting further development; their binding affinity is limited to the micromolar range.

    3. Reviewer #2 (Public review):

      Summary:

      The authors investigate chemical strategies to disrupt the interaction between the SARS-CoV-2 accessory protein Orf9b and the host mitochondrial receptor Tom70, an interaction implicated in suppression of type-I interferon responses. They employ two discovery approaches: a crystallographic fragment screen against the Orf9b homodimer and a high-throughput fluorescence polarization screen for compounds that compete with Orf9b binding to Tom70. The study identifies fragment-binding hotspots on Orf9b, develops lipidated analogs that stabilize the Orf9b homodimer, and discovers Tom70-binding compounds with low micromolar activity that inhibit Orf9b binding in vitro.

      Strengths:

      An impressive amount of work using a variety of complementary approaches and methods to validate binding (FP, SPR, and computational modelling and SAR). The combination of crystallographic fragment screening on Orfb9 and HTS on Tom70 provides two independent routes for perturbing the Orf9b-Tom70 interaction. The structural work seems to be of very high-quality. The fragment campaign is extensive, yielding a substantial number of fragment-bound structures and identifying biologically meaningful binding hotspots on Orf9b.

      Finally, the screen results in reporting useful chemical starting points. Although potency remains modest, the study provides tractable scaffolds and a clear framework for future optimization.

      Weaknesses:

      General comment:

      (1) Although there is already an incredible amount of data presented, one limitation of this study is the lack of cellular validation - do these drugs enter cells, restore interferon signalling, reduce viral loads, or alter Orfb9 localization?

      (2) The logic of locking Orfb9 as a dimer is that the monomer binds Tom70 - thus, a more stable dimer means less monomer. In Figure 2, the Orfb9 homerdimer stabilization by compounds should reduce binding affinity to Tom70. A direct binding experiment measuring reduced Tom70 binding with compound treatment would better strengthen this claim.

    4. Reviewer #3 (Public review):

      Summary:

      This paper attempts to and succeeds in demonstrating that Orf9b is able to bind small molecules using X-ray fragment screening, SPR and FP assays. Exploration of sites from the fragment screening is performed along with fragment linking with inter-dimer lipid moieties.

      Strengths:

      The experimental work looks strong and well performed. The interpretation of the data is appropriate and was often validated through orthogonal methods and follow-up compounds. The use of Tom70 to find binders that might disrupt interactions between Orf9b and Tom70 is elegant.

      Weaknesses:

      The use of Chai-1 to predict co-folded structures with binding molecules was not properly described - no mention of this in the methods. It was not commented on whether the compounds which were found were attempted to be co-crystallised. If they were but negative data was collected (didn't crystallise, didn't diffract or no additional density was found), then this needs to be stated.

    1. eLife Assessment

      This important study on multisensory learning examines the effect of integrating color vision and olfaction in the center for learning and memory (the mushroom body) in the fruit fly brain, and finds that memory performance is improved, with neurons otherwise selective for color being recruited into odor memories via interneuron interactions. The combination of carefully controlled behavioural experiments, neurogenetic manipulations of individual neurons and detailed analyses of synaptic connectivity in the fly brain connectome convincingly supports the findings. The circuit- and receptor-level demonstration of interactions across sensory modalities in memory performance will be of interest to the broader neuroscience field.

    2. Reviewer #1 (Public review):

      Summary:

      The study investigates how learning with combined visual and olfactory cues strengthens memory in fruit flies. It demonstrates that pairing colours with odours improves later memory performance, even when only one of the two cues is presented during testing. The authors show that multisensory learning recruits visually responsive Kenyon cells in the mushroom body into memory representations that would otherwise primarily encode odours. Their experiments indicate that the serotonergic DPM neuron links sensory representations that are normally separated, while the APL neuron regulates local GABAergic inhibition of separated learning subcircuits. Together, these findings provide a mechanistic explanation for how a single sensory cue can retrieve a broader memory of a multisensory experience.

      Strengths:

      A major strength of the paper is its integration of behavioural experiments, targeted neuronal manipulations, and detailed anatomical analysis to address a clear mechanistic question. The findings are supported by multiple complementary experiments showing that multisensory learning enhances memory and recruits visual pathways into olfactory memory representations. Overall, the work provides a coherent mechanistic framework for how multisensory experiences strengthen subsequent memory.

      Weaknesses:

      A limitation of the paper is that it represents an unusual case, as substantial parts of the broader study were previously published in Nature and subsequently retracted because the physiological findings could not be reproduced. Those physiological experiments would have helped resolve several mechanistic questions raised by the behavioural results and directly test how multisensory information is integrated within the fruit-fly learning circuit. Presenting only the reproducible behavioural and anatomical findings is therefore appropriate and preserves the reliable contribution of the work. Nevertheless, the absence of reproducible physiological evidence makes the mechanistic model less complete and more inferential than it would be in a fully comprehensive study. The conclusions should consequently be framed as a well-supported circuit model rather than a direct demonstration of the underlying physiological processes.

    3. Reviewer #2 (Public review):

      Okray et al. identify a novel form of multisensory memory in Drosophila, where pairing reward with a color+odor together gives a stronger memory than color alone or odor alone. Remarkably, this multisensory enhancement occurs even if only one modality is used during testing (i.e. training color+odor, then testing odor alone gives a stronger memory than training odor alone, then testing odor alone), showing that the two modalities are persistently linked following training. The manuscript presents compelling behavioural genetic evidence that the normally visual-selective gamma-d Kenyon cells acquire a functional role in the retrieval of odor memories following odor+color training, and that this occurs via transfer from gamma-main KCs via the serotonergic interneuron DPM.

      The key pieces of evidence supporting this conclusion are that olfactory retrieval of multisensory memories requires:

      (1) synaptic output from gamma-d KCs during retrieval (but not training);

      (2) synaptic output from gamma-main KCs during training and retrieval (whereas it's only required during retrieval, not training, for pure-olfactory memory);

      (3) synaptic output from DPM during training and retrieval, and expression of the serotonin receptor 5HT2A in gamma-d KCs.

      In the absence of physiological data, the exact nature of the gamma-d KCs' participation in olfactory retrieval following odor+color training remains unclear. For example, do the gamma-d KCs encode the odor identity (i.e., is there an odor-specific pattern of gamma-d KCs activated for a particular odor+color combination), or does their activity provide a general activity boost to other neurons (e.g. gamma-m) that encode odor identity? This will be interesting to address in future studies.

      That being said, the behavioural data are clear and back up the authors' conclusion that signaling between KC subtypes via DPM underlies multisensory integration for multimodal memories in the fly mushroom body.

    4. Author response:

      We thank the reviewers for their time and insightful comments. We are also grateful for their appreciation of the unusual circumstances that led to the publication of the manuscript in its current form.

      In response to reviewer #2’s question about replication, we provide additional details here. The error in our retracted original publication affected only the imaging results. Despite this, we reproduced key behavioural experiments by generating additional datasets (rather than simply rechecking records and authenticating results) and therefore have full confidence in our behavioural findings. We are very happy to share some of these replication experiments below:

      Author response image 1.

      Data showing replication of key experiments. From L-R these data replicate those shown in Figure 2e, Figure 4h, Figure 4c, Figure 4d.

      The conceptual framework of the original study remains valid. It was actually formulated based on the behavioural data and before any physiological recordings were made. We believe that it still represents the most parsimonious explanation for the observed behavioural results. Multiple behavioural findings support a model in which multisensory training leads to the recruitment of visual γd Kenyon cells into an otherwise olfactory memory trace. These include: (1) the requirement for γd KC output during olfactory retrieval following multisensory training; and (2) the sequential learning experiments, which were originally designed to test this model and provide independent evidence for its predictions. We nevertheless agree that the loss of the physiological data reduces the amount of evidence supporting the proposed mechanism. The nature of the physiological changes following multisensory learning remains an important question that we intend to address in future work.

      We also thank the reviewers for identifying the unfortunate typo in the Abstract, which we believe contributed to the confusion over the dopamine receptors tested in this circuit. We selected these receptors based on our in-house single-cell transcriptomic expression data, together with published evidence indicating their specific expression in the neurons of interest. We have also responded to reviewer comments about the clarity of the figures and added additional labels to figure 2, to clarify the experimental paradigm in each case.

    1. eLife Assessment

      This valuable study provides key insights into the role of the G protein-coupled receptor GPR34 in an Alzheimer's disease (AD) model. Notably, its findings differ, at least in part, from those of previous studies, although the underlying reasons for these discrepancies are not investigated or discussed. The data show that GPR34 deficiency enhances the transcriptional disease-associated microglia signature in microglia and provide solid evidence that myelin is a GPR34 ligand, although the relevance of myelin engulfment in AD is unclear. This study would be of interest for neuro-immunologists and AD clinicians.

    2. Reviewer #1 (Public review):

      Summary:

      The authors sought to understand the impact of the decreased expression of the G-protein-coupled receptor GPR34 in Alzheimer´s disease (AD). They analyzed the transcriptional impact of GPR34 deficiency in mice and found that it induced a DAM-like phenotype in control mice and enhanced the DAM signature in the AD model 5xFAD, although it did not result in amyloid plaque clearance or gross changes in microglia or astrocytes. Next, the authors developed an in vitro model of GPR34 deficiency using a CRISPR/Cas9 strategy in human iPSCs to introduce functional mutations that resulted in GPR34 protein deficiency in induced microglial cells. In this model, the authors identified myelin as a ligand of GPR34 and showed that GPR34 deficiency resulted in reduced myelin debris engulfment and transcriptional changes related to lysosomal pathways.

      Strengths:

      The combined strategy of using in vivo and in vitro models of GPR34 depletion is robust, and the transcriptional analyses are thoroughly performed.

      Weaknesses:

      The paper´s two main findings related to the lack of GPR34 (enhancement of DAM signature in vivo and reduced myelin engulfment in vitro) are disconnected. At the very least, the authors should discuss what the relevance of myelin clearance in AD is, but the paper would strongly benefit from a more thorough assessment of the impact of GPR34 deficiency in vivo, particularly because no effects on amyloid clearance were observed. The authors could assess whether GPR34-deficient 5xFAD mice have reduced cognitive performance, which, based on their in vitro findings, could be related to the myelin pathology in AD (previously described: see PMID 36284351). The analysis showing reduced myelin content in GPR34-deficient microglia in vitro is superficial and does not allow for identifying whether GPR34 is related to reduced engulfment or increased degradation, which could be related to the changes in the lysosomal gene CD68 identified in vivo. In addition, it would be interesting to compare the transcriptional profile induced by myelin phagocytosis with that of 5xFAD or AD patients, to gain insight into the impact of the signature. Finally, the transgenic approach to delete GPR34 in vivo could have been complemented with experiments with the GPR34 antagonists (YL-365 or S-E49) or agonist (Compound 4B), possibly helping in identifying the source of discrepancy with previous papers showing that GPR34 promotes amyloid clearance.

    3. Reviewer #2 (Public review):

      Summary:

      Using the 5xFAD model in combination with GPR34 mice, the authors explore the function of microglia in the context of neurodegeneration. Using a broad spectrum of methodology, they show that DAM signatures are increased in KO 5xFAD mice. Using several KO clones of GPR34 KO iMGLs and another set of broad methodologies, the authors show that GPR34 is important for microglia homeostasis,<br /> phagocytosis, specifically of myelin. GPR34 KO iMGLs also show a distinct transcriptional response to myelin. Together, they propose that GPR34 limits microglial activation in neurodegeneration.

      Strengths:

      All methods are state-of-the-art, and the combination of mouse and human microglia responses is a particular strength.

      Weaknesses:

      No weaknesses were identified by this reviewer.

    4. Author response:

      We greatly thank the Reviewing Editor, Senior Editor and the reviewers for their constructive and thoughtful feedback, as well as for recognizing the significance and strengths of our study. We are very encouraged by overall positive assessment and appreciate these insights to strengthen the manuscript. Below, we outline our plans to address the key points raised in reviewer#1’s public review. We also note that reviewer#2 did not identify any weaknesses in the study and are thankful for this positive evaluation.

      Point 1: Relating in vivo and in vitro findings and discussing the relevance of myelin clearance in AD. As suggested, we will elaborate our discussion to cover the relationship between our in vivo and in vitro findings and present a more unified picture of GPR34 function. We will also highlight the relevance of myelin clearance in Alzheimer’s disease independent of amyloid plaque burden.

      Point 2: in vivo assessments and phenotypes. While we recognize the value of expanding our in vitro and cellular findings to in vivo and cognitive measures, we believe this additional assessment is beyond the scope of this current study. At least, we will expand our discussion to relate our findings of GPR34-medilated myelin pathology in the contexts of neurodegeneration and cognitive deficiency in AD.

      Point 3: Myelin engulfment versus lysosomal degradation. We agree that the clear distinction between impaired myelin engulfment and altered lysosomal degradation is an important point. Although additional experiments to address these scenarios are beyond the scope of the study, we will perform targeted pathway analysis of existing transcriptomic datasets focusing on the phagosome and lysosomal degradation pathways along with the in vivo findings related to CD68, which could offer an interpretation.

      Point 4: Comparison of transcriptional signatures. As suggested, we will compare the transcriptional signatures induced by myelin exposure in WT iMGLs with our 5xFAD RNA-seq datasets, as well as publicly available datasets from human Alzheimer’s disease microglia, to unravel the potential converged and distinct features.

      Point 5: Reconciling our findings with previous GPR34 studies. We will expand our discussion to compare and clarify our current findings with the previous studies and cover potential reasons for those different observations. We will also cover the current limitations of existing GPR34 pharmacological tools compounds, including their limited blood-brain permeability, which precludes the effective use of those tools for proposed in vivo studies.

    1. eLife Assessment

      This is a fundamental study of individual variation and the contribution of learning to behavioural individuality. The experimental design of massively parallel behavioural phenotypes is outstanding and the conclusions are supported by a compelling and rigorous analysis across a large number of experiments in thousands of individuals across genotypes and conditions. The dataset further represents an advance in studying visual associative learning thanks to the ability to make longitudinal measurements of many behavioural decisions within the same animals. These results are a major contribution to the understanding of the sources of behavioural individuality.

    2. Reviewer #1 (Public review):

      "Learning is a fundamental source of individuality," by Manna and colleagues, interrogates different sources of variation in individual behavior. The authors place individual flies in a Y-shaped arena, which is a common design in the field, and illuminate the arms of the Y with blue versus green light. They track the color preference of individual animals and also perform operant conditioning, meaning that they teach the fly to avoid a particular color/arm by generating a foot shock when the fly enters that arm. There are a number of things that are impressive about this setup: The authors are able to collect data on thousands of individual flies of many different strain backgrounds, and they demonstrate a strong change in color preference after conditioning. This is nice, because in past papers visual learning ability has been modest and difficult to study. To put a number on it, in this paper animals on average don't show a color preference at the start of the assay, spending around 30% of their time in the one arm illuminated green, and the remaining time in the two arms illuminated blue. After conditioning, the average animal spends only 23% of its time in the green arm.

      The authors run 64 animals through the assay for each of 88 wild type strains (maybe? see Major Point 1 below) and see considerable strain-specific (genetic) variation in the change in time spent in the shocked color after conditioning. Some strains show no learning, while others spend <10% of their time in the shocked color after conditioning. They also, I believe, see that some strains have more variability across individuals, which would suggest that some strains have stronger canalization at the development or circuit function level than others-i.e. some genotypes produce more consistent copies of the individual, others less consistent copies. (Or, some genotypes produce robust circuits, and others produce noisy circuits.)

      Finally, the authors argue statistically that learning itself increases variability in individual performance. This makes a lot of sense to me intuitively. Learning changes the physical/chemical properties of circuits in the brain, and because it evolves over time and interacts with environmental variables, it seems like it should send different animals down different channels. Or, at a conceptual level, if I learn to play the piano and my sister doesn't (because of some genetic difference between us or something stochastic), this learning experience will cause all sorts of other differences in our behavior as time passes. I also think the authors do have enough data to be able to make this finding. However, the presentation of the argument in this portion of the paper is hard for me to understand, and I am not an expert in statistics, so the strength of the result is difficult for me to evaluate.

      Major points:

      (1) It's difficult to track through the paper the number of animals tested for different assays. At the beginning, it says N=5632, which works out to 64 flies for each of the 88 DGRP strains. 64 happens to be the number of parallel Y arenas they have. Later in the methods, there's description of more variation within the set of 64 for each strain-two different parent sets per strain, different sexes, conditioned and un-conditioned. And, while the results text focuses on the color learning, the methods discuss additional assays (place learning, multi-day learning).

      Given the numbers, does each run of the 64 mazes include all the tested flies of one strain, or are flies of many strains included in each batch? Do different flies do different assays (color, place, multi-day) or do they all do all the assays? Perhaps there is a table including this information already in the supplement, but I recommend making it much clearer in the main results text and methods. While the dataset is large, if it is split over many conditions and/or if batch and genotype confound each other, this will affect the robustness of the results and how strong the conclusions can be.

      (2) The data presentation in Figure 1 is elegant and easy to follow, but getting into Figure 2 and subsequently, I get lost in the statistics and have trouble understanding what is being measured. My understanding of the big picture is that while genetics and individual randomness contribute a lot to behavior, the evidence for learning as an amplifier of individuality is that variance in behavior among animals of the same strain increases over time in the conditioned group (i.e. the group that is doing the most learning, or a specific kind of learning), but not in the control group. This idea is illustrated in the flattening distributions in the cartoons in Figure 1A. The authors should include graphs of the real data that use the same format as in that cartoon. Instead, the graphs present "residuals," and I don't know what those are. I suspect it's "variation left over after accounting for effects of strain and individual stochasticity." I see the residuals being tracked per strain over time in Figure 2H, but I don't see the change over time in other graphs. I'm looking for something simple like, "variation within the strain at the beginning of learning and at later time points in learning." (But I'm not sure exactly what instantaneous measurement would be the focus in longitudinal analyses of learning behavior.)

      (3) Figure 3 is a cool stab at tracking down the precise mechanism by which stochastic environment interacts with learning to send individuals along different behavioral routes. But again, like in Figure 2, I don't have the sophisticated understanding of statistics to understand exactly what the graphs are telling me, or how they relate to the underlying measurements. I'm relying on the results text alone to reach a conceptual understanding and just taking the graphs on trust.

      So, overall, the authors have a very nice body of work here and with the potential to add a new facet to our understanding of the origins of diversity in animal behavior. In addition to the interpretations they focus on here, this dataset also represents an advance in studying visual associative learning in general, and quite an amazing ability to make longitudinal measurements of many behavioral decisions within the same animals. Improving the data presentation to make it easier to follow for a larger swathe of researchers, especially in figures 2 and 3, will increase its potential impact.

      Comment on revised version:

      The authors have addressed my main points, including adding description of their statistical analyses and providing more detail about the different assays run and which animals were included in the same assay batches.

    3. Reviewer #2 (Public review):

      Summary:

      The authors set out to test the extent to which differences in learning capacity and experience contribute to behavioural variation in a genetically identical population under identical environmental conditions.

      Strengths:

      The authors developed and used a scaled-up version of a simple two-choice behavioural paradigm allowing them to test thousands of individuals across multiple genotypes. They then deployed clever and powerful statistical analysis methods and provided compelling evidence for a role of variability in learning in the expression of behavioural variation.

      Weaknesses:

      There are no major weaknesses, although some level of longitudinal analysis to strengthen the evidence for a strict definition of individuality would be a welcome extension of a future study. In addition, it would have been very interesting, although understandably beyond the current scope, to delineate a potential source of learning variability in the brain.

    4. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      "Learning is a fundamental source of individuality," by Manna and colleagues, interrogates different sources of variation in individual behavior. The authors place individual flies in a Y-shaped arena, which is a common design in the field, and illuminate the arms of the Y with blue versus green light. They track the color preference of individual animals and also perform operant conditioning, meaning that they teach the fly to avoid a particular color/arm by generating a foot shock when the fly enters that arm. There are a number of things that are impressive about this setup: The authors are able to collect data on thousands of individual flies of many different strain backgrounds, and they demonstrate a strong change in color preference after conditioning. This is nice, because in past papers, visual learning ability has been modest and difficult to study. To put a number on it, in this paper, animals on average don't show a color preference at the start of the assay, spending around 30% of their time in the one arm illuminated green, and the remaining time in the two arms illuminated blue. After conditioning, the average animal spends only 23% of its time in the green arm.

      The authors run 64 animals through the assay for each of 88 wild-type strains (maybe? see Major Point 1 below) and see considerable strain-specific (genetic) variation in the change in time spent in the shocked color after conditioning. Some strains show no learning, while others spend <10% of their time in the shocked color after conditioning. They also, I believe, see that some strains have more variability across individuals, which would suggest that some strains have stronger canalization at the development or circuit function level than others, i.e., some genotypes produce more consistent copies of the individual, others less consistent copies. (Or, some genotypes produce robust circuits, and others produce noisy circuits.)

      Finally, the authors argue statistically that learning itself increases variability in individual performance. This makes a lot of sense to me intuitively. Learning changes the physical/chemical properties of circuits in the brain, and because it evolves over time and interacts with environmental variables, it seems like it should send different animals down different channels. Or, at a conceptual level, if I learn to play the piano and my sister doesn't (because of some genetic difference between us or something stochastic), this learning experience will cause all sorts of other differences in our behavior as time passes. I also think the authors do have enough data to be able to make this finding. However, the presentation of the argument in this portion of the paper is hard for me to understand, and I am not an expert in statistics, so the strength of the result is difficult for me to evaluate.

      Major points

      (1) It's difficult to track through the paper the number of animals tested for different assays. At the beginning, it says N=5632, which works out to 64 flies for each of the 88 DGRP strains. 64 happens to be the number of parallel Y arenas they have. Later in the methods, there's a description of more variation within the set of 64 for each strain, two different parent sets per strain, different sexes, conditioned and unconditioned. And, while the results text focuses on the color learning, the methods discuss additional assays (place learning, multi-day learning).

      Given the numbers, does each run of the 64 mazes include all the tested flies of one strain, or are flies of many strains included in each batch? Do different flies do different assays (color, place, multi-day), or do they all do all the assays? Perhaps there is a table including this information already in the supplement, but I recommend making it much clearer in the main results text and methods. While the dataset is large, if it is split over many conditions and/or if batch and genotype confound each other, this will affect the robustness of the results and how strong the conclusions can be.

      (2) The data presentation in Figure 1 is elegant and easy to follow, but getting into Figure 2 and subsequently, I get lost in the statistics and have trouble understanding what is being measured. My understanding of the big picture is that while genetics and individual randomness contribute a lot to behavior, the evidence for learning as an amplifier of individuality is that variance in behavior among animals of the same strain increases over time in the conditioned group (i.e., the group that is doing the most learning, or a specific kind of learning), but not in the control group. This idea is illustrated in the flattening distributions in the cartoons in Figure 1A. The authors should include graphs of the real data that use the same format as in that cartoon. Instead, the graphs present "residuals," and I don't know what those are. I suspect it's "variation left over after accounting for effects of strain and individual stochasticity." I see the residuals being tracked per strain over time in Figure 2H, but I don't see the change over time in other graphs. I'm looking for something simple like, "variation within the strain at the beginning of learning and at later time points in learning." (But I'm not sure exactly what instantaneous measurement would be the focus in longitudinal analyses of learning behavior.)

      (3) Figure 3 is a cool stab at tracking down the precise mechanism by which a stochastic environment interacts with learning to send individuals along different behavioral routes. But again, like in Figure 2, I don't have the sophisticated understanding of statistics to understand exactly what the graphs are telling me, or how they relate to the underlying measurements. I'm relying on the results text alone to reach a conceptual understanding, and just taking the graphs on trust.

      So, overall, the authors have a very nice body of work here, and with the potential to add a new facet to our understanding of the origins of diversity in animal behavior. In addition to the interpretations they focus on here, this dataset also represents an advance in studying visual associative learning in general, and quite an amazing ability to make longitudinal measurements of many behavioral decisions within the same animals. Improving the data presentation to make it easier to follow for a larger swathe of researchers, especially in figures 2 and 3, will increase its potential impact.

      Reviewer #2 (Public review):

      Summary:

      The authors set out to test the extent to which differences in learning capacity and experience contribute to behavioural variation in a genetically identical population under identical environmental conditions.

      Strengths:

      The authors developed and used a scaled-up version of a simple two-choice behavioural paradigm, allowing them to test thousands of individuals across multiple genotypes. They then deployed clever and powerful statistical analysis methods and provided compelling evidence for a role of variability in learning in the expression of behavioural variation.

      Weaknesses:

      There are no major weaknesses, although some level of longitudinal analysis to strengthen the evidence for a strict definition of individuality would be a welcome extension of a future study. In addition, it would have been very interesting, although understandably beyond the current scope, to delineate a potential source of learning variability in the brain.

      Following our provisional response to the reviewers, we have implemented these additions to the manuscript:

      (1) We have added 7 additional tables (Table 1-6 and table 8) to the supplementary that detail how many individual flies were used in which of the seven separate experiments, how the individuals were distributed across genotypes, replicates and sexes, and how many were filtered out before the final analysis. At the bottom of each table, we added a short description of the type of experiment and a brief explanation of the filtering. The four smaller experiments where we tested the two mutant lines and one wild-type DGRP line were used primarily to test and validate the experimental platform and the behavioural paradigm used for the main experiment. In these four experiments we tested green place learning, blue place learning, green colour learning, blue colour learning in four separate batches of flies. In each of these experiments we used 192 individuals (64 individuals x 3 genotypes x 4 experiments = 768 individuals in total). In the multiday experiment we used 64 flies per genotype per each of the four groups of sequences of learning paradigms, in total 512 individual flies. Here, each of these 512 individuals were retested in different learning paradigms over 4 days (Table 5). The main experiment was the green place learning (Table 6) where we tested all 88 DGRP lines and again the two mutant lines was used to obtain the majority of the main results and conclusions (64 individuals x 90 genotypes = 5760 individuals). Lastly, additional 896 individuals were measured in the experiment using blue place learning paradigm to test the consistency of learning behaviour as opposed to colour bias within genotype (Table 8). In summary, in all experiments, we have always measured behaviour in 64 flies per genotype (full loading of the behavioural platform), and they were distributed almost entirely evenly across replicates, sexes, and conditions (control vs conditioned). No individual was reused across experiments. In most cases, after filtering the data, 60 or fewer individuals were used in final analyses. For the very few deviations from this experimental design (which occurred due to unforeseen events such as dropped/sick vials, flies flying away or accidentally squished during setup, skewed number of males and females etc.) we added a short explanation in the text below the tables. In total, across all reported experiments in this study, we measured behaviour in 7936 individuals.

      (2) We have added a schematic visual representation of classical measurement of individuality (variance of the distribution of behaviour within genotype where genetically identical individuals are raised in the same environment), entropy-based measurement of individuality (residual individuality) and the change in residual individuality, as we use them in this study (Figure 2D). We also provide a list of different DH<sub>resid</sub> measures and what distributions are being compared across the DH<sub>resid</sub> in the same figure. We hope this will serve as a more intuitive explanation of individuality and help readers interpret and follow more easily the results that we report after this figure.

      (3) In the same vein, we added another schematic visual representation to Figure 3 (Figure 3F) where we depict how distributions of individual behaviour may change with every decision and how this change translates to (or can be read out from) the change in residual individuality. We have also renamed the X axis of Figure 3E to “DH<sub>resid</sub> Start”, so that it is clearer what is measured here and matches the explanation in Figure 2D.

      (4) We have added two additional supplementary figures where the reader can inspect in more detail how the distributions of individual behaviour change longitudinally across time for each genotype in control and conditioned (Figure 3 – Figure supplement 2 and  Figure 3 – Figure supplement 3). From these figures one can glance how variance as well as the shapes of the distributions change as the flies learn in the conditioned setting, and how they remain largely the same in the control where flies behave spontaneously. We have added a sentence in the main text to introduce these figures: “We found that the distributions of individual behaviour were broader and their shapes changed substantially over the course of the experiment for the conditioned flies, and not for the control flies Figure 3 - figure supplement 2, Figure 3 - figure supplement 3).”

      (5) As noted in the first provisional response to reviewers, we changed the sentence “In every individual, behaviour is shaped by deterministic, genetic factors and by environmental events throughout lifetime, which may be stochastic and can occur at the molecular, cellular, organismal and even population scales.” to “In every individual, behaviour is shaped by fixed genetic factors and by variable environmental events throughout lifetime, which may be stochastic and can occur at the molecular, cellular, organismal and even population scales.”

      (6) Some sentences were edited in the results so that we can correctly refer to the newly added tables and figures. Context, meaning or interpretation of the results in these sentences was not altered.

      (7) While adding the new table references to the text, we noticed a typo that propagated in the previous version where the number of flies used in the main experiment was stated to be N= 5238, when in fact it should have been N=5239. This is now fixed.

      We once again thank the reviewers for their comments and suggestions – we believe their suggestions helped us improve the presentation and interpretability of our study and we hope the reviewers and readers will agree with this as well.

    1. eLife Assessment

      This short report is an important study that visual acuity declines nonlinearly with cone dropout, while eye motion partially compensates by improving sampling from remaining cones. The method for experimentally simulating cone dropout is compelling, leveraging state-of-the-art imaging and testing in human subjects.

    2. Reviewer #1 (Public review):

      The authors demonstrate an innovative approach to investigate the effect of cone dropout on visual acuity using their newly developed Oz platform. By systematically reducing the coverage of real-world input to the cone photoreceptor mosaic ("cone dropout condition"), the authors are able to assess how having less cones leads to reduced vision, in comparison to existing approaches ("pixel dropout condition").

      The observation of visual acuity maintenance with cone dropout has been a longstanding mystery since the 2013/2018 papers by Ratnam and Foote. The authors should be commended for their approach to address this important question. However, there are some simplifications and assumptions being applied to make this jump (i.e. that a 50% reduction in cone stimulation in a healthy eye is comparable to a 50% reduction in cone density in a patient). It seems unlikely that in a patient eye, with cone dropout, that there will be gaps in the mosaic. Not considering any other non-photoreceptor related reasons for visual acuity loss which can occur in patients, the cone aperture acceptance angle may be different due to changes in cone size or packing; the sensitivity of individual cones may also be reduced due to deficits in the visual cycle recovery which could be affected in disease. Some of these limitations could be addressed and acknowledged more explicitly.

      The capture of a rich dataset including both cone imaging and eye motion is valuable. Since the C stimulus test relies on foveal fixation, and there is a high degree of subject-to-subject variation in peak cone density, the authors may wish to report on peak cone density measurements of the subjects being included in this study. In addition, evaluating whether the eye motion is affected by simulated cone dropout condition can help to rule out whether these observed effects can be attributed to eye motion.

      Overall, this is an impressive study incorporating state-of-the-art technology to probe the fundamental limits of human vision.

      Comments on revised version.

      The authors have nicely addressed my concerns. The additional clarifications and revised text have strengthened the paper. Thank you also for pointing out the inaccuracy of referring to the system as the olo system; this has been corrected.

    3. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      The authors demonstrate an innovative approach to investigate the effect of cone dropout on visual acuity using their newly developed olo system. By systematically reducing the coverage of real-world input to the cone photoreceptor mosaic ("cone dropout condition"), the authors are able to assess how having fewer cones leads to reduced vision, in comparison to existing approaches ("pixel dropout condition").

      The capture of a rich dataset, including cone imaging and eye motion, is valuable. Benchmarking with the prior literature, suggesting that good visual acuity can be maintained despite a 50% loss in cone density, is impressive. However, it is known that cone density varies dramatically from the peak cone density location in the foveal center to even a location a few degrees outside of the fovea. In addition, there is a high degree of subject-to-subject variation in peak cone density. Given that the C stimulus is hollow in the middle, the stimulus does not actually hit the location of the peak cone density but must land slightly outside of it. Therefore, considering the actual cone density of where the stimulus lands will be important to discuss and/or analyze.

      The reviewer is correct that the cone density will vary dramatically with distance from the foveal center. However, importantly, in our experiment the Landolt C stimulus is fixed in the world and the eye is free to move across it. Therefore, the subject can direct their gaze to any part of the letter, rather than it being fixed to the hollow center of the letter. In the worst case, if the subject kept their gaze fixed at the center of the letter and were viewing the largest letter corresponding to the worst acuity measured in our experiments (20/100), the cone density on average would be 13% lower at the edge of the letter than at the center. However, it is unlikely that a subject would have fixated in this manner, and the vast majority of letters shown during the experiments were much smaller than 20/100.

      For completeness, we have calculated the peak cone densities for each of our subjects using the cone density centroid method described by Reiniger et al (2021). We have added these numbers and a description of the method to the Subjects section in Methods and Materials on lines 339-344.

      The observation of visual acuity maintenance with cone dropout has been a longstanding mystery since the 2013/2018 papers by Ratnam and Foote. The authors should be commended for their approach to addressing this important question. However, there are some simplifications and assumptions being applied to make this jump (i.e., that a 50% reduction in cone stimulation in a healthy eye is comparable to a 50% reduction in cone density in a patient). It seems unlikely that, in a patient's eye, with cone dropout, there will be gaps in the mosaic. Not considering any other non-photoreceptor-related reasons for visual acuity loss, which can occur in patients, the cone aperture acceptance angle may be different due to changes in cone size or packing; the sensitivity of individual cones may also be reduced due to deficits in the visual cycle recovery, which could be affected in disease. Some of these limitations could be addressed and acknowledged more explicitly.

      Cone loss does manifest differently in different retinal degenerative diseases, and in this work we implement dropout on a cone-by-cone level. To address the reviewer’s points, we have added a description of the range of spatial manifestations of cone loss across a range of diseases to the Discussion section on lines 320-328, and emphasize that we focus on one particular manifestation in this paper.

      Overall, this is an impressive study incorporating state-of-the-art technology to probe the fundamental limits of human vision.

      We thank the reviewer for their helpful comments and constructive feedback.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      The patient recruitment limitation seems to be a bit artificial here. This is indeed a limitation, but perhaps not the primary limitation or motivating factor. Consider removing/rephrasing this motivation.

      We agree with the reviewer’s comment, and have removed it from the abstract and removed its framing as a limitation in the Introduction section in two instances on lines 31 and 36-37.

      Was the peak cone density quantified, and the location of the peak cone density determined? Reporting the range of eccentricities over which the C stimulus lands relative to the peak cone density location, as well as the actual cone density that is being used to sample the C stimulus on the retina, seems to be important for contextualizing this study. It is a bit too simple to only consider the percentage of cones that are reduced.

      We have added the peak cone density for each subject to the Subjects section in Methods and Materials. This experiment did not require fixation; rather, the Landolt C stimulus was fixed in space and the subject could move their eye freely across it, meaning that different parts of the fovea may have sampled the letter on different trials. At the highest dropout percentage, where acuity was the worst, the letter size was 20/100 at threshold, or 25 arcmin. If the subject were to fixate with their peak cone density at the center of the letter, we have computed that the average decrease in cone density at the edge of the letter (12.5 arcmin away) would be 13%. The majority of trials in the experiment showed letters that were much smaller than this, and would have been subject to even less variation in cone density.

      Do the authors have any idea about the approximate size of the cones in healthy subjects compared to diseased eyes? Importantly, if the cones in patients are larger due to the dropout of their neighbors, then the retinal coverage area would be larger due to their larger size, and the amount of light that can be coupled into larger cones may also be larger. Can this be modeled or discussed?

      In our implementation, we did not emulate a change in cone size, and rather modeled the loss as discrete holes in an otherwise intact retina. We have added text to the Discussion section on lines 320-328 to make the distinction between this form of cone loss and other forms where cones appear to fill in for their neighbors resulting in a contiguous mosaic of lower density overall.

      Acknowledging some of the shortcomings of this approach for simulating the patient condition could be improved. It may be worthwhile to tone down the premise of this paper if these cannot be adequately explained.

      In order to tone down the premise of the paper, we have made the following changes to the text.

      We now emphasize on lines 65-69 in the Introduction section that we focus specifically on the impact of cone loss on acuity without modeling downstream factors.

      In addition to the description of other diseases that we added in response to a previous comment, we have also added the following text on lines 313-318 of the Discussion section:

      “... factors beyond the photoreceptors also play a role in shaping vision under retinal degeneration. In this work, we did not model any downstream factors such as shorter outer segments (Foote 2018), retinal rewiring (Jones 2016, Lee 2021), or ganglion cell hyperactivity (Kramer 2023). Instead, we sought to characterize vision in the presence of cone loss at the lowest possible level, considering only the decrease in sampling power at the retinal input.”

      In the methods, it is not completely clear the rationale for determining the appropriate size of the C stimulus. How is visual acuity determined if the C stimulus size is not changed?

      A more careful explanation of how the C stimulus size is set is warranted.

      In the experiments measuring visual acuity, the C stimulus size was selected by a QUEST staircase on each trial. For each dropout condition, we ran 4 interleaved QUEST staircases with 20 trials each. This is described in the “Acuity Threshold Experiment” section in the main text (lines 99-100) and in Materials and Methods (line 407). To clarify further, we have updated the following sentence on line 423:

      “For each condition, we ran 4 interleaved QUEST staircase procedures (Watson and Pelli, 1983) with 20 trials per staircase, which varied the size of the Landolt C on each trial.”

      What is the clinical visual acuity of the subjects being tested? It seems important to report this if the authors want to use their C stimulus as a proxy for clinical visual acuity.

      The subjects being tested have excellent acuity. In Figure 1, we can see that their adaptive-optics-corrected acuity for the baseline 0% dropout condition ranges from approximately 20/10 to 20/12.5 across the 4 subjects. We have added the following statement to the “Subjects” section in Materials and Methods (line 339):

      “All subjects self-reported to have normal vision.”

      Given that the title of the paper emphasizes the role of eye motion, it seems that a more careful analysis of the magnitude and type(s) of eye motion could be added. There are eye motion data provided in the supplemental figure, but it is not completely clear how this eye motion data is actually being used to derive meaningful information about visual acuity.

      We performed analyses to determine whether there seemed to be a significant difference in eye motion patterns between the cone and pixel dropout conditions, which was described in the section “Analysis of Eye Motion Data” and in Supplementary Figure S1. In that figure, we show that for all 4 subjects there is no significant difference in the iso-density contour area containing 68% of their eye motion data. We suggest in the paper that due to the pseudorandom presentation of trials and the limited duration of those trials, subjects were unlikely to adapt or adjust their eye movement, and that instead their natural eye motion served as a data collector that improved acuity.

      What is the accuracy of the eye motion and cone dropout stimulation delivery in the fovea? Given the small size of the cones, it seems that this is one of the most challenging locations of the eye to test with this new olo technology.

      Eye tracking and targeted light delivery are crucial in the AOSLO system and the reviewer is correct to point out that this is most difficult to achieve at the foveal center. To address this concern, we have done some simple modeling and have added the following text to the Cone-by-Cone Stimulation section in the Methods and Materials.

      “This latency, combined with other factors such as diffraction and residual aberrations, limit the ability to restrict the light to only the targeted cone. Considering a 543-nm focus through a 7.2 mm pupil, a random tracking error with a full-width-at-half maximum (FWHM) of 0.5 arcminutes (Harmening et al. (2014)), a 0.0125 diopter residual defocus error (maximum error given the step sizes of 0.025 diopters in the AOSLO defocus controller), an average cone spacing of 0.5 arcminutes (Wang et al. (2019)), and a Gaussian cone acceptance aperture with a FWHM that is 0.5 times the inner segment diameter (Macleod et al. (1992)), we estimate that each targeted cone receives 5.41 times more light than its nearest neighbor. This means that the ’dead’ cones cannot be fully excluded from the visual processing. Furthermore, the light leakage reported in Fong et al. (2025) further adds to the signal of non-targeted cones.

      Nevertheless, it is important to point out that the information about the stimulus (Landolt C in our case) is sampled at the targeted cone’s location and so, although nearby stimulated cones might detect light, they do not contribute to any increases in the sampling process. This is analogous to adding defocus blur to letters in the pixel dropout condition as neither situation will improve the spatial information.”

      References

      Reiniger, J.L., Domdei, N., Holz, F.G., Harmening, W.M.: Human gaze is systematically offset from the center of cone topography. Current Biology 31(18), 4188–4193 (2021)

    1. eLife Assessment

      This important study provides evidence that plateau pikas, at moderate densities, can facilitate yak nutrition by suppressing a poisonous plant, offering a helpful perspective on reciprocal interactions between small mammal ecosystem engineers and large herbivores. The evidence is convincing, supported by a manipulative field experiment and appropriate measurements of intermediary ecological processes. The work will be of interest to ecologists, conservation biologists, and rangeland managers, particularly those studying grassland herbivore interactions and livestock management.

    2. Reviewer #1 (Public review):

      Summary:

      This is important and significant work because it helps describe the complexity of interactions between system components where 2 herbivores interact with vegetation. Whereas other studies have shown that the larger ungulate (yaks, Bos grunniens, in this case) can facilitate the abundance and population growth of the smaller (the semi-fossorial lagomorph, Ochotona curzoniae, plateau pika hereafter), this study flips the tables, and shows that, at least under some conditions, moderate densities of the plateau facilitate the nutritional condition of yaks.

      Strengths:

      Notably, the strong inference the authors can claim for their results is supported by the careful experimental design. A weaker paper would have simply noted correlations between pika burrow density and yak feeding efficiency without experimental removal. This paper, to its credit, not only used experimental removals but also documented the various intermediary results that support the ultimate conclusions. The statistical approaches used appear to be appropriate. (Readers are encouraged to read the full Materials and Methods, which are available in the Supplementary Materials section).

      Weaknesses:

      Although the study was well designed and executed, and its conclusions appear strongly supported, readers interested in the management implications on the Qinghai-Tibetan Plateau should be mindful of its limitations. First, the study site, at approximately 3,200 m elevation, was relatively low by Qinghai-Tibetan Plateau standards. Stellera chamaejasme becomes less common at elevations > 4,000 m, where a majority of livestock grazing occurs. Thus, it would be instructive to learn, through follow-up studies, whether similar facilitation occurs where unpalatable (and mildly poisonous) species in such genera as Astragalus, Oxytropis, and Thermopsis replace S. chamaejasme as the problematic plant for pastoralists. Second, the authors make no mention of wild ungulates, so it is unclear what, if any, role they may have played in this system. At least one study in Qinghai Province, albeit at a slightly higher elevation, showed that not only pikas, but also Tibetan gazelles (Procapra picticaudata), which were commonly observed on grazed pastures, grazed more frequently on some dicots avoided by domestic sheep than did the livestock themselves (Harris et al. 2015). It would also be instructive to learn if similar facilitation as observed here applied to the other principal livestock species in the area, domestic sheep (which are often herded together with smaller numbers of domestic goats). Finally, as suggested by this study, the interactions between all components of the system are complex and interactive. If pika facilitation of yak nutrition at the densities documented results in herders increasing yak density, might the increased herbivory from the domestic animals provide the conditions for the pika population to increase beyond the densities observed here, and thus toward the levels where facilitation yields to competition?

    3. Reviewer #2 (Public review):

      This study uses a combination of field sampling and manipulative experiments to test for facilitative impacts of pikas on yaks via suppression of a poisonous forb. The authors found that, when Stellera forbs were present, yak weight increases over the growing season were greater in the presence of pikas compared to in their absence. This occurred because, although pikas do not consume Stellera, they clip it and use it in nest/burrow construction, thereby decreasing its relative abundance in the plant community. Thus, overall, the study contributes to our understanding of how herbivores of different size classes indirectly affect each other via use of shared resources.

      It is well known that large herbivores on grasslands impact smaller animals, but the reciprocal interaction is rarely tested. Thus, this study asks a valuable question, and the experiment is well-designed to test it. The authors also do a good of demonstrating the potential conservation impacts of their research.

    4. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      This important study provides evidence that plateau pikas, at moderate densities, can facilitate yak nutrition by suppressing a poisonous plant, offering a helpful perspective on reciprocal interactions between small mammal ecosystem engineers and large herbivores. The evidence is solid, supported by a manipulative field experiment and appropriate measurements of intermediary ecological processes, although some claims about density dependence, competition, and stress-gradient mechanisms are not fully supported by the experimental design. The work will be of interest to ecologists, conservation biologists, and rangeland managers, particularly those studying grassland herbivore interactions and livestock management on the Qinghai-Tibetan Plateau.

      Thank you very much for these positive assessments of our work. Below, we provide point-by-point responses to the comments from the two peer reviewers, and we hope these revisions are satisfied by you and the reviewers.

      Reviewer #1 (Public review):

      Summary:

      This is important and significant work because it helps describe the complexity of interactions between system components where two herbivores interact with vegetation. Whereas other studies have shown that the larger ungulate (yaks, Bos grunniens, in this case) can facilitate the abundance and population growth of the smaller (the semi-fossorial lagomorph, Ochotona curzoniae, plateau pika hereafter), this study flips the tables and shows that, at least under some conditions, moderate densities of the plateau facilitate the nutritional condition of yaks.

      The study was not designed to investigate the reasons that pikas clip Stellera chamaejasme. That said, based on other studies and general knowledge of the ecology of these pikas, it is likely that they clip (although do not eat) this plant because its relatively large size hinders predator detection. This species of pika does better where vegetation height is low than where it is higher.

      Strengths:

      Notably, the strong inference the authors can claim for their results is supported by the careful experimental design. A weaker paper would have simply noted correlations between pika burrow density and yak feeding efficiency without experimental removal. This paper, to its credit, not only used experimental removals but also documented the various intermediary results that support the ultimate conclusions. The statistical approaches used appear to be appropriate. (Readers are encouraged to read the full Materials and Methods, which are available in the Supplementary Materials section.)

      We appreciate these positive comments on our work.

      Weaknesses:

      Although the study was well designed and executed, and its conclusions appear strongly supported, readers interested in the management implications of the Qinghai-Tibetan Plateau should be mindful of its limitations. First, the study site, at approximately 3,200 m elevation, was relatively low by Qinghai-Tibetan Plateau standards. Stellera chamaejasme becomes less common at elevations > 4,000 m, where a majority of livestock grazing occurs. Thus, it would be instructive to learn, through follow-up studies, whether similar facilitation occurs where unpalatable (and mildly poisonous) species in such genera as Astragalus, Oxytropis, and Thermopsis replace S. chamaejasme as the problematic plant for pastoralists.

      Thank you for this suggestion. We have acknowledged this limitation in the Discussion by adding the paragraph below (see the Third point):

      “Despite of these, several questions deserve further investigation. First, our study examined pika–yak interactions only during the summer period, when food resources are most abundant. Whether such facilitative effects weaken or even shift toward competition under more stressful conditions—for example, when forage becomes limited during autumn or winter—remains to be tested. Second, if the documented facilitation of yak nutrition by pikas prompts herders to increase yak densities, could the resulting rise in livestock herbivory push pika populations beyond the levels observed here, potentially toward the threshold where facilitation gives way to competition? Third, our study site is located at approximately 3,200 m elevation, relatively low by Qinghai-Tibetan Plateau standards. Stellera becomes less common at elevations > 4,000 m, where a majority of livestock grazing occurs. It would be instructive to learn, through follow-up studies, whether similar facilitation occurs where unpalatable (and mildly poisonous) species in such genera as Astragalus, Oxytropis, and Thermopsis replace Stellera as the problematic plants for pastoralists (Lu et al., 2012; Li and Zhao, 2025). Finally, it is unclear whether similar facilitation as observed here applied to the other principal livestock species in the area, such as domestic sheep and goats.”

      See these revisions in Line 272-286 in the Discussion section.

      Second, the authors make no mention of wild ungulates, so it is unclear what, if any, role they may have played in this system. At least one study in Qinghai Province, albeit at a slightly higher elevation, showed that not only pikas, but also Tibetan gazelles (Procapra picticaudata), which were commonly observed on grazed pastures, grazed more frequently on some dicots avoided by domestic sheep than did the livestock themselves (Harris et al. 2015).

      Citation:

      Harris RB, Wang, WY, Badinqiuying , Smith AT, Bedunah DJ (2015) Herbivory and Competition of Tibetan Steppe Vegetation in Winter Pasture: Effects of Livestock Exclosure and Plateau Pika Reduction. PLoS ONE 10(7): e0132897. doi:10.1371/journal.pone.0132897

      Thank you for this suggestion. We have added more details about the study site, particularly regarding wild ungulates, in the Methods section. Specifically, we have included the sentence of “Wild ungulates, such as Tibetan gazelles (Procapra picticaudata) (Harris et al., 2015), and other small mammals such as rabbits and zokors, occur rarely in the area.”

      See these revisions in Line 333-335 in the Methods section.

      It would also be instructive to learn if similar facilitation as observed here applied to the other principal livestock species in the area, domestic sheep (which are often herded together with smaller numbers of domestic goats).

      Thank you for the suggestion. We have acknowledged this limitation in the Discussion, by adding a paragraph as: “Finally, it is unclear whether similar facilitation as observed here applied to the other principal livestock species in the area, such as domestic sheep and goats.”

      See these revisions in Line 284-286 in the Discussion section.

      Finally, as suggested by this study, the interactions between all components of the system are complex and interactive. If pika facilitation of yak nutrition at the densities documented results in herders increasing yak density, might the increased herbivory from the domestic animals provide the conditions for the pika population to increase beyond the densities observed here, and thus toward the levels where facilitation yields to competition?

      Thank you for your suggestion. We have acknowledged this limitation in the Discussion, by adding the paragraph as “Second, if the documented facilitation of yaks by pikas prompts herders to increase yak densities, could the resulting rise in livestock herbivory push pika populations beyond the levels observed here, potentially toward the threshold where facilitation gives way to competition (Yang et al., 2026)?”

      See these revisions in Line 276-279 in the Discussion section.

      Reviewer #1 (Recommendations for the authors):

      Although no doubt a bit sensitive, it would have been better to reveal a bit more about how pikas were removed.

      We have provided more details about how pikas were removed in the no-pika treatment, by adding “For the no-pika treatment, pikas were trapped once every two weeks using 30 live traps (25 cm high × 25 cm wide × 40 cm long) within each plot and relocated elsewhere in the study site.” in the Methods section. We didn’t recorded how many pikas were removed from the corresponding plots, so no data were available for this point.

      See these revisions in Line 411-413 in the Methods section.

      The authors also missed a few relevant papers worth citing, including Badingqiuying, R. B. Harris, and A. T. Smith. 2018. Summer habitat use of plateau pikas (Ochotona curzoniae) in response to winter livestock grazing in the alpine steppe Qinghai-Tibetan Plateau. Arctic, Antarctic, and Alpine Research 50 (1): e1447190

      We have cited this key paper in Line 75 in the Introduction section.

      Reviewer #2 (Public review):

      Summary:

      This study uses a combination of field sampling and manipulative experiments to test for facilitative impacts of pikas on yaks via suppression of a poisonous forb. The authors found that, when Stellera forbs were present, yak weight increases over the growing season were greater in the presence of pikas compared to in their absence. This occurred because, although pikas do not consume Stellera, they clip it and use it in nest/burrow construction, thereby decreasing its relative abundance in the plant community. Thus, overall, the study contributes to our understanding of how herbivores of different size classes indirectly affect each other via the use of shared resources.

      Strengths:

      It is well known that large herbivores on grasslands impact smaller animals, but the reciprocal interaction is rarely tested. Thus, this study asks a valuable question, and the experiment is well-designed to test it. The authors also do a good job of demonstrating the potential conservation impacts of their research.

      We appreciate these positive comments on our work.

      Weaknesses:

      What the authors tested is really cool, but their claims go far beyond what they can say based on their experimental design. For example, the authors claim to show that pika impacts on yaks display density-dependent transitions from competition to facilitation. However, their experiment only looked at the presence (at moderate densities) and absence of pikas, and they only tested for facilitation, not competition.

      The paper would also benefit from changes to the framing in the introduction and discussion. For example, the authors pitch the work as a test of the stress-gradient hypothesis. However, there is no abiotic stress gradient in the study, which is an essential component of the SGH. They also pitch the work in terms of density dependence, but there is no significant variation in population densities beyond the presence-absence binary. The paper would be stronger if they focused their framing around the literature on facilitative interactions across mammals of different size classes, especially indirect facilitation via use of shared resources, which is what this paper is really about.

      We agree that our work had explored only the facilitative effects of pikas on yaks, rather than the Stress Gradient Hypothesis (SGH). Thus, we deleted the description on SGH. However, the finding of a humped relation between yak weight gains and pika burrow densities (Figure 3C) is very important which provides evidence that moderate densities of pikas has the best beneficial effects on yak growth. We added a separate paragraph in discussion to have a clear discussion.

      We have made the major revisions below to address these concerns.

      (1) We have revised the title into “Small mammalian herbivores at moderate densities facilitate livestock growth by improving vegetation composition in grasslands ”.

      (2) We have deleted all the statements about facilitation and competition predicted by the SGH in the Abstract (Line 56-59), Introduction (Line 88-91), Discussion (Line 231-233, the whole paragraph about SGH was removed here), and the References sections.

      (3) We added a paragraph in discussion (Line 248-259) to have a clear discussion on the humped relation between yak weight gains and pika burrow densities as “Because of the natural variations in pika density in the pika-present treatment, we were able to obtain a hump-shaped relationship between yak weight gains and pika burrow densities in these plots. Compared with the absence of pikas, the facilitative effect reached its maximum at approximately 200 burrows/ha but became competitive at densities exceeding 400 burrows/ha (Figure 3C). This result reveals that pika density modulates the net outcome for yak weight gain, with a facilitation peak at ~200 burrows/ha and a competition onset above 400 burrows/ha. Our findings offer empirical evidence for the non-monotonicity theory, under which the competition-facilitation balance varies with population density: facilitation dominates at low densities, competition at high densities, and these density-dependent shifts may underpin community stability and productivity (Zhang, 2003; Zhang et al., 2015). The theory further holds that the facilitation threshold, not the competition-facilitation transition, is the critical factor governing the stability of interacting species or communities (Zhang et al., 2015).”

      Most importantly, there are inconsistencies in what is visualized in the figures compared to the model results. For example, the results section in several places notes a lack of significant interaction terms in the model but shows interactions in the p-values on the figures.

      In the Results section, there are only two places where we discussed non-significant interactions: Line 175–177 “Pikas and Stellera had no interactive effects on abundance of sedges, forbs, and neutral detergent fiber (NDF) of total forage for yaks (Figure 3F, I and Appendix 1—figure 1, table 5,8).” and Line 190–192 “Pikas and Stellera had no interactive effects on yaks’ foraging efficiency on forbs (Appendix 1—figure 2, table 10).”.

      We have cross-checked both the Results section and the Figures sections mentioned above, and confirmed that they are consistent now.

      The authors also plot smoothed lines rather than their model results and then draw interpretations from those lines that cannot be tested in the models that they used.

      Thank you for the suggestion. Now we have added the Appendix 1—table 3 and Appendix 1—table 7 for the model results of generalized additive models (GAMs) for Figure 2C and Figure 3C that plotted with smoothed lines in Appendix 1.

      There are also missing details that are important for model interpretation, including the distributions used and the sample sizes.

      We have provided the Appendix 1—table 13 to summarize all statistical models used in the study, including the distributions used and the sample sizes in the Appendix 1.

      We have also added a sentence of “A summary of all statistical models used in the study is available in Appendix 1 table 13.” in Line 475-476 in the Statistical analyses section to indicate this information.

      Another major concern with experimental design is in the forage nutrient analyses. The authors picked plants along a grazing trail, then measured nutrient content without standardizing based on plant species, so any differences across treatments could be because of what they happened to grab rather than overall forage quality.

      We have revised this section to provide more details on how forage samples were collected and their quality were analyzed. Specifically, five forage samples were collected per grazing plot, focusing on the two dominant plant species—one sedge and one grass—that were most frequently grazed by yaks. To ensure comparability across plots and treatments, we mixed the two species at equal dry mass (5 g). We have revised this section as below.

      “To assess forage quality, five forage samples were collected from each grazing plot to quantify their nutritive values. To obtain samples that reflect the forage actually consumed by yaks, we tracked the animals along their grazing paths and collected the plant tissues of the two most frequently consumed species: the dominant sedge Kobresia humilis and the dominant grass Elymus nutans (Figure 2B; Pan et al., 2019). The collected tissues of each species were dried in a forced-air oven at 60 °C for 48 h, then ground through a 1-mm mesh. Subsequently, 5 g of each dried and ground species were combined in a 1:1 dry mass ratio, and the resulting mixture was stored in plastic bags for subsequent analyses.”

      See these revisions in Line 439-447 in the Methods section.

      Reviewer #2 (Recommendations for the authors):

      (1) Introduction

      Line 53 - I wouldn't describe small mammals like rodents as keystone species. They can have strong impacts on ecosystems, but not disproportionate relative to population size, which is a key part of that definition. It may be true when you talk specifically about pikas later on, but not small mammals as a general category.

      We have replaced this term with “consumers” here, see Line 63.

      Lines 58-61 - Good hook

      Thank you for this positive comment.

      Lines 69-74 - I don't think the stress gradient hypothesis is the right pitch for this work. The SGH posits that facilitation increases with abiotic stress, but no abiotic stressors were measured in this study. Population density interacts with abiotic factors in the SGH, but population density in and of itself is not an abiotic stressor. The papers you cite here all look at the interaction between population density and abiotic stressors (e.g., water availability). So these lines set me up to expect a stress gradient in your experiment that didn't exist, then left me confused later on. It would be better to highlight the strength of your work (lines 70-76 pose interesting questions and predictions) rather than trying to make it fit within the SGH.

      Thank you for pointing out this problem. We have deleted the statements about SGH here.

      See these revisions in Line 88-91 in the Introduction section.

      (2) Materials and Methods

      Lines 498-509 - Did you verify beforehand that no plants within the enclosures had been grazed on? How did you know that the consumption or clipping was specifically from those pikas?

      We have clarified here by adding “Before cage installation, we carefully checked the plants within each plot and removed those that had been previously grazed or damaged by herbivores.” in the Methods section. In this case, we made it sure that the consumption or clipping was specifically from those pikas with the cages.

      See the revisions in Line 349-350 in the Methods section.

      Lines 509-512 - I would be careful calling this preference. It's really just a record of what they consumed along paths they were walking, which could be about accessibility and convenience as much as preference.

      We have replaced “diet preferences” with “diet composition” here, see Line 359 in the Methods section.

      Line 514 - Clarify the specific question or hypothesis you're testing with this field survey, beyond just generally testing associations.

      We have modified the sentences here as “In July 2021, we investigated the potential facilitation of pikas on yaks mediated by the poisonous Stellera forbs under unmanipulated field conditions in the study site.”.

      See Line 365-366 in the Methods section.

      Lines 532-534 - The intro for this paper sets it up to be about density-dependent movement from competition to facilitation, but the experiment here is set up to compare pika presence/absence. The framing of the paper needs to be adjusted to better align with this experimental design.

      The same issue as mentioned above. We agree that our work had explored only the facilitative effects of pikas on yaks, rather than the balance between competition and facilitation as predicted by the Stress Gradient Hypothesis (SGH). However, the finding of a humped relation between yak weight gains and pika burrow densities (Figure 3C) is very important which provides evidence that moderate density of pika has the best benefical effect on yak. We added an separate paragraph in discussion to have a clear discussion about this point in the Discussion section.

      We have made the major revisions below to address this concern.

      (1) We have revised the title as “Small mammalian herbivores at moderate densities facilitate livestock growth by improving vegetation composition in grasslands”.

      (2) We have deleted the statements about facilitation and competition and the SGH in the Abstract (Line 56-59), Introduction (Line 88-91), Discussion (Line 231-233, the whole paragraph about SGH was removed here), and the References sections.

      (3) We kept the discussion on the humped relation between yak weight gains and pika burrow densities (Figure 3C). We added an separate paragraph in discussion to have a clear discussion about this point. For details, see Line 248-259 in the Discussion section.

      Lines 566-569 - Why did you need to simulate pika clipping when you already had pika presence/absence treatments?

      We conducted Stellera removal treatment by simulating poisonous plant clipping behaviors of pikas because we want to confirm that the shifts in abundance of this dominant poisonous plant species is the key mechanism in driving pika-yak facilitation in our system. If we simply looked the differences in yak weight gain in the pika presence/absence treatments, it should be difficult to secure the underlying mechanism. In addition to the reduction in abundance of the poisonous Stellera, pikas may cause a variety of shifts vegetation properties including plant productivity and diversity, and soil disturbances that can exert direct and indirect effects on yak foraging activities, and thus their weight gains.

      Lines 566-569 - Is there another citation you can give showing that Stellera forbs taller than 20cm are both preferred by pikas and exert greater impacts on plant/animal communities? Those are big assumptions that need to be better supported or explained. If there was a logistical reason that you didn't remove all of the smaller forbs, that needs to be laid out as well.

      We have added one new citation here to support this method here. We have revised this section as “To simulate the clipping behavior of pikas, we clipped only those Stellera forbs exceeding 20 cm in height. This threshold was chosen based on previous observations that pikas preferentially target large forbs of this size (Liu et al., 2009).”

      Liu W, Zhang Y, Wang X, Zhao JZ, Xu QM, Zhou L. 2009. The relationship of the harvesting behavior of plateau pikas with the plant community (In Chinese). Acta Theriologica Sinica 29:40-49.

      Also, we have deleted the sentence of “and can exert significant impacts on the plant community and on yak grazing behaviors (Z.Z., field observations)” mentioned above, because these patterns were observed only by the authors in the field and lack supporting data.

      See these revisions in Line 419-422 in the Methods section.

      Line 587 - sample size per treatment? Was it consistently one species that you measured, and if so, which one? If you collected different species or a mix of species across treatments, then you can't really compare the nutrient values because you haven't accounted for interspecific variation.

      We have revised this section to provide more details on how forage samples were collected and their quality were analyzed. Specifically, five forage samples were collected per grazing plot, focusing on the two dominant plant species—one sedge and one grass—that were most frequently grazed by yaks. To ensure comparability across plots and treatments, we mixed the two species at equal dry mass (5 g).

      We have revised this section as below:

      “To assess forage quality, five forage samples were collected from each grazing plot to quantify their nutritive values. To obtain samples that reflect the forage actually consumed by yaks, we tracked the animals along their grazing paths and collected the plant tissues of the two most frequently consumed species: the dominant sedge Kobresia humilis and the dominant grass Elymus nutans (Fig. 2B; Pan et al., 2019). The collected tissues of each species were dried in a forced-air oven at 60 °C for 48 h, then ground through a 1-mm mesh. Subsequently, 5 g of each dried and ground species were combined in a 1:1 dry mass ratio, and the resulting mixture was stored in plastic bags for subsequent analyses.”

      See these revisions in Line 439-447 in the Methods section.

      Lines 603-613 - Please provide the dependent variables in each model, as well as any interaction terms, in addition to the random effects. Please also state explicitly what you were trying to test with each of these models.

      There are two models that use the tweedie family (forb and sedge bite rate). Indeed, we need to include the tweedie power parameter to help understand the mixture of the three families. We have included the p-value (power) in the Appendix 1—table 13 in the Appendix 1. We chose to use tweedie because a normal gaussian family model fitted the results poorly.

      We have added the Appendix 1—table 13 in the Appendix 1, which provides the summary of all statistical models used in the study, including response variables, model type, distribution family (with Tweedie power parameter where applicable), interaction terms, random effects structure, and sample sizes.

      Also, we have added a sentence of “A summary of all statistical models used in the study is available in Appendix 1—table 13.” in Line 475-476 in the Statistical analyses section to indicate this information.

      Lines 613-614 - Tweedie is a category of distributions that includes quite a few different options, including Gaussian, Poisson, and Gamma distributions, some of which are normal and some of which are not. So, justifying the use of Tweedie distributions in your model structure doesn't really make sense, and it doesn't really tell me which distribution each model pulled from. Please clarify specifically which distributions you used for each model and why.

      We have now added the reason why we used Tweedie distributions by adding the sentence of “There were two models that used the tweedie family (forb and sedge bite rate). We chose to use tweedie because a normal gaussian family model fitted the results poorly” in Line 468-470 in the Statistical analyses section.

      We have also included the p-value (power) for forb and sedge bite rate in the Appendix 1—table 13 in the Appendix 1.

      Lines 615-618 - Provide citations for R packages described in the text.

      We have provided all the related citations for all R packages described in the text, as listed below.

      glmmTMB: Brooks, M. E., Kristensen, K., van Benthem, K. J., Magnusson, A., Berg, C. W., Nielsen, A., Skaug, H. J., Maechler, M., & Bolker, B. M. (2017). glmmTMB Balances Speed and Flexibility Among Packages for Zero-inflated Generalized Linear Mixed Modeling. The R Journal, 9(2), 378–400. https://doi.org/10.32614/RJ-2017-066

      mgcv: Wood, S.N. (2017). Generalized Additive Models: An Introduction with R (2nd edition). CRC Press. AND Wood, S.N. (2011). Fast stable restricted maximum likelihood and marginal likelihood estimation of semiparametric generalized linear models. Journal of the Royal Statistical Society: Series B (Statistical Methodology), 73(1), 3–36. https://doi.org/10.1111/j.1467-9868.2010.00749.x

      DHARMa: Hartig, F. (2022). DHARMa: Residual Diagnostics for Hierarchical (Multi-Level / Mixed) Regression Models. R package version 0.4.6. https://CRAN.R-project.org/package=DHARMa

      tidyverse: Wickham, H., Averick, M., Bryan, J., Chang, W., McGowan, L.D., François, R., Grolemund, G., Hayes, A., Henry, L., Hester, J., Kuhn, M., Pedersen, T.L., Miller, E., Bache, S.M., Müller, K., Ooms, J., Robinson, D., Seidel, D.P., Spinu, V., Takahashi, K., Vaughan, D., Wilke, C., Woo, K., & Yutani, H. (2019). Welcome to the tidyverse. Journal of Open Source Software, 4(43), 1686. https://doi.org/10.21105/joss.01686

      See these revisions in Line 472-475 in the Statistical Analyses section.

      (3) Results

      Line 127 - Somewhere in the intro or methods, describe what clipping is and why the pikas do it.

      We have provided more details about the clipping behaviors of pikas by adding “Notably, pikas often clip (although do not eat) the wolf poison S. chamaejasmehas because its relatively large size hinders predator detection (Fan et al., 1998).” in Line 330-332 in the Methods section.

      Line 128 - Change from "preferred" to "consumed greater proportions of"

      See the correction in Line 146 in the Results section.

      Line 136 - You measured yak weight once a month, so give the result in monthly weight gain rather than daily.

      Here we preferred to keep the unit of daily weight gains (converted from monthly ones), as this is the standard presentation for livestock growth performance, also see Fig. 1 in Odadi et al., 2011 Science’s paper.

      W. O. Odadi, M. K. Karachi, S. A. Abdulrazak, T. P. Young, Science 333, 1753–1755 (2011).

      Lines 138-140 - The linear models, as you described them in the methods (lines 603-620), don't test for hump-shaped relationships. Please update the methods to explain how you tested the density relationship and how you got this interpretation.

      The description for Figure 3C here showed estimates from a GAM (family: gaussian), and we did not use a linear model in this figure.

      We have now added a new model summary Appendix 1—table 7 for this Figure 3C in the Appendix 1.

      Line 145 - I don't think you can claim that the total available forage was more nutritious for yaks, as you picked plants that yaks happened to be chewing along a grazing path. You would need to take samples from a consistent set of plant species at random locations to make this claim.

      Sorry for this confusion. We modified “the total available forage” as “the major available forage” here (see Line 172) because we collected the same forage plant species and analyzed their nutrients.

      The same as mentioned above, to clarify the sampling methods, we have also revised this point in the Methods section to provide more detail on how forage samples were collected and their quality were analyzed. See these revisions in Line 439-447 in the Methods section.

      (4) Discussion

      Line 166 - Need to address inconsistency in how you talk about pika density. Here you talk about the impacts of pikas at moderate densities, which I think is a fair claim. Elsewhere, you talk about density-dependence, which I don't think you really measured, given that all your sampling was either in the absence of pikas or within a narrow window of densities that can all be categorized as moderate.

      Done! As mentioned above, we have removed the term of “density-dependence” in the whole manuscript, but keep the term of “moderate density” in the Abstract (Line 56-59), Introduction (Line 88-91), Discussion (Line 231-233, the whole paragraph was removed here), and the References section.

      Line 176-178 - This is a really cool finding.

      Thank you for this positive comment.

      Lines 179-181 - You didn't test competition between plant species, and you didn't measure light, soil moisture, or soil nutrients. So you can suggest competition as a potential mechanism, but you can't say definitively that's what is happening.

      We now have lowered our tone here as “We speculate that these improvements in food availability and nutrition for yaks may be due to the release of grasses and sedges from competition with the forbs for limiting above- and below-ground resources”.

      See the revision in Line 208-211 in the Discussion section.

      Lines 184-186 - You did a good job of it here, suggesting a likely potential mechanism at play without claiming it is for sure happening when it hasn't been measured.

      Thank you for this positive comment!

      Lines 188-199 - Strong paragraph. The impacts of large herbivores on smaller animals are well-studied, but reciprocal impacts are often overlooked.

      Thank you for this positive comment!

      Lines 201-205 - You didn't test the stress gradient hypothesis because there was no abiotic gradient. You also did not take any measurements during outbreaks, so you cannot claim to have compared low-moderate to outbreak pika densities. I think the paper would be much stronger if you removed the stress gradient hypothesis and instead focused more on the literature around facilitation between mammals of different body sizes, as you do in lines 205-209.

      We agreed that our work didn’t specifically design to test the stress gradient hypothesis (SGH) between pikas and yaks, so we have deleted this paragraph here, see Line 231-233 in the Discussion section.

      Line 209-215 - Again, you didn't test a competition-facilitation balance because you never tested or demonstrated competition. One of the main strengths of this paper is demonstrating facilitation, so build on that strength rather than referencing things you didn't measure.

      The same as mentioned above, we have deleted this paragraph here, see Line 231-233 in the Discussion section.

      Lines 218-222 - Not an accurate description of the relationship between herbivore diet and body size. Larger herbivores typically tolerate lower-quality plants in order to consume sufficient calories, but plenty of them do this via mixed feeding. Grazing in large herbivores and livestock is usually due to specifics of the digestive tract (e.g., hindgut fermentation) rather than specifically about body size.

      We have deleted this description of the relationship between herbivore diet and body size here. Instead, we have modified this statement as “The coexistence of a diverse of herbivore species with different diet selections and size classes can lead to an “compensatory effect” on grass and forb biomass that helps to maintain a balance and diverse plant community” in the Discussion section.

      See these revisions in Line 234-237 in the Discussion section.

      Lines 245-251 - Paragraph addresses an important point. Lines 247-249, though, overstate what you measured. There's no measurement of livestock production or biodiversity in the study.

      We have replaced the term of “livestock production and biodiversity” with “livestock growth performance” here. See Line 288-298 in the Discussion section.

      (5) Figures

      Figure 2C-D - This applies to all figures, but you need to plot the best-fit line generated from your model instead of using geom_smooth, which is what these lines look like. You can do this using functions like predict or ggpredict. Using these smoothed lines implied non-linear relationships that you didn't actually test for.

      We have redrawn Figures 2C and 2D to use model estimates directly and have included the related model summaries as Appendix 1—table 3 and table 4 in the Appendix 1.

      Figure 3B - This figure doesn't show yak weight gain in the presence of pikas. Instead, it shows weight loss when pikas are absent. It's a subtle difference, but very important for interpretation. Yak can maintain weight just fine without pikas as long as Stellera are absent, too. Your results consistently show no Pika x Stellera interactions, but that doesn't match your significance values here. Need to double-check and explain that.

      We have revised the descriptions for Figure 3 and 4 in the Results section, by emphasizing that the absence of pikas REDUCED weight gains of yaks, INCREASED toxic plant abundance, and REDUCED the quantity and quality of palatable grasses and sedges. We have revised these sections as below:

      Abstract (see Line 51-53)

      “Compared to the pika-present treatment, pika removal dramatically increased cover of the poisonous Stellera forbs by two-fold, reducing the abundance and protein content of palatable grasses and sedges, yak foraging efficiency, and yak weight gain by up to 42%.”

      Results (see Line 153-167, Line 169-175, Line 184-192)

      Also, we did find significant Pika x Stellera interactions for yak weight gains, we have provided these details in the Appendix 1—table 5 and table 6 in the Appendix 1.

      (6) Recommendations

      Lines 245-257 - Would recommend combining the last two paragraphs into one.

      We have combined the last two paragraphs, see Line 288-298 in the Discussion section.

      Line 530 - Can you replace large with a more precise measure of area?

      We are unable to provide a more precise measure of area here, so we have deleted the description of “in a large area”, but we have also added the note of “in the study site” by the end of the sentence to better describe the location of the plots.

      See the revision in Line 413 in the Methods section.

      Line 541 - Does this mean +/- 7.8 standard deviations? If so, how big is that range in kg?

      Here should be “115±7.8 kg”, we have done this correction in Line 392 in the Methods section.

      Figure 3 - Would be helpful to use "Stellera/pika present" and "Stellera/pika absent" rather than saying "No Stellera/pika" since you also use the No. abbreviation for numbers a lot in this figure.

      We have revised all the related Figures for this issue in Figure 3, 4, and Appendix 1—figure 1,2.

      Figure 3F and I - Show letters for significance on these two plots as well, even if it is just a row of a's.

      We have redrawn Figure 3F,I to address this issue.

      Figure S1 - Applies to all boxplots. Be consistent about showing significance, even if it is a row of a's indicating no difference between treatments.

      We have re-drawn Appendix 1—figure 1,2 to address this issue.

      Table S1 - Something happened with the line numbers, so they are in the table instead of on the left side.

      We have fixed this problem for Appendix 1—table 1 in the Appendix 1.

      Table S1 - Applies to all tables. Include the type of model that you ran (including distribution if not Gaussian) in the table legend.

      We have provided an summary of all statistical models used in the study in Appendix 1—table 13 in the Appendix 1.

      Also, we have added a sentence of “A summary of all statistical models used in the study is available in Appendix 1—table 13.” in Line 475-476 in the Statistical analyses section to indicate this information.

      Table S9 - This legend has a good description, including the type of model you used and what you were testing. Apply this more detailed legend to the rest of the tables.

      Again! We have provided an summary of all statistical models used in the study in Appendix 1—table 13 in the Appendix 1.

      Also, we have added a sentence of “A summary of all statistical models used in the study is available in Appendix 1—table 13.” in Line 475-476 in the Statistical analyses section to indicate this information.

    1. eLife Assessment

      This important study provides new insights into the regulation of cell organization and division in Trypanosoma brucei through the phosphorylation-dependent control of a kinesin motor protein by a polo-like kinase. The authors present convincing evidence, combining rigorous biochemical, cell biological, and imaging analyses, demonstrating that phosphorylation modulates kinesin localization and function, thereby influencing cellular organization and cytokinesis. The findings advance our understanding of the molecular mechanisms governing trypanosome cell division and will be of broad interest to researchers studying trypanosomes, cytoskeletal regulation, and eukaryotic cell division.

    2. Reviewer #1 (Public review):

      I thank the authors for the revised manuscript and for the detailed responses.

      I think the main points raised in the review have now been addressed. In particular, the new experiment with TbPLK inhibition and mass spectrometry is an important addition, as it provides direct evidence that phosphorylation of KIN-G at Thr301 and Ser569 depends on TbPLK activity in cells.

      I also appreciate that the authors have toned down the interpretation of the Golgi phenotype. The revised text now makes clear that the fluorescence data show altered Golgi/ERES organization or duplication, but do not prove a structural defect in Golgi biogenesis.

      The added discussion of the T301A result is also helpful. The finding that only a small fraction of KIN-G is phosphorylated at Thr301 in asynchronous cells makes the lack of a strong T301A phenotype more understandable.

      Overall, I am happy with the revision of the beautiful manuscript.

    3. Reviewer #2 (Public review):

      Summary:

      The authors identify KIN-G as an in vitro substrate for phosphorylation by TbPLK and show that several of the in vitro P-ated sites, including T310, overlap with P-ation sites seen in live cells. The authors further show that PLK-mediated P-ation inhibits KIN-G binding to microtubules in vitro, as does a KIN-G-T301D mutant, and that expression of a KIN-G-T301D Phospho-mimic in T. brucei phenocopies KIN-G RNAi knockdowns, producing defects in cell division, morphogenesis of the centrin arm, FAZ and other cellular structures, as well as misplaced cytokinesis furrow.

      Understanding cytoskeletal rearrangements that drive cell division in T. brucei is an important and unresolved problem, so the work addresses important questions that are of great interest. PLK and KIN-G have previously been shown to be important for cell division and morphogenesis of cytoskeletal structures that drive cell division in T. brucei. The current work advances our understanding by suggesting a potential mechanism by which PLK and KIN-G might participate, namely through PLK-dependent P-ation to control KIN-G MT binding activity.

      Strengths:

      The authors use a rigorous combination of biochemistry, phosphoproteomics, cell biology, and mutant analysis to support their conclusion that PLK-mediated P-ation of KIN-G negatively regulates KIN-G microtubule binding and this may explain the observation that a KIN-G T301 phosphomimic mutant blocks cell division and perturbs biogenesis of cytoskeletal structures that drive cell division and morphogenesis. Combining rigorous and informative in vitro studies with mutant analysis in live cells is a great strength. The work is solid and important, though a few pieces are needed to fully connect the in vitro findings with the in vivo observations, as detailed below.

      Weaknesses:

      Overall, I find this work to be solid, and to provide an important addition to our understanding of mechanisms controlling cell division in T. brucei. The biochemistry, in particular, is rigorous and convincingly demonstrates PLK can P-ate KIN-G, altering its MT-binding ability. Analysis of phospho-mutants of KIN-G in live T. brucei support the conclusion that P-ation of KIN-G at T301 negatively affects KIN-G function in vivo. I think, however, that the results fall short of supporting the title, because, although the data convincingly show that PLK can phosphorylate KIN-G at T301 in vitro, and that T301 is P-ated in vivo, they do formally demonstrate (nor even test) whether PLK is the kinase responsible for this phosphorylation in vivo (experiments to address this seem quite feasible). I also do not see where the authors try to reconcile the absence of phenotype for KIN-G-T301A with the implied importance of KIN-G phosphorylation by PLK in cell division, which calls into question the need for P-ation of KIN-G-T301 in cell division. Suggestions for addressing these concerns are provided below.

      My two main questions are:

      (1) What is the biological relevance of KIN-G P-ation at T301?<br /> a. The authors report no defect for the KIN-G-T301A mutant, so what then is the need for T301 P-ation, if the cell gets along fine without it? One step toward addressing this would be to ask what fraction of KIN-G shows P-ation at T301. Although published studies indicate P-ation at T301, it isn't known what percentage of KIN-G in the cell is P-ated. One might anticipate, for example, that T301-P is a small minority of the population in asynchronous cultures and that T301 P-ation increases at specific cell cycle stages.<br /> b. Published work links PLK to cell division, FAZ elongation, etc... The current work suggests that one role of PLK is to P-ate KIN-G at T301. In contrast, however, the current work also indicates that P-ation of KIN-G at T301 is unnecessary for normal cell division, FAZ elongation, etc....<br /> c. Some experiments or at least commentary on points a and b above would strengthen the paper.<br /> - The authors have now addressed this question by assessing what % of KING is phosphorylated at T301 and adding commentary on this point in the revised paper.<br /> - I would suggest that the model (new figure 8) include a dephosphorylation step, as that is proposed by the authors in the text. Also include in the legend some commentary on the role of phosphorylation, which is the center point of this paper, but not currently mentioned.

      (2) Is PLK the kinase that P-ates Kin-G T301 in vivo?<br /> a. The authors show PLK P-ates T301 (and other residues) in vitro, and that T-301 is P-ated in vivo. To bring the analysis full circle, it would be informative to examine KIN-G P-ation in a PLK mutant or upon inhibition of PLK with published inhibitors. This seems to be a very doable experiment with the tools available.<br /> - The authors have addressed this question by demonstrating that T301 phosphorylation is reduced upon treatment with a PLK inhibitor, thus supporting that PLK phosphorylated T301 in vivo. It is noted that one might consider an alternate kinase is also able to phosphorylate T301 in absence of PLK activity, as that could explain the relatively low (~27%) reduction in phosphorylation by PLK inhibitor treatment.

    4. Reviewer #3 (Public review):

      Summary:

      Here the authors investigate the role of the Trypanosoma brucei polo-like kinase TbPLK in the function of flagellum-associated cellular structures in trypanosomes. They set out to test the hypothesis that a key substrate of TbPLK is the kinesin protein KIN-G, and that TbPLK phosphorylation of KIN-G regulates its functions in cells.

      Strengths:

      Using in vitro biochemistry with purified proteins, the authors convincingly demonstrate that TbPLK phosphorylates KIN-G at 29 sites. Moreover, they convincingly show that phosphorylation at one site, T301, impairs the binding of purified KIN-G to purified microtubules. They further confirm that inhibition of TbPLK in cells reduces KIN-G phosphorylation at T301 (and S569). Using immunofluorescence-based imaging approaches, they also show that TbPLK colocalizes with KIN-G at centrin arms during early S-phase of the cell cycle. Centin arms are structures that are located near the basal body and flagellum and are important for new flagellum biogenesis, Golgi positioning, and cell division. To evaluate the function of KIN-G phosphorylation in cells, they depleted KIN-G by RNAi, simultaneously expressed phospho-mimetic (T301D) and phospho-ablative mutant proteins, and used immunofluorescene to examine the impact on flagellum-associated cellular structures. They show that expression of the phospho-mimetic mutant KIN-G-T301D causes the following defects: reduced cell proliferation, disruption of centrin arm and Golgi biogenesis, impairment of FAZ elongation and flagellum positioning, and misplacement of the cell division plane. The data convincingly support the conclusion that KIN-G phosphorylation on T301 plays an important role in regulating the cellular functions of this kinesin motor protein.

      Weaknesses:

      The authors have addressed prior weaknesses in the manuscript through additional experimentation and rewording of the conclusions.

    5. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      This important study provides new insight into the regulation of cell organization and division in Trypanosoma brucei through the control of a kinesin motor protein by a polo-like kinase. The authors present solid evidence from rigorous biochemical and imaging analyses showing that phosphorylation modulates kinesin function and cellular organization. However, direct in vivo evidence that PLK phosphorylates kinesin-G is lacking.

      We performed experiments to investigate the effect of PLK inhibition on the phosphorylation of KIN-G in vivo in trypanosome cells by immunoprecipitation and mass spectrometry. The new results showed that treatment of trypanosome cells with GW843682X, a potent TbPLK inhibitor validated previously in procyclic trypanosomes, reduced the phosphorylation levels on Thr301 and Ser569 of KIN-G by ~27% and 100%, respectively. The partial reduction in Thr301 phosphorylation after GW843682X treatment could be attributed to slower dephosphorylation of phosphorylated Thr301 after GW843682X was added to the cell culture. Nonetheless, these new results demonstrated that KIN-G is an in vivo substrate of PLK.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This manuscript identifies the orphan kinesin KIN-G as a substrate of Polo-like kinase (TbPLK) in Trypanosoma brucei and demonstrates that phosphorylation of Thr301 inhibits KIN-G microtubule binding and disrupts its cellular function. Using a combination of in vitro kinase assays, phosphosite mapping, microtubule binding and gliding assays, and in vivo complementation with phosphomimetic and phosphodeficient mutants, the authors link TbPLK-mediated regulation of KIN-G to defects in centrin arm integrity, FAZ elongation, Golgi organization, flagellum positioning, and division plane placement. The study provides a mechanistic advance in understanding how TbPLK regulates centrin arm biogenesis and integrates KIN-G into the growing regulatory network controlling hook complex and FAZ assembly. Overall, the work is technically strong, internally consistent, and builds logically on previous studies from this group and others.

      Strengths:

      A major strength of the manuscript is the clear mechanistic link between phosphoryltion of Thr301 and loss of microtubule binding activity. The use of phosphomimetic (T301D) and phosphodeficient (T301A) mutants in an RNAi-rescue framework provides a clean and convincing demonstration of functional relevance in vivo. The integration of biochemical assays with detailed cell biological phenotyping (centrin arm length, FAZ elongation, basal body segregation, and cytokinesis markers) is particularly effective and makes the central conclusion robust. The observed phenotypic cascade from centrin arm defects to FAZ and division plane abnormalities is also well aligned with existing models of trypanosome morphogenesis.

      Weaknesses:

      My (more or less main) concern relates to the interpretation of the Golgi phenotype. The conclusion that phosphorylation of KIN-G "impairs Golgi biogenesis" is currently based on fluorescence microscopy using TbGRASP and Sec13 markers and on quantification of the number and distribution of Golgi/ERES puncta in binucleated cells. While these data convincingly demonstrate altered Golgi/ERES number and spatial organization, they do not distinguish between true defects in Golgi biogenesis or duplication and alternative possibilities such as fragmentation, vesiculation, or mislocalization of Golgi membranes. Given the central role of Golgi-centrin arm organization in the proposed model, ultrastructural analysis (for example, by EM or electron tomography) would greatly strengthen this aspect of the study by providing direct evidence for structural alterations of the Golgi and its association with the centrin arm and ERES. Such data would elevate this part of the manuscript from a descriptive fluorescence phenotype to a true structural cell biological insight. I appreciate that this experiment goes beyond the current dataset, but it would substantially enhance the mechanistic depth of the Golgi-related conclusions and strengthen the causal chain linking centrin arm defects to Golgi abnormalities. However, I have to confess, the inclusion of such data would make this reviewer particularly enthusiastic about the work. If this is not feasible, I would recommend tempering the wording of "Golgi biogenesis" to a more conservative description, such as altered Golgi organization or duplication, and explicitly acknowledging the limitations of fluorescence-based analysis for this conclusion.

      Thanks for these very constructive comments, which are very well taken. We totally agree with this reviewer on these points. Since it is not feasible for us to perform EM or electron tomography, we have revised the manuscript to describe the effect of KIN-G phosphorylation on the Golgi as “altered Golgi duplication” rather than “Golgi biogenesis”. We also explicitly acknowledge the limitations of fluorescence-based analysis of the Golgi for this conclusion and suggest that further characterization with EM or electron tomography would allow one to reveal the potential structural alterations of the Golgi and its association with the centrin arm.

      An additional conceptual point concerns the dual role of TbPLK in centrin arm regulation. TbPLK is known to promote centrin arm biogenesis through phosphorylation of TbCentrin2, yet in this study, TbPLK phosphorylation of KIN-G negatively regulates centrin arm assembly. This dual positive and negative regulatory role is intriguing but could be discussed more explicitly. The manuscript would benefit from a clearer conceptual framework addressing how phosphorylation of KIN-G might serve as a temporal or spatial switch to restrain KIN-G activity at specific stages of centrin arm assembly.

      This is a great point. However, we are not sure whether the previous work on TbPLK phosphorylation of TbCentrin2 could lead to the conclusion that TbPLK promotes centrin arm biogenesis through phosphorylation of TbCentrin2. In the published work (de Graffenried et al., MBoC, 2013), trypanosome cells expressing the phospho-deficient mutant TbCentrin2-S54A only showed minor growth defects, exhibiting growth defects after 5 days (de Graffenried et al., MBoC, 2013). Cells expressing the phosphomimic mutant TbCentrin2S54D, however, showed very strong growth defects (de Graffenried et al., MBoC, 2013). The effects of TbPLK phosphorylation on TbCentrin2 appear to be quite similar to that of TbPLK phosphorylation on KIN-G, although the KIN-G-T301A mutant does not have growth defects (up to 5 days in our experiments). It appears that the primary role of TbPLK in regulating TbCentrin2 and KIN-G is negative regulation. Nonetheless, we have added more discussion on these regulatory roles of TbPLK in the revised manuscript.

      Finally, a schematic model summarizing the proposed regulatory pathway from TbPLK phosphorylation of KIN-G to centrin arm assembly, FAZ elongation, division plane placement, and Golgi organization would aid the reader.

      Thanks for this suggestion. We made a schematic model to summarize the roles of KIN-G and its regulation by TbPLK. This is included in Figure 8.

      Reviewer #2 (Public review):

      Summary:

      The authors identify KIN-G as an in vitro substrate for phosphorylation by TbPLK and show that several of the in vitro P-ated sites, including T310, overlap with P-ation sites seen in live cells. The authors further show that PLK-mediated P-ation inhibits KIN-G binding to microtubules in vitro, as does a KIN-G-T301D mutant, and that expression of a KIN-G-T301D Phospho-mimic in T. brucei phenocopies KIN-G RNAi knockdowns, producing defects in cell division, morphogenesis of the centrin arm, FAZ and other cellular structures, as well as a misplaced cytokinesis furrow.

      Understanding cytoskeletal rearrangements that drive cell division in T. brucei is an important and unresolved problem, so the work addresses important questions that are of great interest. PLK and KIN-G have previously been shown to be important for cell division and morphogenesis of cytoskeletal structures that drive cell division in T. brucei. The current work advances our understanding by suggesting a potential mechanism by which PLK and KIN-G might participate, namely through PLK-dependent P-ation to control KIN-G MT binding activity.

      Strengths:

      The authors use a rigorous combination of biochemistry, phosphoproteomics, cell biology, and mutant analysis to support their conclusion that PLK-mediated P-ation of KIN-G negatively regulates KIN-G microtubule binding, and this may explain the observation that a KIN-G T301 phosphomimic mutant blocks cell division and perturbs biogenesis of cytoskeletal structures that drive cell division and morphogenesis. Combining rigorous and informative in vitro studies with mutant analysis in live cells is a great strength. The work is solid and important, though a few pieces are needed to fully connect the in vitro findings with the in vivo observations, as detailed below.

      Weaknesses:

      Overall, I find this work to be solid and to provide an important addition to our understanding of mechanisms controlling cell division in T. brucei. The biochemistry, in particular, is rigorous and convincingly demonstrates PLK can P-ate KIN-G, altering its MT-binding ability. Analysis of phospho-mutants of KIN-G in live T. brucei supports the conclusion that P-ation of KIN-G at T301 negatively affects KIN-G function in vivo. I think, however, that the results fall short of supporting the title, because, although the data convincingly show that PLK can phosphorylate KIN-G at T301 in vitro, and that T301 is P-ated in vivo, they do not formally demonstrate (nor even test) whether PLK is the kinase responsible for this phosphorylation in vivo (experiments to address this seem quite feasible). I also do not see where the authors try to reconcile the absence of phenotype for KIN-G-T301A with the implied importance of KIN-G phosphorylation by PLK in cell division, which calls into question the need for P-ation of KIN-G-T301 in cell division. Suggestions for addressing these concerns are provided below.

      My two main questions are:

      (1) What is the biological relevance of KIN-G P-ation at T301?

      (a) The authors report no defect for the KIN-G-T301A mutant, so what then is the need for T301 P-ation, if the cell gets along fine without it? One step toward addressing this would be to ask what fraction of KIN-G shows P-ation at T301. Although published studies indicate P-ation at T301, it isn't known what percentage of KIN-G in the cell is P-ated. One might anticipate, for example, that T301-P is a small minority of the population in asynchronous cultures and that T301 P-ation increases at specific cell cycle stages.

      This is a great point that is very well taken. We also had been puzzled by the observation of no growth defects of T301A mutant. This comment enlightened us. From the new experiments we performed to compare the phosphorylated peptides of KIN-G in cells treated and non-treated with the PLK inhibitor GW843682X, we calculated the percentage of phosphorylated Thr301 in non-GW843682X-treated cells. We found that the percentage of peptides containing the phosphorylated Thr301 is ~14% of the total Thr301-containing peptides (Fig. 1H). This result indicates that T301-P is indeed a small minority of the population in the asynchronous trypanosome cells. It is possible that T301 phosphorylation may occur at a specific cell cycle stage such as early S-phase, during which PLK and KIN-G co-localize at the centrin arm.

      (b) Published work links PLK to cell division, FAZ elongation, etc.. The current work suggests that one role of PLK is to P-ate KIN-G at T301. In contrast, however, the current work also indicates that P-ation of KIN-G at T301 is unnecessary for normal cell division, FAZ elongation, etc..

      Yes, previous work discovered essential roles of TbPLK in basal body segregation, centrin arm biogenesis, FAZ elongation, and cytokinesis. These functions of TbPLK correlate with TbPLK’s localization to multiple subcellular structures, the basal body, the centrin arm, and the new FAZ tip, and are attributed to the regulation of its substrates at these structures. At the basal body, TbPLK phosphorylates SPBB1, which is required for basal body segregation. Defects in basal body segregation can lead to defective flagellum positioning and FAZ elongation. At the new FAZ tip, TbPLK regulates the cytokinesis regulator CIF1, which is required for cytokinesis. At the centrin arm, TbPLK phosphorylates TbCentrin2 at S54 and KIN-G at T301 (and S569, which was newly identified as an in vivo TbPLK site and has not yet been characterized). However, cells expressing TbCentrin2-S54A have very weak growth defects, and cells expressing KIN-G-T301A have no detectable growth defects. In contrast, cells expressing TbCentrin2-S54D and cells expressing KIN-G-T301D have strong growth defects. Therefore, the essential role of TbPLK in centrin arm biogenesis apparently is not attributed to the phosphorylation of TbCentrin2 and KIN-G. It is possible that phosphorylation of other centrin arm-localized protein(s) by TbPLK may be essential for centrin arm biogenesis, but this possibility remains to be explored.

      (c) Some experiments or at least commentary on points a and b above would strengthen the paper.

      We performed experiments and presented the data in Fig. 1H. We also included commentary in the revised manuscript on the points about the potential role of T301 phosphorylation. Thanks for these great comments that significantly improved the manuscript.

      (2) Is PLK the kinase that P-ates Kin-G T301 in vivo?

      (a) The authors show PLK P-ates T301 (and other residues) in vitro, and that T-301 is P-ated in vivo. To bring the analysis full circle, it would be informative to examine KIN-G P-ation in a PLK mutant or upon inhibition of PLK with published inhibitors. This seems to be a very doable experiment with the tools available.

      We treated trypanosome cells with a potent PLK inhibitor GW843682X, which was previously demonstrated to inhibit TbPLK activity in vitro and mimic TbPLK knockdown in vivo in trypanosomes, and immunoprecipitated KIN-G for mass spectrometry. We compared the KIN-G peptides identified by mass spectrometry from trypanosome cells treated with or without GW843682X, and found that two phosphosites (T301 and S569) were reduced by ~27% and 100%, respectively, after GW843682X treatment (Fig. 1G). These results provided evidence to support that TbPLK phosphorylates KIN-G in vivo.

      Reviewer #3 (Public review):

      Summary:

      Here, the authors investigate the role of the Trypanosoma brucei polo-like kinase TbPLK in the function of flagellum-associated cellular structures in trypanosomes. They set out to test the hypothesis that a key substrate of TbPLK is the kinesin protein KIN-G, and that TbPLK phosphorylation of KIN-G regulates its functions in cells.

      Strengths:

      Using in vitro biochemistry with purified proteins, the authors convincingly demonstrate that TbPLK phosphorylates KIN-G at 29 sites. Moreover, they convincingly show that phosphorylation at one site, T301, impairs the binding of purified KIN-G to purified microtubules. Using immunofluorescence-based imaging approaches, they also show that TbPLK colocalizes with KIN-G at centrin arms during the early S-phase of the cell cycle. Centrin arms are structures that are located near the basal body and flagellum and are important for new flagellum biogenesis, Golgi positioning, and cell division. To evaluate the function of KIN-G phosphorylation in cells, they depleted KIN-G by RNAi, simultaneously expressed phospho-mimetic (T301D) and phospho-ablative mutant proteins, and used immunofluorescence to examine the impact on flagellum-associated cellular structures. They show that expression of the phospho-mimetic mutant KIN-G-T301D causes the following defects: reduced cell proliferation, disruption of centrin arm and Golgi biogenesis, impairment of FAZ elongation and flagellum positioning, and misplacement of the cell division plane. The data convincingly support the conclusion that KIN-G phosphorylation on T301 plays an important role in regulating the cellular functions of this kinesin motor protein.

      Weaknesses:

      Some of the broader conclusions are not directly supported by the data. For example, the title states "Polo-like kinase phosphorylation of the orphan kinesin KIN-G negatively regulates centrin arm biogenesis in Trypanosoma brucei," but the data do not directly address the specific role of TbPLK in phosphorylating KIN-G in cells. Moreover, some of the more specific conclusions in the paper, for example, that "phosphorylation of KIN-G" causes various cellular defects, are a bit of an overstatement. The supporting data rely on the expression of a phospho-mimetic mutant of KIN-G. Presumably, phosphorylation in cells is a normal part of KIN-G regulation, and it is not just phosphorylation, but rather hyperphosphorylation that is being mimicked by the mutant. Some rewording of the specific conclusions is warranted, and the broader conclusion would be better supported with additional experimental evidence.

      This is a great point that is very well taken. We performed new experiments to address the in vivo phosphorylation of KIN-G by TbPLK. We treated cells with a potent PLK inhibitor, GW843682X, and then immunoprecipitated KIN-G for mass spectrometry to identify changes in phosphorylation. We found that the phosphorylation levels of T301 and S569 were reduced by ~27% and 100%, respectively, confirming that these two sites are in vivo TbPLK phosphosites.

      We also calculated the ratio of phospho-T301 versus non-phospho-T301 in non-treated cells and found that phospho-T301 accounts for ~14% of the total KIN-G protein. This new result suggests that it is the hyperphosphorylation that causes growth defects. We have revised the manuscript accordingly to reflect this point.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Several statements use rather strong causal language (for example, "thereby impairing Golgi biogenesis, FAZ elongation, and division plane placement"). While the phenotypic correlations are convincing, direct causality is largely inferred from prior literature. Slightly tempering this wording could improve precision. It would also be helpful to state explicitly in figure legends the number of cells analyzed per condition and the number of independent experiments for each quantified phenotype.

      Thanks for these constructive comments, which we agree and appreciate greatly. We have revised the manuscript accordingly to improve precision. The total number of cells analyzed, and the number of independent experiments were included in the figure legends.

      Reviewer #2 (Recommendations for the authors):

      Minor comments for improving the text are:

      (1) The paper overall is clearly written. However, the Discussion starts with a solid sentence, then becomes a bit diffuse in discussing a wide range of PLK activities that were not addressed in the current work. That detracts attention a bit from the central contributions of this paper.

      Thanks for this comment. We have deleted the discussion about TbPLK activities that were published previously.

      (2) At least two places in the text state apparent contradictions.

      (a) p.5 and Figure 2C. The authors say microtubule gliding speed was "...insignificantly reduced..." by the TbPLK-K70R mutant, yet they then state that motility was "interfered with". If the effect is "insignificant", why do they claim there is an effect?

      (b) p6 and Figure 3C. The authors report KIN-G-T301A impact on microtubule gliding activity is insignificant, but then say this mutation reduces the motility of KIN-G. These statements are contradictory.

      We meant to say that there was a slight but insignificant effect. We agree that such statements are somewhat contradictory and, hence, have been deleted. Thanks.

      (3) p. 8, and Figure 7. "ventral side" and "leading edge" are not defined but are used to describe the KIN-G RNAi phenotype.

      We have deleted the wording “ventral side”, as it is not necessary. Thanks.

      (4) Figure 7B. Please explain the labeling - the new flagellum daughter is indicated as having the old posterior, while the old flagellum daughter cell is indicated as having the new cell posterior. This is counterintuitive to a reader not intimately familiar with the T. brucei cell division process.

      Thanks very much. We included two sentences in the revised manuscript to explain this point.

      The sentences read as follows: “The nascent posterior is formed near the mid-portion of the NFD cell through microtubule bundling and cytoskeleton remodeling during late stages of the cell cycle (Wheeler et al., 2013). Consequently, the NFD cell inherits the old, existing cell posterior, whereas the OFD cell inherits the newly formed or nascent cell posterior.”

      (5) Figure 4, 5, and 7: "% Cells" is reported. Please indicate the total number of cells that were examined.

      The total number of cells were included in the figure legends.

      Reviewer #3 (Recommendations for the authors):

      (1) The manuscript should be carefully edited for minor grammatical errors.

      Thanks. We have carefully proofread the manuscript and corrected the grammatical errors.

      (2) A general conclusion is that TbPLK phosphorylation of KIN-G in cells is critical for regulating its motor activity. However, this relies on the expression of phospho-mimetic mutants, which bypass TbPLK. Thus, there really is no direct evidence provided to support the specific role of TbPLK other than the in vitro phosphorylation data. Some additional experiments to assess the specific role of TbPLK in phosphorylating KIN-G in cells would lend support for the general conclusion. Is it possible to deplete or inhibit TbPLK and show that this impacts the phosphorylation of KIN-G in cells?

      This is a great point that is very well taken. We performed a new experiment by inhibiting TbPLK with a potent PLK inhibitor, GW843682X, and then immunoprecipitating KIN-G for mass spectrometry to identify the phosphorylation levels before and after GW843682X treatment. We were able to confirm that T301 and S569 phosphorylation was reduced after treatment. This confirms that TbPLK phosphorylates T301 and S569 of KIN-G in vivo in trypanosome cells.

      Specific:

      (1) Figure 2: For microtubule gliding assays, representative videos should be included as supplementary data. Also, when the data do not show a significant difference between KIN-G and KIN-G + TbPLD-K70R, then the authors should not state there is a slight difference, as this is not supported by the data.

      We have deleted the statement. We have included the representative videos for all the experiments presented in Figures 2 and 3. Thanks.

      (2) Figure 3: As stated above, for microtubule gliding assays, representative videos should be included. Moreover, for KIN-G-T301A, the authors say that the gliding activity was "moderately, but insignificantly, reduced." Again, if the difference is not significant, it cannot be concluded that there is a difference compared with the wild-type protein.

      We have deleted the statement. Thanks.

      (3) Some of the headings in the Results section are not accurate. Regarding the in vivo results, the heading "Phosphorylation of Thr301 in KIN-G by TbPLK causes defective cell proliferation" seems to be an overstatement. It may be more accurate to state that "Expression of a phospho-mimetic mutant of KIN-G causes defective cell proliferation." The same comment applies to the other headings that follow this one. Presumably, there is a population of phosphorylated KIN-G in cells, and phosphorylation/dephosphorylation is a normal part of its regulation. In the text, the authors might more accurately conclude that hyperphosphorylation causes the defects they are seeing.

      This is a great point. We agree and as we responded above, we have revised the manuscript, from the title to the main text, to reflect the point that hyperphosphorylation of T301 by TbPLK causes the defects. Thanks very much for these great comments!

    1. eLife Assessment

      This important study demonstrates that nutrient resorption efficiency in the widespread wetland grass Phragmites australis is largely canalized by phylogeographic lineage, ecotype and geographic origin, rather than responding plastically to short-term salt stress. The findings have implications beyond plant ecophysiology because they suggest that predictions of wetland nutrient cycling under increasing salinization should account for the genetic composition of plant populations. The evidence supporting the central conclusions is compelling, based on a well-designed common-garden experiment involving 110 genotypes, paired control and salinity treatments, and convergent metabolomic, ionomic and whole-plant evidence confirming that the treatment imposed substantial physiological stress. The interpretation is nevertheless limited to one growing season and a single, moderate salinity treatment, leaving open whether chronic, more severe or multigenerational exposure might induce plastic or transgenerational responses; further clarification of the relationships among phylogeographic lineage, ecotype and latitude, and a more cautious interpretation of the variance explained by latitude, would strengthen the manuscript.

    2. Reviewer #1 (Public review):

      Summary:

      This study demonstrates that nutrient resorption efficiency (NuRE) in Phragmites australis is genetically canalized rather than plastic to salt stress. Using 110 genotypes in a common garden, the authors show that intraspecific variation in NuRE is explained by phylogeographic lineage, ecotype, and latitude, not by effective salinity. Element-specific regulatory strategies further reveal how N, P, and K resorption are differentially controlled. At the population level, this is an important study that fundamentally advances our understanding of plant functional trait evolution and its implications for ecosystem nutrient dynamics under global change.

      Strengths:

      This study is the first to demonstrate genetic determination of a key nutrient conservation trait under effective salt stress in a widespread macrophyte, directly testing the 'plastic acclimation versus inherent conservatism' paradigm in a non-nutrient stress context. The experimental design is rigorous: each genotype was paired across control and salt treatments, and multilevel stress effectiveness (metabolomics, biomass, Na accumulation) was confirmed before evaluating NuRE. The large sample size of a macrophyte and dual classification (phylogeography + ecotype) allow robust disentangling of genetic versus plastic sources of variation.

      The analysis comprehensively tests three resorption control hypotheses using appropriate SMA regression, revealing element-specific and condition-dependent patterns. The latitudinal gradient and variation partitioning provide strong evidence that genetic origin and geographic context outweigh short-term plasticity, with important implications for predicting ecosystem nutrient cycling under global change. This study provides a clear empirical demonstration that a key nutrient conservation trait can remain homeostatic under non-nutrient stress, and that intraspecific variation is primarily a product of population differentiation rather than short-term plasticity.

      Weaknesses:

      First, the salinity treatment spanned only one growing season. The conclusion of genetic canalization therefore specifically refers to the absence of plasticity to an acute salt shock. Whether long-term, multigenerational chronic salinity could act as a selective agent or induce transgenerational plasticity remains an open and interesting question for further research. Likewise, the physiological mechanisms underlying the observed lack of plastic increase in NuRE (for example, phloem loading or senescence gene expression) are not directly resolved, leaving some inference about trade-offs versus true unresponsiveness. These points do not weaken the study's main conclusion. Instead, they suggest productive future directions, such as longer-term field manipulations and targeted molecular investigations.

      Second, the test of nutrient limitation control relies on resorbed N:P and N:K ratios as proxies, an established but indirect approach. Direct nutrient addition experiments would provide stronger causal evidence. Also, the metabolomic analysis is used primarily to validate stress effectiveness; deeper integration of specific metabolites with NuRE variation across genotypes could have offered mechanistic insights but was not pursued. Additionally, the potential collinearity between ecotype and phylogeographic lineage among Chinese populations is not quantitatively addressed. None of these considerations undermines the main finding, which is supported by a robust experimental design and widely accepted analytical approaches.

    3. Reviewer #2 (Public review):

      Summary:

      The study finds that nutrient resorption efficiency in Phragmites australis shows no plastic response to salinity stress but is canalized by phylogeographic lineage, ecotype, and latitude. In a common garden with 110 genotypes, salinity induced stress, yet no plastic change occurred for N, P, or K resorption. The authors conclude that intraspecific variation is historical and geographic; thus, predictions of wetland nutrient cycling need to account for phylogeographic composition.

      Strengths:

      The core finding that NuRE shows no plastic response to salinity, but is instead evolutionarily canalized by lineage and latitude, challenges a key assumption of broad trait plasticity. This conclusion is firmly supported by a robust common garden design with 110 genotypes, rigorous multi-level stress validation, and element-specific resorption analyses. The work provides compelling evidence that intraspecific variation in this critical nutrient cycling trait is shaped by phylogeographic history rather than short-term acclimation. The implications for predicting wetland responses to salinization are significant, as ecosystem-level nutrient dynamics may be constrained by the genetic composition of plant populations.

      Weaknesses:

      The experiment covers only one growing season, with salinity applied in June and measurements in December. While the stress is clearly effective, longer-term or multi-year stress might reveal acclimation or epigenetic effects that are not captured. Given the author team's expertise in parental and transgenerational effects in clonal plants, this limitation is particularly relevant and warrants more thorough discussion in the manuscript.

      The salinity treatment uses a single moderate level of 10 ppt, which does not allow assessment of whether more extreme stress might trigger a plastic response. A dose-response design across a gradient would have provided stronger inference about the threshold at which NuRE canalization might be overcome. Additionally, the ecotype analysis in Figure 4 applies only to Chinese populations, as classification was not available for non-Chinese populations, which should be stated more explicitly in the Results.

      The variation partitioning shows latitude as a significant predictor, but the R² values are relatively low, indicating that much variance remains unexplained. The manuscript should avoid overinterpreting latitude's explanatory power and more openly acknowledge the role of unmeasured factors. The interpretation of slopes greater than 1 for the resorbed N:P versus green N:P relationship, labeled as "inverted limitation", also needs further explanation regarding its functional significance.

    4. Author response:

      We sincerely thank the editors and reviewers for the positive assessment of our work and for the constructive and insightful feedback. We are grateful that the experimental design and the evidence for genetically canalized nutrient resorption efficiency (NuRE) were recognized as compelling and important, with clear implications for predicting wetland nutrient cycling under salinization. We fully agree with the major points raised in the public reviews and outline below our planned revisions to address them, with particular attention to the weaknesses noted.

      Reviewer #1 raised two important concerns regarding the scope and generality of our conclusions. First, the salinity treatment spanned only one growing season, so the genetic canalization we document specifically refers to the absence of a plastic response to an acute salt shock; whether long-term, chronic or multigenerational salinity could act as a selective agent or induce transgenerational plasticity remains an open and interesting question. Second, the test of nutrient limitation control relies on resorbed N:P and N: K ratios as proxies rather than direct nutrient manipulation, and the metabolomic analysis is used primarily to validate stress effectiveness. We accept these criticisms and will address them as follows.

      Regarding the temporal scope of the treatment, we will explicitly state in the Discussion that our conclusion of canalization pertains to short-term acclimation to an acute salt shock, and we will discuss the scenarios under which chronic, more severe or multigenerational exposure could trigger plastic, acclimatory or transgenerational responses. Given our team’s prior work on parental and transgenerational effects in clonal plants, we will frame these as testable hypotheses for future research. We will also acknowledge that the physiological mechanisms underlying the lack of a plastic increase in NuRE (for example, phloem loading or senescence-associated gene expression) are not directly resolved in the present study and will propose targeted molecular investigations as a natural next step.

      Regarding the nutrient limitation tests, we will clearly acknowledge in the Discussion that the resorbed N:P and N:K ratios provide an established but indirect proxy for nutrient limitation, and we will discuss how direct nutrient addition experiments could provide stronger causal evidence, while deepening the integration of the metabolomic profiles with genotype-level NuRE variation where feasible. We will also quantitatively assess the potential collinearity between ecotype and phylogeographic lineage among the Chinese populations.

      Reviewer #2 raised three substantive issues regarding the interpretation and presentation of our results. First, although latitude emerged as a significant predictor in the variation partitioning, the relatively low R² values indicate that much variance remains unexplained, and our interpretation should be more cautious; in addition, the functional significance of the “inverted” nutrient limitation (slopes greater than 1 for resorbed versus green N:P) needs further explanation. Second, the single moderate salinity level (10 ppt) does not allow assessment of whether more extreme stress might trigger a plastic response. Third, the ecotype analysis applies only to the Chinese populations, as ecotype classification was not available for non-Chinese populations, and this should be stated explicitly in the Results. We fully agree with these points and will revise accordingly.

      To address the concern about latitude, we will temper the language in the Results and Discussion, explicitly noting the limited proportion of variance explained by latitude and acknowledging the role of unmeasured factors, while retaining the study’s central message that genetic and phylogeographic origin outweigh short-term plasticity in shaping NuRE. We will also expand the Discussion to explain the functional significance of the “inverted” nutrient limitation, that is, why P and K are resorbed more completely relative to N, whether as a strategy to maintain optimal N:P:K ratios or a reflection of the higher costs and lower availability of N.

      To address the salinity gradient concern, we will acknowledge in the Discussion that the single moderate salinity level may not have been severe enough to trigger a plastic response and will justify future dose-response experiments to identify the threshold at which canalization of NuRE might be overcome.

      To address the ecotype limitation, we will explicitly state in the Results that the ecotype analysis in Figure 4 is based only on the Chinese populations, for which ecotype classification was available.

    1. eLife Assessment

      This important study introduces an open-source software package that improves the computational efficiency of whole-brain modelling and facilitates individualised model fitting in large neuroimaging cohorts. The evidence is convincing for the main computational and practical claims, supported by evaluations across multiple optimisation strategies, speed benchmarks, and demonstrations using human neuroimaging data. Its modular design and documentation should make it a resource that is of value to researchers interested in scalable and individualised brain network modelling.

    2. Reviewer #1 (Public review):

      Summary:

      The manuscript introduces cuBNM, a GPU‑accelerated Python package for whole‑brain modeling. The authors demonstrate that running simulations on GPUs provides substantial benefits in computational speed, cost-efficiency, and scalability compared to traditionally used CPUs, making large‑scale and individualized brain network modeling computationally feasible. The usage of cuBNM has been demonstrated by running optimization of group-level and individualized low- and high-dimensional models. By investigating the test-retest reliability and heritability of simulated and empirical measures in the Human Connectome Project dataset, the authors showed that simulated features were fairly reliable and significantly heritable.

      Strengths:

      This study is timely and presents an important contribution to the field of whole-brain computational modeling. A major strength is that the authors go beyond introducing a GPU-accelerated framework by demonstrating its utility through comprehensive benchmarking and biologically relevant applications, including individualized model fitting, comparisons of homogeneous and heterogeneous models, and analyses of test-retest reliability and heritability.

      The computational performance is evaluated comprehensively, assessing speed, computational cost, energy consumption, and scalability across different simulation settings. The Human Connectome Project dataset is used to demonstrate that the software enables individualized whole-brain modeling in large datasets.

      Finally, the software is modular, open-source, and well-documented, and can facilitate the broader adoption of GPU-accelerated whole-brain modeling within the neuroscience community.

      Weaknesses:

      The test-retest reliability and heritability are estimated using high-quality Human Connectome Project data. The manuscript would benefit from discussion and/or demonstrations regarding how the software performs under more challenging conditions, such as clinical datasets, shorter data acquisitions, higher-motion datasets, or multi-site datasets.

      Apart from demonstrating the benefits of GPUs over CPUs, the manuscript would benefit from a more direct comparison between cuBNM and other whole-brain modeling software, such as The Virtual Brain.

      The manuscript demonstrates that heterogeneous models improve the fit to empirical functional connectivity. However, the biological interpretation of this improvement could be expanded. The heterogeneous models are also more complex than homogeneous models, and some improvement in model fit may be explained by the increased model flexibility.

      In whole-brain brain network modeling, different parameter combinations can result in similar empirical functional connectivity measures. The manuscript would benefit from a discussion of how this influences the interpretation of individualized parameter estimates.

    3. Reviewer #2 (Public review):

      Summary:

      The authors aim to address a major problem in brain network modeling: the high computational cost of simulating and fitting brain activity models, particularly for large samples, individualized models, and broad parameter searches. They introduce cuBNM, an open software package that uses graphics processing units to accelerate model simulation, fitting, and calculation of simulated brain activity features.

      The manuscript is primarily a methods and software contribution, rather than a paper providing novel neurobiological insights. The authors demonstrate the tool using human imaging data, showing examples of group-level and individualized model fitting, comparisons between homogeneous and heterogeneous model parameterizations, and analyses of repeated-measurement stability and genetic influences of simulated features. They also provide speed and scaling tests to support the claim that the software can make large-scale and individualized brain network modeling more practical for the field.

      Strengths:

      A major strength of this work is that it addresses a clear computational bottleneck in brain network modeling. The authors provide an open software package that combines a user-friendly Python interface with an accelerated back-end, making large numbers of simulations and model fits more practical for other researchers.

      The demonstrations are broad and relevant to real use cases. The authors show group-level and individualized model fitting, different optimization strategies, and comparisons between homogeneous and heterogeneous models, rather than limiting the paper to a narrow technical benchmark. The benchmarking and openness of the work further increase its value. The comparisons across hardware and network sizes give readers a practical sense of the tool's speed and scalability, while the availability of code, documentation, tutorials, and containers should make the method easier for the community to test and adopt.

      Weaknesses:

      (1) The benchmarking provides solid evidence for substantial speed improvements within the authors' implementation, but the generality of the performance claims is more limited. The largest reported speed-ups are measured relative to a single central processing unit thread, and the study does not fully benchmark cuBNM against other optimized brain modeling frameworks. This makes the results useful as evidence of strong acceleration in the tested setting, but less definitive as a general comparison across available implementations.

      (2) The comparison between homogeneous and heterogeneous models is informative, but it is not fully controlled for model complexity. The best-fitting node-based heterogeneous model has more free parameters than the homogeneous and map-based alternatives, so its improved fit may partly reflect greater flexibility rather than a more biologically valid parameterization. As a result, the model comparison supports the conclusion that this parameterization fits better under the current setup, but not necessarily that it is generally superior or more biologically realistic.

      (3) The reliability and heritability analyses are valuable demonstrations of what scalable individualized modeling can enable, but they do not establish the simulated features as validated biological mechanisms. Because these simulated features are derived from individualized structural and functional imaging data, their stability across repeated measurements and genetic influences may partly reflect information already present in the empirical inputs or fitting targets. These results therefore support a more cautious conclusion: the simulated features retain stable and genetically structured variation, but their biological interpretation remains model dependent.

      (4) The empirical demonstrations are narrower than some of the broader claims made in the manuscript. Most analyses rely on one human imaging dataset, one cortical parcelation, one main brain model, and a specific fitting objective, while broader claims refer to diverse populations, dense networks, high-dimensional models, and biological applications. The current results show that cuBNM is a useful and scalable tool in the tested setting, but the extent to which the findings generalize across datasets, model classes, network resolutions, or clinical contexts remains to be established.

    1. eLife Assessment

      This is a valuable study addressing a debated question in brain repair research by testing whether NeuroD1 can convert brain immune cells into nerve cells using a virus-free genetic approach and live imaging. The solid evidence presented here supports that, under the conditions tested, NeuroD1-expressing cells do not become nerve cells and instead retain their original identity; however, some aspects of expression level, injury timing, and cell-type composition would benefit from further clarification. This work will be of interest to neuro-immunologists.

    2. Reviewer #1 (Public review):

      Summary:

      This study revisits an important and controversial question in brain repair: whether NeuroD1 can convert brain immune cells into nerve cells in vivo. Using a virus-free genetic system, in vivo imaging, injury experiments, and single-cell profiling, the authors provide convincing evidence that NeuroD1-expressing cells do not become nerve cells under the tested conditions. Instead, these cells largely retain their original immune-cell identity, and some appear to undergo cellular stress or loss.

      Strengths:

      The main strength of the work is that it tests this question with a cleaner genetic strategy, avoiding some of the concerns associated with viral delivery and unintended cell labeling. Although the overall conclusion is consistent with the authors' previous work, the current study adds useful independent evidence, particularly through the virus-free fate-mapping system and live imaging in the brain.

      Weaknesses:

      There are some limitations. In the injury experiment, the labeled cells may include both resident brain immune cells and blood-derived immune cells recruited after injury, so the authors should be cautious when referring to all labeled cells as microglia. The level of NeuroD1 expression achieved by the genetic system is also not fully defined, which matters because the effects of such a cell-fate regulator may depend on expression level. Finally, the tested time window may not fully address very delayed or incomplete neuronal differentiation.

      Overall, this is a useful and careful study that supports the conclusion that NeuroD1 does not drive brain immune cells to become nerve cells in the tested settings. It should be valuable for researchers studying brain repair, cell fate conversion, and genetic fate mapping, and it provides a clear caution against overinterpreting reprogramming results based only on viral labeling.

    3. Reviewer #2 (Public review):

      Summary:

      In vivo glia-to-neuron conversion emerges as a potential regeneration-based therapeutic strategy for neural injuries and diseases. However, controversies exist in this exciting field, largely arising from the non-stringent methods employed for analyzing in vivo neuronal conversions. The study by Li et al. directly addressed this controversy regarding Neurod1-mediated microglia-to-neuron conversion. They took advantage of two transgenic mouse lines to specifically express Neurod1 in the microglia of adult mouse brains. Results from immunohistochemistry, in vivo live-cell imaging, and scRNA-seq convincingly demonstrate that microglia cannot be converted in vivo to neurons by ectopic Neurod1 expression under both normal and injury conditions. Instead, it induces microglia death, consistent with their earlier findings. These solid results, though negative, are critical additions to the field and further support that stringent lineage tracing methods are essential for studying in vivo cell reprogramming. Overall, the studies are rigorously designed and executed. Only minor issues need to be dealt with.

    1. eLife Assessment

      This study uses a technically challenging long-term in vivo recordings and computational modeling to investigate whether hawkmoth olfactory receptor neurons show circadian modulation of spontaneous firing. The study provides solid evidence that ORNs have a circadian pattern of spontaneous firing that is Orco-dependent and yet that Orco transcript abundance does express a circadian rhythmicity and that cAMP can modulate all Orco-dependent activity. Together with computational work, the work is valuable in proposing a provocative hypothesis that an Orco-centred post-translational feedback loop generates the circadian rhythm.

    2. Joint Public Review:

      The revised manuscript is much clearer, and the additional analyses address several of the original concerns. RAIN analyses (Rhythmicity Analysis Incorporating Nonparametric methods) now detects circadian rhythmicity in 7/11 recordings under light-dark conditions and 8/12 recordings in constant darkness, compared with 2/12 following treatment with the Orco antagonist. This supports circadian modulation of spontaneous firing and a role for Orco in its normal expression. The expanded qPCR analysis of Orco also supports the conclusion that Orco transcript abundance is not circadian, and the cAMP experiment shows that cAMP can modulate Orco-dependent activity.

      The remaining issue concerns the mechanistic interpretation. The lack of rhythmic Orco transcript abundance does not distinguish an autonomous post-translational feedback-loop (PTFL) clock from a model in which the canonical transcriptional-translational (TTFL) clock acts upstream through cAMP, calcium, kinases, phosphatases, channel trafficking, or related pathways to regulate Orco.

      Similarly, the new Figure 10 provides a useful representation of the authors' hypothesis, but the proposed delayed feedback and coupling mechanisms are not experimentally demonstrated.

      We do not think any further experiments are necessary for the present study. Instead, we recommend that the manuscript should clearly distinguish between what the data show and what remains proposed. The data support circadian modulation of ORN firing a role for Orco in its normal expression, non-circadian Orco transcript abundance, and cAMP-sensitive modulation of Orco-dependent activity. The proposal that an Orco-centred membrane feedback loop generates the rhythm is intriguing and may remain a hypothesis generated through this study that needs formal testing in the future. This should be explicitly stated. While this has been done in the discussion section, elsewhere, including in the abstract and elsewhere, the original claim remains.

    3. Author response:

      The following is the authors’ response to the previous reviews

      eLife Assessment

      This valuable study uses technically compelling long-term in vivo recordings and computational modeling to investigate whether hawkmoth olfactory receptor neurons show circadian modulation of spontaneous firing. The authors further propose the provocative model that post-translational mechanisms, rather than the transcriptional-translational processes, may contribute to circadian regulation of neuronal excitability.

      We are pleased that our study is recognized as valuable and that our technically very challenging long-term in vivo recordings and computational modeling are appreciated. We agree that we are proposing a provocative model that opposes the current hypothesis in chronobiology, which suggests that all observed biological circadian rhythms are outputs of a transcriptional-translational feedback loop (TTFL) clock. Instead, we suggest that a cell comprises, in addition to the TTFL clock, other posttranslational feedback loop (PTFL) clocks without the need of daily transcription and daily degradation of its core elements. While the circadian TTFL clock is entrained to the daily light-dark cycle, the circadian PTFL clocks are suggested to be entrained to other daily cues such as to the availability of pheromone, or the daily changing levels of hormones and second messenger levels. Our novel hypothesis proposes that the TTFL and PTFL clocks are coupled and linked, constituting an adaptive, plastic network that can tune and phase-lock to different external and internal Zeitgeber signals. However, we certainly do not claim that the TTFL circadian clock is not at all involved in the circadian control of the ORN’s circadian membrane potential rhythms. We clarified our manuscript accordingly.

      However, the evidence for circadian firing in these neurons […] remains incomplete.

      As requested by the reviewers during the previous round of review, we had provided the results of RAIN analysis (Thaben and Westermark, 2014) of individual animals in our first revision (Fig. 4A), which clearly shows that two-thirds of the population express circadian rhythms in key attributes that we used to characterize the spontaneous spiking activity. In the initial review it was assumed by the reviewers that phase alignment of dispersed rhythms would bias interpretations of rhythmicity. After having established with RAIN that individuals show circadian rhythms, albeit dispersed across the population due to the lack of a zeitgeber in DD conditions, phase-alignment of recordings from DD animals is a valid next step to prepare the data for statistical analysis across the population. Phase-alignment of desynchronized rhythms is a generally accepted and necessary method employed in chronobiology (e.g., for insect ORNs: Gosh et al., 2024). It is proven as prerequisite to find and statistically analyze rhythmicity in complex, desynchronized data.

      As we explained in the previous rebuttal and in our first revision, rhythms in electrical activity of insect ORNs cannot be easily synchronized by the light-dark cycle alone, but appear to require daily cycles of pheromone, as shown in other moth species (Gosh et al., 2024). Please be aware that the animals that we used here have never been exposed to pheromone, as we state in the Methods. Therefore, this lack of pheromone exposure can explain why about one third of our experimental population is not expressing any daily or circadian rhythmicity in spiking attributes. This is an important result of our manuscript, reported for the first time for Manduca sexta, providing evidence for our hypothesis that it is not the LD-entrained circadian TTFL clock that governs electrical activity rhythms in ORNs. As we explained in the first revision, and clarified further here in the second revision, we cannot phase-synchronize our animals with cycles of pheromone application in our experimental paradigm because we are researching circadian rhythms in spontaneous spiking activity and not pheromone responses. We failed to obtain phase-alignment with a single pheromone pulse the night before the experiments started. These data were added as supplementary Figure to Fig. 3 in the first revision. Here, we further revised our manuscript to clarify this important finding.

      Thus, as requested by the reviewers in the initial review, we could successfully confirm our previous results of circadian firing in ORNs and the disruption of these circadian rhythms with Orco antagonist OLC15 with RAIN. In the current review, the reviewers raise no further specific critical points or comments that would doubt our careful rhythm analysis of our long-term recordings. Thus, we conclude that we provided clear evidence for our central, exciting new finding. For the first time we demonstrated an unexpected new task for Orco: Orco controls the circadian firing pattern in the spontaneous activity, and thus, of the ORN’s membrane potential, via its property as leak/pacemaker channel.

      However, the evidence […] for post-translational modification of Orco as the underlying mechanism remains incomplete.

      We agree with the reviewers that there are many more experiments and combined efforts of biochemists, structural biologists, and electrophysiologists required to provide complete evidence for post-translational modification of Orco and to reveal the underlying mechanism of its circadian control. It is beyond the scope of the current manuscript to provide all details of post-translational control of Orco.

      The reviewers asked previously for additional evidence that Orco transcription is not controlled via the TTFL clock. As requested, we provided extended qPCR–based evidence that Orco, in contrast to timeless, is not controlled by the TTFL clock on the transcriptional level (Fig. 6 in Revision 1). Furthermore, we added a new result in Revision 1 to demonstrate cAMP-dependent post-translational modulation of Orco open-time probability (Fig 9 in Revision 1) at a ZT at which antennal cAMP levels are low (Schendzielorz et al., 2015). We already showed in Flecke et al., 2010, that the addition of cAMP at different ZTs increases the spontaneous spiking activity only at specific ZTs. Here, we show that the effect of cAMP depends on Orco. Since Orco’s circadian role is not mediated via TTFL control, it can be concluded that post-translational mechanisms provide daily/circadian temporal control. In this additional Figure we provide statistically significant proof that, in agreement with our model-prediction, Orco’s circadian control of the ORN spontaneous activity could be mediated via the second messenger cAMP. ZT-dependent input for Orco would be provided via daily changes in cAMP levels (Schendzielorz et al., 2015).

      In contrast, the study does provide strong evidence that the application of cyclic nucleotides can modulate Orco-dependent activity at a single time point, and reports that the temporal pattern of Orco transcript abundance is not circadian.

      We appreciate that the reviewer confirms that our revised manuscript with additional experiments now provides strong evidence that cAMP modulates Orco-dependent spontaneous activity of M. sexta ORNs. Since we already published that cAMP levels expresses daily rhythms in M. sexta antennae (Schendzielorz et al., 2015), and in vivo cAMP infusion increases spontaneous activity and sensitizes pheromone detection (Flecke et al., 2010), and our computational model here proves that circadian modulation of open time probability of Orco is sufficient to explain our experimental data, it is sufficient for our conclusions to test just the one specific zeitgeber time when endogenous cAMP levels are low and pharmacological cAMP increase has the strongest impact. To further reveal complete ZT-dependence of cAMP modulation of Orco´s control of spontaneous activity is beyond the scope of the current manuscript and not part of the current research question. The structure of Orco is extraordinarily conserved during evolution, thus, the cited experimental results from other laboratories and other species showing that Orco is a hub for posttranslational modification are very likely generalizable to different insect species. We clarified the manuscript accordingly. In Drosophila, Orco has at least 5 phosphorylation sites for protein kinase C (PKC), is cAMP-dependently sensitized, and has a Ca<sup>2+</sup>/calmodulin binding site that orchestrates the localization of the OR-Orco heteromer to the cilia. However, in fruit fly and other insects, so far, it can only be speculated how circadian control is provided for Orco, since there are no other publications that examine the circadian regulation of Orco in detail. We clarified our manuscript accordingly.

      To summarize, the logical conclusion based on our newly provided data is that the current hierarchical hypothesis in chronobiology based solely on a circadian TTFL clock that controls Orco transcription does not explain our findings in hawkmoth ORNs. Therefore, we suggest a new systemic hypothesis based upon coupled TTFL and PTFL circadian clocks that can also reconcile otherwise inconsistent data published for insect and mammalian circadian clocks (please see reviewed data in: Stengl and Schneider, 2024). We clarified our manuscript in the second revision and added a new Figure 10 to illustrate our novel hypothesis.

      However, the findings are incomplete to exclude a role for transcriptional-translational mechanisms and their associated multi-layered controls in circadian regulation.

      We certainly do not imply excluding a role for the TTFL clock in (indirectly) affecting circadian control of the membrane potential of ORNs. The new qPCR experiments added in the first revision clearly show that the circadian control mediated via Orco is not an output of the TTFL clock via transcriptional control of Orco. Instead, we predict links between a PTFL membrane clock comprising Orco as hub to integrate posttranslational control and the TTFL nuclear clock. We clarified our manuscript accordingly, adding a new Figure 10 to further illustrate and visualize our hypothesis. The predictions of this systemic hypothesis will be challenged in further experiments that, however, are beyond the scope of the current manuscript.

      Joint Public Review:

      This manuscript puts forward the provocative idea that a posttranslational feedback loop regulates daily and ultradian rhythms in neuronal excitability. The authors used in vivo long-term tip recordings of the long trichoid sensilla of male hawkmoths to analyze spontaneous spiking activity indicative of the ORNs' endogenous membrane potential oscillations. This firing pattern was disrupted by pharmacological blockade of the Orco receptor. They then use these recordings together with computational modeling to predict that Orco receptor neuron (ORN) activity is required for circadian, not ultradian, firing patterns. Orco did not show a circadian expression pattern in a qPCR experiment, and its conductance was proposed to be regulated by cyclic nucleotide levels. This evidence led the authors to conclude that a post-translational feedback loop (PTFL) clockwork, associated with the ORN plasma membrane, allows for temporal control of pheromone detection via the generation of multi-scale endogenous membrane potential oscillations. The findings will interest researchers in neurophysiology, circadian rhythms, and sensory biology. However, the manuscript has limited experimental evidence to support its central hypothesis and is undermined by several assumptions that underlie their data analysis and model builds, as well as insufficient biological data including critical controls to validate and/or fully justify the model the authors are proposing.

      We want to remind our reviewers that we used “ORN” as abbreviation for olfactory receptor neuron (= sensory receptor neurons, a.k.a. olfactory sensory neuron (OSN)) and not for Orco receptor neuron, although we focus on the function of Orco. Accordingly, our central finding is that Orco as ion channel is required for the daily/circadian modulation of spontaneous action potential activity generated by the olfactory receptor neurons in the absence of pheromone stimulation.

      We do not understand the specific basis for the conclusions of the reviewers. Therefore, we ask to please specify what experimental evidence is missing to support our central hypothesis that Orco is not directly TTFL- but PTFL clock-controlled, and to name specifically what the “several assumptions” are that undermine our careful data analysis and model builds. Which specific argument in our previous rebuttal was wrong, was not conclusive? Furthermore, please specify your claim that “critical controls are missing”. Which controls are missing for which experiments? We did add a new figure panel in the first revision to demonstrate that neither the addition of DMSO (the OLC15 solvent) nor the repeated attachment of the recording electrode, which was necessary to obtain paired datasets, altered the spontaneous spiking activity (Fig 1B in Revision 1). Furthermore, we expanded the time series of qPCR data and added tim as positive control to Orco (Fig 7), strengthening our argument that Orco expression is not under TTFL control. Dose-response curves of various Orco agonists and antagonists have been published before (see our references in Revision 1) and are therefore not repeated here.

      Our newly added data confirm what our modeling predicted: cAMP increases spontaneous ORN activity dependent on Orco. Previous publications provide evidence for daily rhythms in cAMP concentrations in hawkmoth antennae (Schendzielorz et al., 2012).

      As is true for any other hypothesis, a hypothesis can only be falsified but not validated and needs to be tested by many experiments from many laboratories over a long time until it will be replaced by the next hypothesis that better explains accumulating contradicting evidence. We are very much looking forward to experimental challenges of our provocative new hypothesis by colleagues in the field of olfaction and of chronobiology. We are convinced that our manuscript will greatly stimulate the field, possibly provoking a paradigm switch in chronobiology and in olfactory research.

      Strengths:

      The authors raise several intriguing model-based hypotheses regarding the mechanisms that underlie the generation of olfactory rhythms. The electrophysiological approach and the long-term recording paradigm are elegant and technically impressive. In the revised version, the authors have added additional qPCR data supporting the lack of rhythmic Orco transcript expression and included a new figure suggesting that cAMP can modulate Orco conductance.

      We thank the reviewers for their acknowledgement of our careful work and hope that our further revisions and clarifications help to argue our case.

      Major weaknesses:

      (1) The cAMP experiment was only conducted at one time-point, which is insufficient to support the central claim that "AMP and cGMP may have ZT-dependent effects on Orco conductivity".

      We agree with the reviewers and revised our discussion accordingly to clarify that in this manuscript it is not our central claim that cAMP and cGMP may have ZT-dependent effects on Orco conductivity. Instead, our data show for the first time that Orco controls circadian rhythms of spontaneous activity of ORNs and that the circadian rhythmicity of spontaneous activity is lost when Orco is blocked. Therefore, we provide novel experimental evidence that Orco is a prerequisite to the circadian rhythmicity of spontaneous activity and thus, to circadian rhythms in the membrane potential of ORNs. Furthermore, as requested by the reviewers we provided clear evidence in the first revision that Orco is not controlled at the transcriptional level by the TTFL clock, in contrast to the TTFL clock protein TIMELESS. Thus, it follows logically that Orco is under post-transcriptional control. Since cAMP levels show circadian oscillations and Orco is gated by cAMP (Fig 9 in Revision 1), we used our computational model to show that a cAMP-dependent increase in Orco conductance alone, via daily oscillating concentrations of cAMP, is sufficient to explain our findings. Therefore, we propose here that daily/circadian oscillations of cAMP modify spontaneous spiking activity via Orco on a posttranscriptional level. But it certainly does not provide all evidence for respective mechanisms of how this cAMP modulation of Orco is obtained, since this is beyond the scope of the current manuscript.

      Since we realized that it is difficult for our readers to visualize a circadian PTFL membrane clock we added a new hypothesis-Figure (Figure 10) and considerably focused and clarified especially the discussion of our manuscript. We pointed out that a membrane-associated signalosome that comprises delayed negative feedback mechanisms, and, thus, constitutes an oscillator, a “membrane clock” that generates oscillations. Based upon our data we propose a membrane-associated signalosome constituting a PTFL circadian clock with Orco as central element. This PTFL membrane clock generates superimposed ultradian and circadian rhythms in its outputs: rhythms in the membrane potential, Ca<sup>2+</sup>, and cAMP levels. The PTFL clock comprises positive feedforward elements that upregulate its outputs, resulting in more depolarization, higher Ca<sup>2+</sup>- and higher cAMP levels. Via the clock’s delayed negative feedback mechanisms these outputs are downregulated, again, resulting in hyperpolarization, decreasing Ca<sup>2+</sup>- and cAMP levels. This signalosome comprises the pacemaker channel Orco as a central hub that is controlled via changes in voltage, Ca<sup>2+</sup>, and cAMP levels. Nevertheless, we predict coupling between the multiscale PTFL membrane clock and the TTFL circadian clock in the nucleus to obtain stable circadian rhythms. As likely mechanism of coupling we predict that Ca<sup>2+</sup>- and cAMP-dependent kinases interlink both types of clocks, thereby obtaining robust and at the same time flexible interlinked cellular rhythms.

      We hope to now successfully clarify and to visualize our central hypothesis of our manuscript that Orco is not directly controlled by a TTFL circadian clock but is a central element of a membrane-associated posttranslational feedback loop clock (PTFL) clock that is linked to but not forced by the TTFL clock which is predicted to control intracellular Ca<sup>2+</sup> homeostasis in a circadian rhythm.

      (2) The revised manuscript continues to rely heavily on prior publications or defers key mechanistic questions (or important manipulations) to future studies. In its current form, the evidence presented remains insufficient to support the central claim that a PTFL constitutes the primary underlying circadian clock mechanism. The proposed model is intriguing, but the data provided do not yet directly demonstrate the novel mechanism.

      We do not understand why the reviewers considers it to be problematic that we “continue to rely heavily on prior publications”. Certainly, we built upon previous publications of our lab as well as on manifold experimental data published by other laboratories in the field of insect olfaction. Our ample citations demonstrate that we have an overview both of the current state of literature and relevant previous literature, dating back to the very first experiments that pioneered pheromone transduction in insects. Based upon our extensive knowledge and experimental data collected in insect olfaction and based on very careful, critical, rigorous analysis of our data and data published by others, we were able to come up with a novel interpretation of the current literature about insect olfaction that differs considerably from the current main views. We consider this to be our strength and judge it as good scientific practice and not a flaw of our work. However, since we do not focus on OR-Orco heteromers and their function in pheromone/odor transduction in the cilia in the current manuscript, we considerably shortened this part of the discussion, avoiding pointing out that highly sensitive moth pheromone transduction greatly differs from less sensitive general odor transduction in Drosophila. Furthermore, since here we focus on cAMP, but not on cGMP-dependent modulation of Orco, we also deleted/considerably shortened this part of our discussion.

      We certainly agree with the reviewers that, while the data provided in the current manuscript are a logical basis for developing our novel hypothesis, they are not a direct and sufficient demonstration of proof and we are not able yet to directly demonstrate and explain the novel mechanism predicted. We would like to point out that if we provided this final proof, it would not be any more a novel hypothesis, but only a novel finding.

      We agree with the reviewers that our provocative hypothesis requires rigorous testing by many further experiments, hopefully not only by our laboratory, but hopefully stimulating new experimental challenges by other laboratories employing different species. But certainly, these experiments with proof-of-principle will take many years and are beyond the scope of our current research paper.

      As per eLife’s assessment system we would like to ask the reviewers to provide detailed feedback as to which experiments/results within the scope of this manuscript would complete this work, or how they think this study should be framed in the light of the results that we obtained. Nevertheless, we hope that with our current careful review the reviewers will be more convinced by our arguments and experiments as valid basis for our provocative new hypothesis.

    1. eLife Assessment

      This fundamental work significantly advances our understanding of how contact-dependent antagonism enables keystone bacteria to establish and maintain their niche over time. The evidence obtained is convincing, supporting most of the conclusions drawn. This work will be of significant interest to the microbiome research community.

    2. Reviewer #1 (Public review):

      Summary:

      In this study, the authors investigate the physiological role of the Type VI secretion system (T6SS) in a naturally evolved gut microbiome derived from wild mice (the WildR microbiome). Focusing on Bacteroides acidifaciens, the authors use newly developed genetic tools and strain replacement strategies to test how T6SS-mediated antagonism influences colonization, persistence, and fitness within a complex gut community. They further show that the T6SS resides on an integrative and conjugative element (ICE), is distributed among select community members, and can be horizontally transferred, with context-dependent effects on colonization and persistence. The authors conclude that the T6SS stabilizes strain presence in the gut microbiome while imposing ecological and physiological constraints that shape its value across contexts.

      This study is likely to have significant impact on the microbiome field by moving experimental tests of T6SS function out of simplified systems and into a naturally co-evolved gut community. The WildR system, together with the strain replacement strategy, ICE-seq approach, and genetic toolkit, represents a powerful and reusable platform for future mechanistic studies of microbial antagonism and mobile genetic elements in vivo.

      The datasets-including isolate genomes, metagenomes, and ICE distribution maps-will be valuable community resources, particularly for researchers interested in strain-resolved dynamics, horizontal gene transfer, and ecological context dependence. Even where mechanistic resolution is incomplete, the work provides a strong experimental foundation upon which such questions can be directly addressed.

      Overall, this study occupies a space between system building and mechanistic dissection. The authors demonstrate that the T6SS influences persistence and community structure in vivo, but the physiological basis of these effects remains unresolved. Interpreting the results as evidence of fitness costs or selective advantage therefore requires caution, as multiple ecological and host-mediated processes could produce similar abundance trajectories.

      Placing the findings within the broader literature on microbial antagonism, particularly work emphasizing measurable costs, benefits, and tradeoffs, would help readers better contextualize what is directly demonstrated here versus what remains an open question. Viewed in this light, the principal contribution of the study is to show that such questions can now be addressed experimentally in a realistic gut ecosystem.

      Strengths:

      A major strength of this study is that it directly interrogates the physiological role of the T6SS in a naturally evolved gut microbiome, rather than relying on simplified pairwise or in vitro systems. By working within the WildR community, the authors advance beyond descriptive surveys of T6SS prevalence and address function in an ecologically relevant context.

      The authors provide clear genetic evidence that Bacteroides acidifaciens uses a T6SS to antagonize co-resident Bacteroidales, and that loss of T6SS function specifically compromises long-term persistence without affecting initial colonization. This temporal separation is well designed and supports the conclusion that the T6SS contributes to maintenance rather than establishment within the community.

      Another strength is the identification of the T6SS on an integrative and conjugative element (ICE) and the demonstration that this element is distributed among, and exchanged between, community members. The use of ICE-seq to track distribution and transfer provides strong support for horizontal mobility and adds mechanistic depth to the study.

      Finally, the transfer of the T6SS-ICE into Phocaeicola vulgatus and the observation of context-dependent colonization benefits followed by decline is a compelling result that moves the study beyond simple "T6SS is beneficial" narratives and highlights ecological contingency.

      Weaknesses:

      Despite these strengths, there is a mismatch between the precision of the claims and the precision of the measurements, particularly regarding fitness costs, physiological burden, and mechanistic role of the T6SS.

      First, while the authors conclude that the T6SS "stabilizes strain presence" and that its value is constrained by fitness costs, these costs are not directly measured. Persistence, abundance trajectories, and eventual loss are informative outcomes, but they do not uniquely identify fitness tradeoffs. Decline could arise from multiple non-exclusive mechanisms, including community restructuring, host-mediated effects, incompatibilities of the ICE in new hosts, or ecological retaliation, none of which are disentangled here.

      Second, the manuscript frames the T6SS as having a defined physiological role, yet the data do not resolve which physiological processes are under selection. The experiments demonstrate that T6SS activity affects persistence, but they do not distinguish whether this occurs via direct killing, resource release, niche modification, or higher-order community effects. As a result, "physiological role" remains underspecified and risks being conflated with ecological outcome.

      Third, although the authors emphasize context dependence, the study offers limited quantitative insight into what aspects of context matter. Differences between native and recipient hosts, or between early and late colonization phases, are described but not mechanistically interrogated, making it difficult to generalize beyond the specific cases examined.

      Fourth is the lack of engagement with recent experimental literature demonstrating functional roles of the T6SS beyond simple interference competition. While the authors focus on persistence and competitive outcomes, they do not adequately situate their findings within recent work demonstrating that T6SS-mediated antagonism can serve additional physiological functions, including resource acquisition and DNA uptake, thereby linking killing to measurable benefits and tradeoffs. The absence of this literature makes it difficult to place the authors' conclusions about physiological role and fitness cost within the current conceptual framework of the field. Without this context, the physiological interpretation of the results remains incomplete, and alternative functional explanations for the observed dynamics are underexplored.

      A further limitation concerns the taxonomic scope of the functional analysis. The authors state the role of the T6SS in the murine environment is functionally investigated using genetically tractable Bacteroides species, citing lack of genetic tools for Mucispirillum schaedleri. While this is a reasonable practical choice, it means that a substantial fraction of T6SS-encoding species in the WildR community are not experimentally interrogated. Consequently, conclusions about the role of the T6SS in the murine gut necessarily reflect the subset of taxa that are genetically accessible and may not fully capture community-level or niche-specific functions of T6SS activity. Given that M. schaedleri is represented as a metagenome-assembled genome, its isolation and genetic manipulation would be technically challenging. Nonetheless, explicitly acknowledging this limitation and slightly tempering claims of generality would strengthen the manuscript.

      Finally, several interpretations would benefit from more cautious language. In particular, claims invoking fitness costs, selective advantage, or physiological burden should be explicitly framed as inferences from persistence dynamics, rather than as direct measurements, unless supported by additional quantitative fitness or growth assays.

      Comments on revised version.

      The authors have addressed my main concerns by more clearly distinguishing ecological outcomes from directly measured physiological mechanisms. They have moderated claims about fitness costs and benefits, replaced "physiological" with "ecological" where appropriate, expanded the discussion of potential downstream benefits of T6SS-mediated killing, and acknowledged the limited taxonomic scope of the functional analyses. The persistence trajectories support context-dependent relative fitness effects, although they do not identify the specific physiological basis of those effects. The revised manuscript now generally maintains this distinction. These revisions substantially improve the precision and balance of the manuscript.

    3. Reviewer #2 (Public review):

      Summary:

      In this study, the authors set out to determine how a contact-dependent bacterial antagonistic system contributes to the ability of specific bacterial strains to persist within a complex, native gut community derived from wild animals. Rather than focusing on simplified or artificial models, the authors aimed to examine this system in a biologically realistic setting that captures the ecological complexity of the gut environment. To achieve this, they combined controlled laboratory experiments with animal colonization studies and sequencing-based tracking approaches that allow individual strains and mobile genetic elements to be followed over time.

      Strengths:

      A major strength of the work is the integration of multiple complementary approaches to address the same biological question. The use of defined but complex communities, together with in vivo experiments, provides a strong ecological context for interpreting the results. The data consistently show that the antagonistic system is not required for initial establishment but plays a critical role in long-term strain persistence, an insight that moves beyond traditional invasion-based views of microbial competition. The observation that transferable genetic elements can confer only temporary advantages, and may impose longer-term costs depending on community context, adds important nuance to current understanding of microbial fitness.

      Weaknesses:

      Overall, the study is not a lack of evidence, but a deliberate trade-off between ecological realism and mechanistic resolution, which leaves some causal pathways open to interpretation.

      Comments on revised version.

      The authors have addressed all previous concerns thoroughly and satisfactorily.

    4. Reviewer #3 (Public review):

      Summary:

      In this work, the authors investigate the contribution of the type VI secretion system of Bacteroidales to gut microbiome assembly and the targeting of closely related species. They demonstrate that B. acidifaciens relies on T6SS-mediated antagonism to prevent displacement by co-resident Bacteroidales and other members of the microbiome, allowing it to persist in the gut. They also developed new tools for analyzing the distribution of mobile genetic elements. This study advances our understanding of how molecular systems contribute to shaping complex microbial communities.

      Strengths:

      The use of a gnotobiotic model colonized with a wild-mouse microbiome is a significant strength of this study. This approach allows tracking of microbiome changes over time and evaluating the targeting by Bacteroidales carrying T6SS in a more natural setting. The development of ICE-seq for mapping the distribution of the T6SS in the microbiome is remarkable, enabling the study of how this bacterial weapon is transferred between microbiome members without requiring long-read metagenomics methods.

      Weaknesses:

      Some conclusions are based on a limited number of mice per condition. This could be due to the complexity of using a gnotobiotic mouse model, but this should be considered when interpreting the data.

      Overall, the authors successfully achieved their objectives, and their experimental design and results support their findings. As mentioned in the discussion, it would be important to investigate the role of the T6SS in resilience to microbiome disturbances, such as antibiotics, diet, or pathogen invasion. This work represents a step forward in understanding how contact-dependent competition influences the gut microbiome in relevant ecological contexts.

    5. Author response:

      The following is the authors’ response to the original reviews.

      We appreciate that the reviewers provided an overall positive assessment of our manuscript and offered constructive suggestions for improvement. All three reviewers noted that a key strength of our study is the implementation of a gut microbiome model for the characterization of interbacterial antagonism pathways such as the type VI secretion system (T6SS) that approaches natural complexity. They note our work represents a significant advance in microbiome research, and generates resources that will be of use to many researchers in the field. Two of the reviewers point out that the complexity of our model limits the nature of measurements we can make, and suggest we temper the strength of the some of the conclusions we draw. As noted in more detail below, in our revised manuscript, we have used more precise wording to characterize our findings, and we are more explicit about the connection between the measurements we made and what we can conclude about the physiological role of the T6SS in the gut microbiome.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      In this study, the authors investigate the physiological role of the Type VI secretion system (T6SS) in a naturally evolved gut microbiome derived from wild mice (the WildR microbiome). Focusing on Bacteroides acidifaciens, the authors use newly developed genetic tools and strain-replacement strategies to test how T6SS-mediated antagonism influences colonization, persistence, and fitness within a complex gut community. They further show that the T6SS resides on an integrative and conjugative element (ICE), is distributed among select community members, and can be horizontally transferred, with context-dependent effects on colonization and persistence. The authors conclude that the T6SS stabilizes strain presence in the gut microbiome while imposing ecological and physiological constraints that shape its value across contexts.

      This study is likely to have a significant impact on the microbiome field by moving experimental tests of T6SS function out of simplified systems and into a naturally coevolved gut community. The WildR system, together with the strain replacement strategy, ICE-seq approach, and genetic toolkit, represents a powerful and reusable platform for future mechanistic studies of microbial antagonism and mobile genetic elements in vivo.

      The datasets, including isolate genomes, metagenomes, and ICE distribution maps, will be a valuable community resource, particularly for researchers interested in strainresolved dynamics, horizontal gene transfer, and ecological context dependence. Even where mechanistic resolution is incomplete, the work provides a strong experimental foundation upon which such questions can be directly addressed.

      Overall, this study occupies a space between system building and mechanistic dissection. The authors demonstrate that the T6SS influences persistence and community structure in vivo, but the physiological basis of these effects remains unresolved. Interpreting the results as evidence of fitness costs or selective advantage, therefore, requires caution, as multiple ecological and host-mediated processes could produce similar abundance trajectories.

      Placing the findings within the broader literature on microbial antagonism, particularly work emphasizing measurable costs, benefits, and tradeoffs, would help readers better contextualize what is directly demonstrated here versus what remains an open question. Viewed in this light, the principal contribution of the study is to show that such questions can now be addressed experimentally in a realistic gut ecosystem.

      We thank the reviewer for this thoughtful summary of our study. We were glad to read they conclude our work will have a significant impact on the microbiome field and that the resources we have developed will be of value to the community.

      Strengths:

      A major strength of this study is that it directly interrogates the physiological role of the T6SS in a naturally evolved gut microbiome, rather than relying on simplified pairwise or in vitro systems. By working within the WildR community, the authors advance beyond descriptive surveys of T6SS prevalence and address function in an ecologically relevant context.

      The authors provide clear genetic evidence that Bacteroides acidifaciens uses a T6SS to antagonize co-resident Bacteroidales, and that loss of T6SS function specifically compromises long-term persistence without affecting initial colonization. This temporal separation is well designed and supports the conclusion that the T6SS contributes to maintenance rather than establishment within the community.

      Another strength is the identification of the T6SS on an integrative and conjugative element (ICE) and the demonstration that this element is distributed among, and exchanged between, community members. The use of ICE-seq to track distribution and transfer provides strong support for horizontal mobility and adds mechanistic depth to the study.

      Finally, the transfer of the T6SS-ICE into Phocaeicola vulgatus and the observation of context-dependent colonization benefits followed by decline is a compelling result that moves the study beyond simple "T6SS is beneficial" narratives and highlights ecological contingency.

      We appreciate this detailed and nuanced characterization of the strengths of our study.

      Weaknesses:

      Despite these strengths, there is a mismatch between the precision of the claims and the precision of the measurements, particularly regarding fitness costs, physiological burden, and the mechanistic role of the T6SS.

      We acknowledge that in some places, our manuscript could benefit from greater precision in the language we use when linking the outcomes we observe in our study to their potential underlying causes. Specific revisions we made to address this concern are described below.

      First, while the authors conclude that the T6SS "stabilizes strain presence" and that its value is constrained by fitness costs, these costs are not directly measured. Persistence, abundance trajectories, and eventual loss are informative outcomes, but they do not uniquely identify fitness tradeoffs. Decline could arise from multiple nonexclusive mechanisms, including community restructuring, host-mediated effects, incompatibilities of the ICE in new hosts, or ecological retaliation, none of which are disentangled here.

      We agree that multiple mechanisms could explain why populations of certain species carrying a T6SS decline over time, and why for others, the T6SS contributes to long-term persistence. Our use of the term “fitness cost” to describe the phenomenon of decline observed for P. vulgatus carrying the T6SS was not meant to imply any particular underlying mechanism, but was rather our attempt to characterize the phenotypic outcome we observed in simplified terms. We note that ecological context is an important determinant of the fitness cost or benefit of any given trait, and our study sheds light on the importance of the presence of the WildR community and the mouse intestinal environment to the fitness contribution of the T6SS to B. acidifaciens and P. vulgatus. Nonetheless, to avoid implying an overly simplistic interpretation of our results, we have modified our language in the manuscript in several places when describing the role of the T6SS in species persistence in mice colonized with the WildR community.

      Second, the manuscript frames the T6SS as having a defined physiological role, yet the data do not resolve which physiological processes are under selection. The experiments demonstrate that T6SS activity affects persistence, but they do not distinguish whether this occurs via direct killing, resource release, niche modification, or higher-order community effects. As a result, "physiological role" remains underspecified and risks being conflated with ecological outcome.

      We acknowledge that our study does not fully resolve the physiological processes under selection that mediate role of the T6SS in maintaining B. acidifaciens populations in WildR-colonized mice. Indeed, several of the outcomes of T6SS activity the reviewer lists, such as target cell killing and nutrient release, are inextricably linked and thus inherently difficult to disentangle. We note that we did attempt to measure higher-order community effects of T6SS activity with metagenomic sequencing, but acknowledge that this approach may not have been sufficiently sensitive to detect small community shifts mediated by a relatively low-abundance species. To address the concern that our current framing implies more of a mechanistic understanding that our study achieves, we have substituted “ecological” for “physiological” where appropriate throughout the manuscript.

      Third, although the authors emphasize context dependence, the study offers limited quantitative insight into what aspects of context matter. Differences between native and recipient hosts, or between early and late colonization phases, are described but not mechanistically interrogated, making it difficult to generalize beyond the specific cases examined.

      We are not entirely clear what the reviewer means by “differences between native and recipient hosts”, but we agree that additional quantitative studies will be needed to address the generalizability of our findings. Future studies are also needed to address the mechanistic basis for the difference in the benefit conferred by the T6SS that we observed between P. vulgatus and B. acidifaciens.

      Fourth is the lack of engagement with recent experimental literature demonstrating functional roles of the T6SS beyond simple interference competition. While the authors focus on persistence and competitive outcomes, they do not adequately situate their findings within recent work demonstrating that T6SS-mediated antagonism can serve additional physiological functions, including resource acquisition and DNA uptake, thereby linking killing to measurable benefits and tradeoffs. The absence of this literature makes it difficult to place the authors' conclusions about physiological role and fitness cost within the current conceptual framework of the field. Without this context, the physiological interpretation of the results remains incomplete, and alternative functional explanations for the observed dynamics are underexplored.

      We thank the reviewer for specifically highlighting the potential pertinence of this literature to our study. Indeed, we did not cite studies indicating a link between T6SS activity and the uptake of DNA and other resources released by targeted cells. As we note above, the release of intracellular contents from target cells is an inevitable consequence of the delivery of lytic effectors. Thus, distinguishing between fitness benefits conferred from the elimination of competitor species and those arising from scavenging the nutrients released during this process is not straightforward. Measuring the benefits deriving from the uptake of certain released molecules, such as DNA, was not immediately feasible in the system employed in this study and instead we focused on the direct lytic consequences of the effectors delivered via the T6SS. We revised the Discussion to include reference to these possible downstream benefits of T6SS activity (Lines 476-479).

      A further limitation concerns the taxonomic scope of the functional analysis. The authors state that the role of the T6SS in the murine environment is functionally investigated using genetically tractable Bacteroides species, citing the lack of genetic tools for Mucispirillum schaedleri. While this is a reasonable, practical choice, it means that a substantial fraction of T6SS-encoding species in the WildR community are not experimentally interrogated. Consequently, conclusions about the role of the T6SS in the murine gut necessarily reflect the subset of taxa that are genetically accessible and may not fully capture community-level or niche-specific functions of T6SS activity. Given that M. schaedleri is represented as a metagenome-assembled genome, its isolation and genetic manipulation would be technically challenging. Nonetheless, explicitly acknowledging this limitation and slightly tempering claims of generality would strengthen the manuscript.

      The reviewer points out that studying the T6SS activity in M. schadleri would potentially expand the generality of our claims. We agree that having an isolate of this species along with genetic tools for its manipulation would allow us to probe the importance of the T6SS in the gut microbiome more broadly. At the suggestion of the reviewer, we have added explicit mention of the potential benefit of studying the T6SS in this organism to the Discussion (lines 538-539), an endeavor that lies outside of the scope of the current study.

      Finally, several interpretations would benefit from more cautious language. In particular, claims invoking fitness costs, selective advantage, or physiological burden should be explicitly framed as inferences from persistence dynamics, rather than as direct measurements, unless supported by additional quantitative fitness or growth assays.

      We agree with the reviewer that invoking fitness costs, selective advantages or physiological burdens should be done cautiously, and have made revisions to our manuscript where we acknowledge that more precise language was needed (line 43, 416, line 417). However, we would also argue invoking fitness costs and benefits when describe strain persistence dynamics in mice has substantial precedent in the literature (Feng et al. 2020, Brown et al. 2021, Park et al. 2022, Segura Munoz et al. 2022), to list a handful of representative examples published by different groups). It is unclear to us what additional in vivo growth measurements could be taken to substantiate our claim that the T6SS provides a fitness benefit to B. acidifaciens during prolonged gut colonization, or that carrying the ICE imposes a fitness cost on P. vulgatus during longterm colonization. Our in vitro experiments evaluating the competitiveness conferred by T6SS activity provide a measure of insight into its fitness benefits, but as our in vivo strain persistence data and the work of many others show, in vitro measurements do not necessarily capture in vivo parameters.

      Reviewer #2 (Public review):

      Summary:

      In this study, the authors set out to determine how a contact-dependent bacterial antagonistic system contributes to the ability of specific bacterial strains to persist within a complex, native gut community derived from wild animals. Rather than focusing on simplified or artificial models, the authors aimed to examine this system in a biologically realistic setting that captures the ecological complexity of the gut environment. To achieve this, they combined controlled laboratory experiments with animal colonization studies and sequencing-based tracking approaches that allow individual strains and mobile genetic elements to be followed over time.

      Strengths:

      A major strength of the work is the integration of multiple complementary approaches to address the same biological question. The use of defined but complex communities, together with in vivo experiments, provides a strong ecological context for interpreting the results. The data consistently show that the antagonistic system is not required for initial establishment but plays a critical role in long-term strain persistence. This insight that moves beyond traditional invasion-based views of microbial competition. The observation that transferable genetic elements can confer only temporary advantages, and may impose longer-term costs depending on community context, adds important nuance to current understanding of microbial fitness.

      We thank the reviewer for the positive feedback and are glad they agree our study provides new insight into the role of interbacterial antagonism in natural communities.

      Weaknesses:

      Overall, there is not a lack of evidence, but a deliberate trade-off between ecological realism and mechanistic resolution, which leaves some causal pathways open to interpretation.

      The reviewer makes a good point that the complexity of the experimental system we employ precludes some lines of experimentation that would yield more mechanistic information. As the reviewer notes, we were aware of the tradeoff between mechanistic resolution and ecological realism when selecting our experimental system. Our deliberate choice to favor biological complexity over mechanistic clarity in this study stemmed from our perception that a major gap in understanding of the T6SS and other antagonism pathways lies in defining their ecological function in complex microbial communities.

      Reviewer #3 (Public review):

      Summary:

      Shen et al. investigate the contribution of the type VI secretion system of Bacteroidales in the gut microbiome assembly and targeting of closely related species. They demonstrate that B. acidifaciens relies on T6SS-mediated antagonism to prevent displacement by co-resident Bacteroidales and other members of the microbiome, allowing B. acidifaciens to persist in the gut.

      Strengths:

      Using a gnotobiotic model colonized with a wild-mouse microbiome is a significant strength of this study. This approach allows tracking of microbiome changes over time and directly examining targeting by Bacteroidales carrying T6SS in a more natural setting. The development of ICE-seq for mapping the distribution of the T6SS in the microbiome is remarkable, enabling the study of how this bacterial weapon is transferred between microbiome members without requiring long-read metagenomics methods.

      We thank the reviewer for their enthusiasm toward our study.

      Weaknesses:

      Some conclusions are based on only four mice per condition. The author should consider increasing the sample size.

      We agree that in some experiments it would be beneficial to increase the sample size from four mice. However, the experiments we performed for this study are time and resource-intensive. Additionally, the experiments on which we base our primary conclusions were all independently replicated with similar results. Given these factors, we determined that the extra confidence that might be afforded by increasing our sample size did not merit the delay in publication and investment in resources that would be required.

      Overall, the authors successfully achieved their objectives, and their experimental design and results support their findings. As mentioned in the discussion, it would be important to investigate the role of the T6SS in resilience to disturbances in the microbiome, such as antibiotics, diet, or pathogen invasion. This work represents a step forward in understanding how contact-dependent competition influences the gut microbiome in relevant ecological contexts.

      We agree that investigating the role of the T6SS during perturbations of the microbiome is a key next step for this work and thank the reviewer for highlighting this important future direction.

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      Beth A. Shen et al. present a comprehensive and carefully executed study investigating the ecological role of the type VI secretion system (T6SS) in maintaining bacterial strains within a native, complex gut microbiome derived from wild mice. By integrating genetic manipulation, metagenomic and sequencing-based tracking approaches, in vitro competition assays, and gnotobiotic colonization experiments, the authors provide compelling evidence that the T6SS functions primarily as a persistence factor rather than a determinant of initial colonization.

      The study is conceptually strong and addresses an important gap in our understanding of how interbacterial antagonistic systems operate in complex, native microbial communities. The manuscript is generally well organized, the data are clearly presented, and the main conclusions are supported by robust experimental evidence. In particular, the demonstration that T6SS-encoding ICEs confer context- and hostdependent fitness effects, including transient benefits and potential long-term costs, adds important nuance to prevailing models of microbial competition.

      That said, several aspects of the study would benefit from clarification and deeper mechanistic discussion. Addressing the points below would further strengthen the rigor and interpretability of the work. Overall, this is a strong and interesting manuscript, requiring some revisions.

      We greatly appreciate the positive summary of our work by the reviewer, which highlights the multi-faceted approach we took to address gaps in our understanding of interbacterial antagonism in the microbiome. It is our hope that the reviewer agrees that our revisions of the manuscript, based on their feedback, clarify our methods and strengthen the interpretability of our work.

      Major comments

      (1) The competition assays in Figure 2B suggest that T6SS-dependent fitness effects are most pronounced among members of the order Bacteroidales. However, these experiments primarily measure population-level competitive outcomes rather than direct T6SS-mediated targeting events. In addition, the limited number of non-Bacteroidales strains included in the assay makes it difficult to conclude that T6SS activity is strictly restricted to closely related taxa.

      The authors should either temper their conclusions regarding target specificity or clarify that these data reflect competitive outcomes rather than direct evidence of targeting. Expanding the discussion to acknowledge these limitations would improve interpretative accuracy.

      With regards to the measurement we employed for assessing T6SS-mediated targeting, we acknowledge that this is, to a degree, an indirect way of determining T6SS targeting. However, there is extensive precedent in the literature for the use of similar assays in assessing targeting by many contact-dependent antagonism systems including the T6SS in Bacteroidales (Russell et al. 2014, Chatzidaki-Livanis et al. 2016, Wexler et al. 2016) and many Proteobacteria (e.g. (Hood et al. 2010)), the T4SS in Xanthomonas citri (Souza et al. 2015), the CDI system in Escherichia coli (Aoki et al. 2005), and the Esx system in Streptococcus intermedius (Whitney et al. 2017). In these studies, targeting was demonstrated by specific depletion of the competitor strain in the presence of a strain encoding an active antagonism system. We acknowledge that the competitive index we report Figure 2B reflects the relative population levels of both species in the assay, and thus does not directly show target species depletion. We opted to use this metric to display the data in the manuscript as a way of efficiently encapsulating and comparing many strain combinations in a single figure, and because the competitive index differences we observed in these derive from differences in target species growth yields (see Author response image 1, indicating growth yields from a representative strain pairing).

      Author response image 1.

      The T6SS of B. acidifaciens targets a WildR-derived P. vulgatus strain. CFUs indicate populations of the indicated strains after co-culture of wild-type or T6SS-inactivated B. acidifaciens with P. vulgatus. Data represent means and standard errors (n=3, *P<0.01, t-test with log ><0.01, t- test with log transformed data)

      We additionally acknowledge that more extensive testing is needed to fully understand the target range of the Bacteroidales T6SS. In our study, we assessed targeting of every WildR species that was readily culturable, which to the best of our knowledge, represents the broadest panel of targets for the Bacteroidales T6SS to be tested to date. We limited our testing to these strains, as the goal of these experiments was to gain insight into which co-residents of the WildR could be targeted by B. acidifaciens. We agree that testing of a broader cross-section of potential targets has merits, but this would require targeted cultivation strategies to obtain these organisms, and lies outside the scope of the current study. We have revised the manuscript to clarify that the target range testing encompassed the diversity of isolates available (p. 10, lines 231-236).

      (2) The bae1 gene encoded in Bacteroides caecimuris F12 contains a frameshift mutation. It would be valuable for the authors to comment on whether such frameshift mutations are a common genomic feature among gut-associated Bacteroides species in murine models. In addition, comparative analysis of human gut metagenomic datasets could reveal whether homologous effector proteins are present in commensal Bacteroides populations, and whether these homologs exhibit similar disruptive mutations.

      More broadly, the manuscript would benefit from a discussion of whether expression of a fully functional bae1 effector might impose a fitness cost on Bacteroidales members, for example, through metabolic burden or altered resource allocation. This is particularly relevant in light of recent studies demonstrating that T6SS effectors can drive physiological trade-offs by modulating metabolic dynamics (PMID: 40592326). Integrating this perspective would strengthen the evolutionary interpretation of effector mutagenesis.

      We agree with the reviewer that the functional and evolutionary significance of the point mutation in bae1 merits further investigation. Following the reviewer's suggestion, we looked in our own datasets and available public datasets from mouse and human microbiomes for evidence of bae1 inactivation. Unfortunately, the gene is present at a low enough frequency that these analyses were inconclusive. In our own metagenomic data from WildR mice, we did not obtain sufficient sequencing depth to assess the frequency at which bae1 is inactivated across genomes. We found a single complete copy of bae1 identical to that of B. acidifaciens in one published mouse microbiome-derived MAG, and detected fragments of the gene in a number of publicly available isolate and MAG genomes, but these were too low of quality to assess whether or not the gene was intact.

      As to whether or not bae1 expression imposes a fitness cost in the producing organism, we think this is unlikely to be significant, given that the impacts of Bae1 will be neutralized by the accompanying immunity protein. We speculate that the point mutation in the B. caecimuris gene is more likely to have arisen through genetic drift than as a result of selection.

      (3) Quantification and tracking of ICE transfer in vivo. In Figure 4D, the authors assess the abundance of resident P. vulgatus populations in germ-free mice co-gavaged with wild-type strains and derivatives carrying either the intact ICE or ICE ΔtssC. Because both ICE variants are capable of horizontal transfer, it is essential to clearly describe how the authors distinguish between (i) the original wild-type strain, (ii) engineered donor strains, and (iii) recipient strains that have newly acquired the ICE or ICE ΔtssC.

      Clarification of the specific molecular or sequencing-based strategies used to discriminate these populations is necessary to ensure accurate interpretation of the colonization dynamics.

      In this experiment, the P. vulgatus strains we introduced which carried the ICE (either the wild-type version or ICE DtssC) also contained an erythromycin resistance cassette (ermG) inserted distal to the ICE insertion site. Populations of the ICE-containing strain were quantified by either qPCR targeting the ermG gene (Fig. 4D, Supplemental Fig. 4D) or by plating on erythromycin-containing media (Fig. 4F). Endogenous P. vulgatus populations were quantified by qPCR targeting the ermG insertion site, which is disrupted in the marked strain. These methodological details have been added to the figure legend for clarity. We acknowledge that transfer of the ICE between introduced and endogenous populations is possible, and would not be detected by these metrics. To assess whether this occurs, we performed ICE-seq analyses on samples collected from mice colonized by the WildR and P. vulgatus ICE at early (7 days) and late (56 days) time points. These analyses revealed that overall, ICE distribution in this experiment was similar to that observed in mice colonized with the WildR alone (Figure 4A and Author response image 2). They additionally provided corroborating evidence that the population of ICE-containing P. vulgatus declined over the course of the experiment. Importantly, the only ICE insertion site we detected in P. vulgatus in these samples was that found in the introduced P. vulgatus strain. Previous studies show that GA1-containing ICE can insert at numerous locations in Bacteroides sp. genomes, a finding supported by our mapping of the ICE insertion sites from in vitro transfer experiments (Supplemental Fig. 4C) (Garcia-Bayona et al. 2021). Thus, our ICE-seq detection of a sole P. vulgatus ICE insertion site indicates that transfer of the element between P. vulgatus populations is likely not occurring in our experiments.

      Author response image 2.

      ICE-seq analysis indicates that introduction of P. vulgatus ICE into WildR-colonized mice has little impact on ICE distribution among endogenous strains. Graphs show frequency of mapped ICE junctions deriving from the indicated species as determined by 5¢ or 3¢ ICE-Seq analysis of DNA extracted from fecal samples collected either 7 or 56 days post-gavage of the WildR and P. vulgatus ICE into germ-free mice.

      (4) The analysis of fitness trade-offs associated with ICE acquisition in P. vulgatus convincingly demonstrates that the benefits of ICE transfer are transient and contextdependent. However, the mechanistic basis of these trade-offs remains underexplored. While the study primarily attributes both benefits and costs to T6SS-mediated antagonism, the ICE likely encodes additional genes that could influence metabolism, regulation, or stress responses.

      We agree with the reviewer that there are many mechanistic questions remaining regarding the benefits and costs associated with ICE acquisition, and acknowledge that we have not investigated the fitness contributions of ICE-encoded genes other than the T6SS. Indeed, as we noted in our discussion of the results from introducing P. vulgatus carrying the ICE into WildR-carrying mice, our data suggest that ICE genes outside the T6SS may be beneficial (lines 463-465). At the reviewer’s suggestion, we have reiterated the importance of considering the fitness contribution of genes beyond the T6SS in determining ICE distribution in the WildR community (line 527).

      Minor comments:

      (1) In lines 319 and 333, the manuscript refers to "Supplemental Figure 3F" and "Supplemental Figure 3G," respectively. However, the provided Supplemental Figure 3 appears to end at panel E. Please clarify or correct these references.

      We have modified the text to reference the correct figure panels.

      (2) Line 1043: The notation for "OD600" should be corrected for consistency and accuracy.

      The notation for OD600 has been updated to be consistent throughout the manuscript.

      Reviewer #3 (Recommendations for the authors):

      Minor comments:

      (1) Line 144. I would be careful of using "strong correlation, in this sentence. Although it shows a higher correlation than lab mice. Also, the labels in Figure 1A for mouse WildRF7 are confusing and not well explained in the figure legend.

      We modified line 147 (new line in edited manuscript) to say “positive correlation” rather than “strong correlation” to better represent the result. We also revised the legend for Figure 1A to better explain the samples of WildR F7 that were analyzed.

      (2) Line 155. It's unclear which strains were isolated from the WildR community, and the reason for isolating only 15 strains. Also, Supplemental Figure 1 shows 17 isolates, not 15.

      We apologize for the confusion here. We isolated 17 strains, which is the number of distinct strains we were able to readily culture from this community. We obtained genome sequences for 15 of these, and were able to assemble a genome for one more of the strains from metagenomic data.

      (3) Line 236. Is it known what makes B. uniformis resistant to B. acidifaciens carrying a T6SSS? Does it have an orphan immunity protein?

      We do not know why B. uniformis is not targeted by B. acidifaciens under the conditions of our experiments. It does not encode homologs of the immunity genes bai1 or bai2, and does not appear to be intrinsically resistant to targeting by this T6SS given that it is effectively targeted by P. vulgatus carrying the ICE (Fig.4b).

      (4) Line 295. There is a consistent decline in C. acid abundance after 27 days in Figures 3B and 3C. How do you explain this? Is the endogenous B. acid expanding to outcompete C. acid exo since the total C. acid exo remains constant when gavaging 100x B. acid exo?

      We believe that the eventual decline in the introduced population of B. acidifaciens is likely due to a fitness cost imposed by the erm resistance marker we employed. We noted this phenomenon when describing the results depicted in Fig. 3F-H, but neglected to include this explanation earlier. This oversight has been corrected (lines 315-317).

      References

      Aoki, S. K., R. Pamma, A. D. Hernday, J. E. Bickham, B. A. Braaten and D. A. Low (2005). "Contact-dependent inhibition of growth in Escherichia coli." Science 309(5738): 1245–1248.

      Brown, E. M., H. Arellano-Santoyo, E. R. Temple, Z. A. Costliow, M. Pichaud, A. B. Hall, K. Liu, M. A. Durney, X. Gu, D. R. Plichta, C. A. Clish, J. A. Porter, H. Vlamakis and R. J. Xavier (2021). "Gut microbiome ADP-ribosyltransferases are widespread phage-encoded fitness factors." Cell Host Microbe 29(9): 1351-1365 e1311.

      Chatzidaki-Livanis, M., N. Geva-Zatorsky and L. E. Comstock (2016). "Bacteroides fragilis type VI secretion systems use novel effector and immunity proteins to antagonize human gut Bacteroidales species." Proc Natl Acad Sci U S A 113(13): 3627– 3632.

      Feng, L., A. S. Raman, M. C. Hibberd, J. Cheng, N. W. Griffin, Y. Peng, S. A. Leyn, D. A. Rodionov, A. L. Osterman and J. I. Gordon (2020). "Identifying determinants of bacterial fitness in a model of human gut microbial succession." Proc Natl Acad Sci U S A 117(5): 2622-2633.

      Garcia-Bayona, L., M. J. Coyne and L. E. Comstock (2021). "Mobile Type VI secretion system loci of the gut Bacteroidales display extensive intra-ecosystem transfer, multispecies spread and geographical clustering." PLoS Genet 17(4): e1009541.

      Hood, R. D., P. Singh, F. Hsu, T. Guvener, M. A. Carl, R. R. Trinidad, J. M. Silverman, B. B. Ohlson, K. G. Hicks, R. L. Plemel, M. Li, S. Schwarz, W. Y. Wang, A. J. Merz, D. R. Goodlett and J. D. Mougous (2010). "A type VI secretion system of Pseudomonas aeruginosa targets a toxin to bacteria." Cell Host Microbe 7(1): 25–37.

      Park, S. Y., C. Rao, K. Z. Coyte, G. A. Kuziel, Y. Zhang, W. Huang, E. A. Franzosa, J. K. Weng, C. Huttenhower and S. Rakoff-Nahoum (2022). "Strain-level fitness in the gut microbiome is an emergent property of glycans and a single metabolite." Cell 185(3): 513-529 e521.

      Russell, A. B., A. G. Wexler, B. N. Harding, J. C. Whitney, A. J. Bohn, Y. A. Goo, B. Q. Tran, N. A. Barry, H. Zheng, S. B. Peterson, S. Chou, T. Gonen, D. R. Goodlett, A. L. Goodman and J. D. Mougous (2014). "A type VI secretion-related pathway in Bacteroidetes mediates interbacterial antagonism." Cell Host Microbe 16(2): 227–236.

      Segura Munoz, R. R., S. Mantz, I. Martinez, F. Li, R. J. Schmaltz, N. A. Pudlo, K. Urs, E. C. Martens, J. Walter and A. E. Ramer-Tait (2022). "Experimental evaluation of ecological principles to understand and modulate the outcome of bacterial strain competition in gut microbiomes." ISME J 16(6): 1594-1604.

      Souza, D. P., G. U. Oka, C. E. Alvarez-Martinez, A. W. Bisson-Filho, G. Dunger, L. Hobeika, N. S. Cavalcante, M. C. Alegria, L. R. Barbosa, R. K. Salinas, C. R. Guzzo and C. S. Farah (2015). "Bacterial killing via a type IV secretion system." Nat Commun 6: 6453.

      Wexler, A. G., Y. Bao, J. C. Whitney, L. M. Bobay, J. B. Xavier, W. B. Schofield, N. A. Barry, A. B. Russell, B. Q. Tran, Y. A. Goo, D. R. Goodlett, H. Ochman, J. D. Mougous and A. L. Goodman (2016). "Human symbionts inject and neutralize antibacterial toxins to persist in the gut." Proc Natl Acad Sci 113(13): 3639–3644.

      Whitney, J. C., S. B. Peterson, J. Kim, M. Pazos, A. J. Verster, M. C. Radey, H. D. Kulasekara, M. Q. Ching, N. P. Bullen, D. Bryant, Y. A. Goo, M. G. Surette, E. Borenstein, W. Vollmer and J. D. Mougous (2017). "A broadly distributed toxin family mediates contact-dependent antagonism between gram-positive bacteria." Elife 6(Jul 11): e26938.

    1. eLife Assessment

      This study presents a valuable insight into the mechanism of action of With-No-lysine (K) kinases (WNK) in the insulin-dependent signaling in the context of learning and memory. The study uses solid methodological approaches that align with the current standards in the field. The revised manuscript, which includes diverse models ranging from mouse to cell systems assessed by a broad range of complementary techniques, convincingly argues that WNK signaling contributes to neuronal function in insulin-sensitive brain regions such as the hippocampus through its effects on GLUT4 trafficking.

    2. Reviewer #1 (Public review):

      [Editor's Note: this version has been assessed by the Reviewing Editor without further input from the original reviewers. When experimentally feasible, the authors have adequately addressed the concerns of the reviewers in the revised manuscript to support the conclusions of the study.

      Summary:

      The study by Akita B. Jaykumar et al. explored an interesting and relevant hypothesis whether serine/threonine With-No-lysine (K) kinases (WNK)-1, -2, -3, and -4 engage in insulin-dependent glucose transporter-4 (GLUT4) signaling in the murine central nervous system. The authors especially focused on the hippocampus as this brain region exhibits high expression of insulin and GLUT4. Additionally, disrupted glucose metabolism in the hippocampus has been associated with anxiety disorders, while impaired WNK signaling has been linked to hypertension, learning disabilities, psychiatric disorders or Alzheimer's disease. The study took advantage of selective pan-WNK inhibitor WNK 643 as the main tool to manipulate WNK 1-4 activity both in vivo by daily, per-oral drug administration to wild-type mice, and in vitro by treating either adult murine brain synaptosomes, hippocampal slices, primary cortical cultures, and human cell lines (HEK293, SH-SY5Y). Using a battery of standard behavior paradigms such as open field test, elevated plus maze test, and fear conditioning, the authors convincingly demonstrate that the inhibition of WNK1-4 results in behavior changes, especially in enhanced learning and memory of WNK643-treated mice. To shed light on the underlying molecular mechanism, the authors implemented multiple biochemical approaches including immunoprecipitation, glucose-uptake assay, surface biotylination assay, immunoblotting, and immunofluorescence. The data suggest that simultaneous insulin stimulation and WNK1-4 inhibition results in increased glucose uptake and the activity of insulin's downstream effectors, phosphorylated Akt and phosphorylated AS160. Moreover, the authors demonstrate that insulin treatment enhances the physical interaction of the WNK effector OSR1/SPAK with Akt substrate AS160. As a result, combined treatment with insulin and the WNK643 inhibitor synergistically increases the targeting of GLUT4 to the plasma membrane. Collectively, these data strongly support the initial hypothesis that neuronal insulin- and WNK-dependent pathways do interact and engage in cognitive functions.

      In response to our initial comments, the authors mildly revised the manuscript, which did not improve the weaknesses to a sufficient level. Our follow-up comments are labeled under "Revisions 1".

      Strengths:

      The insulin-dependent signaling in the central nervous system is relatively understudied. This explorative study delves into several interesting and clinically relevant possibilities, examining how insulin-dependent signaling and its crosstalk with WNK kinases might affect brain circuits involved in memory formation and/or anxiety. Therefore, these findings might inspire follow-up studies performed in disease models for disorders that exhibit impaired glucose metabolism, deficient memory, or anxiety, such as Diabetes mellitus, Alzheimer's disease, or most of psychiatric disorders.

      The graphical presentation of the figures is of high quality, which helps the reader to obtain a good overview and to easily understand the experimental design, results, and conclusions.

      The behavioral studies are well conducted and provide valuable insights into the role of WNK kinases in glucose metabolism and their effect on learning and memory. Additionally, the authors evaluate the levels of basal and induced anxiety in Figures 1 and 2, enhancing our understanding of how WNK signaling might engage in cognitive function and anxiety-like behavior, particularly in the context of altered glucose metabolism.

      The data presented in Figures 3 and 4 are notably valuable and robust. The authors effectively utilize a variety of in vivo and in vitro models, combining different treatments in a clear manner. The experimental design is well-controlled, efficiently communicated, and well-executed, providing the reader with clear objectives and conclusions. Overall, these data represent particularly solid and reproducible evidence on the enhanced glucose uptake, GLUT4 targeting, and downstream effectors' activation upon insulin and WNK/OSR1 signaling crosstalk.

      Weaknesses:

      (1) The study used a WNK643 inhibitor as the only tool to manipulate WNK1-4 activity. This inhibitor seems selective; however, it has been reported that it exhibits different efficiency in inhibiting the individual WNK kinases among each other (e.g. PMID: 31017050, PMID: 36712947). Additionally, the authors do not analyze nor report the expression profiles or activity levels of WNK1, WNK2, WNK3, and WNK4 within the relevant brain regions (i.e. hippocampus, cortex, amygdala). Combined, these weaknesses raise concerns about the direct involvement of WNK kinases within the selected brain regions and behavior circuits. It would be beneficial if the authors provided gene profiling for WNK1, 2, 3, and -4 (e.g. using Allen brain atlas). To confirm the observations, the authors should either add results from using other WNK inhibitors or, preferentially, analyze knock-down or knock-out animals/tissue targeting the single kinases.

      Revisions 1: The authors added Fig. S1A during the revisions to show expression of Wnt1-4. While the expression data from humans is interesting, the experimental part of the study is performed in mice. It would be more informative for the authors to add expression profiles from mice or overview the expression pattern with suitable references in the introduction to address this point. The authors did not add data from knock down or knockout tissue targeting the single kinases.

      (2) The authors do not report any data on whether the global inhibition of WNKs affects insulin levels as such. Since the authors demonstrate the synergistic effect of simultaneous insulin treatment and WNK1-4 inhibition, such data are missing.

      Revisions 1: The authors added Fig. S5A to address this point. It is appreciated that authors performed the needed experiment. Unfortunately, no significant change was found, therefore, the authors still cannot conclude that they demonstrate a synergistic effect of simultaneous insulin treatment and WNT1-4 inhibition. It is a missed opportunity that the authors did not measure insulin in the CSF or tissue lysate to support the data.

      (3) The study discovered that the Sortilin receptor binds to OSR1, leading the authors to speculate that Sortilin may be involved in the insulin-dependent GLUT4 surface trafficking. The authors conclude in the result section that "WNK/OSR1/SPAK influences insulin-sensitive GLUT4 trafficking by balancing GLUT4 sequestration in the TGN via regulation of Sortilin with GLUT4 release from these vesicles upon insulin stimulation via regulation of AS160." However, the authors do not provide any evidence supporting Sortilin's involvement in such regulation, thus, this conclusion should be removed from the section. Accordingly, the first paragraph of the discussion should be also rephrased or removed.

      Revisions 1: The authors added Fig. 5M-N to address this point. The new experiment is appreciated. However, the authors still do not show that sortilin is involved in insulin or WNK-dependent GLUT4 trafficking in their set up since the authors do not demonstrate any changes in GLUT4 sorting or binding. The conclusions should therefore be rephrased or included purely in the discussion. Moreover, the discussion was not adjusted either, leading to over interpretation based on the available data.

      (4) The background relevant to Figure 5, as well as the results and conclusions presented in Figure 5 are quite challenging to follow due to the lack of a clear introduction to the signaling pathways. Consequently, understanding the conclusions drawn from the data is also difficult. It would be beneficial if the authors addressed this issue with either reformulations or additional sections in the introduction. Furthermore, the pulldown experiments in this figure lack some of the necessary controls.

      Revisions 1: The Authors insufficiently addressed this point during the revisions and did not rewrite the introduction as suggested.

      (5) The authors lack proper independent loading controls (e.g. GAPDH levels) in their immunoblots throughout the paper, and thus their quantifications lack this important normalization step. The authors also did not add knock-out or knock-down controls in their co-IPs. This is disappointing since these improvements were central and suggested during the revision process.

      (6) The schemes that represent only hypotheses (Fig. 1K, 4A) are unnecessary and confusing and thus should be omitted or placed at the end of each figure if the conclusions align.

      (7) Low-quality images, such as Fig. 5H should be replaced with high-resolution photos, moved to the supplementary, or omitted.

    3. Reviewer #2 (Public review):

      This study by Jaykumar and colleagues seeks to expand the field's appreciation of insulin responses in the brain, specifically by implicating WNK kinase function in various neuronal responses, ranging from behavioral / memory changes to GLUT4 trafficking to the cell surface with subsequent glucose uptake. This revised study is now comprehensive and presents a logical and reasonably documented cascade of molecular interactions responsible in part for GLUT4 trafficking under the regulation of WKK and insulin. Additional data allow the authors to dissect a plausible WNK/OSR1/SPAK-sortilin pathway for the modulation of GLUT4 trafficking, in part by capitalizing on an overlay of various techniques and systems. The data - much of it in vivo or ex vivo - showing a potential role for WNK function in brain glucose utilization remains a compelling part of the story, with the dissection of the signaling cascade and a potential role for sortilin in mediating WNK function via effects on GLUT4 cellular localization now more convincing.

      Initially, the group shows that oral WNK463 treatment - an inhibitor of WNKs broadly - in mice augments a number of memory readouts. These findings fit within the context of the overall story the authors present: that WNK function is critical to brain glucose utilization, which impacts learning. Multiple approaches are used to show that WNK463 treatment, i.e. inhibition of WNKs, increases glucose uptake, including labeled 2-deoxyglucose uptake in vivo in the brain and in isolated synaptosome, and uptake in ex vivo hippocampal slices. These findings are solid and consistent. With the exception of some relatively minor comments regarding the data presentation made to the authors and now fully addressed, the findings showing that WNK463 treatment increases GLUT4-mediated glucose uptake and surface localization of GLUT4 are reasonable, with the hippocampal slice data being particularly relevant.

      While the details of the WNK signaling cascade is dense, in the revised application one clearly appreciates the molecular interrogation and interactions the group is dissecting, supported by the use of multiple models. With the additional findings, these systems and the data now reinforce each other, presenting a strongly documented overall story.

      A limitation of the study with the initial submission was the authors' reliance upon a single pharmacological tool (WNK463) to inhibit WNK kinases. WNK463 apparently has substantial specificity for WNKs and WNK463 treatment lessened OSR1 phosphorylation (a WNK substrate). Nevertheless, the cohesiveness of the findings in terms of the broader pathway engagement (GLUT4 trafficking, glucose uptake) is consistent with the author's proposed mechanisms and conclusions. The authors have additionally addressed this concern in the revised manuscript with more information supporting the specificity of WNK463 as well as the multiple approaches to confirm the effect of WNK463 on the WNK signaling pathway of interest.

      The final few paragraphs of the discussion that weave the author's findings into the field more broadly, including Sortilin function and neurological disorders, are appreciated. Additional clarity in the Methods section is also helpful.

    4. Author response:

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #1 (Public review):

      (1) The study used a WNK643 inhibitor as the only tool to manipulate WNK1-4 activity. This inhibitor seems selective; however, it has been reported that it exhibits different efficiency in inhibiting the individual WNK kinases among each other (e.g. PMID: 31017050, PMID: 36712947). Additionally, the authors do not analyze nor report the expression profiles or activity levels of WNK1, WNK2, WNK3, and WNK4 within the relevant brain regions (i.e. hippocampus, cortex, amygdala). Combined, these weaknesses raise concerns about the direct involvement of WNK kinases within the selected brain regions and behavior circuits. It would be beneficial if the authors provided gene profiling for WNK1, 2, 3, and -4 (e.g. using Allen brain atlas). To confirm the observations, the authors should either add results from using other WNK inhibitors or, preferentially, analyze knock-down or knock-out animals/tissue targeting the single kinases.

      Revisions 1: The authors added Fig. S1A during the revisions to show expression of Wnk1-4. While the expression data from humans is interesting, the experimental part of the study is performed in mice. It would be more informative for the authors to add expression profiles from mice or overview the expression pattern with suitable references in the introduction to address this point. The authors did not add data from knock down or knockout tissue targeting the single kinases.

      Thank you for the excellent suggestion. We have added mouse in situ hybridization data curated from Allen Brain Atlas and found mRNA encoding WNK1 and WNK2 highly expressed in the hippocampus compared to WNK3 and WNK4. We also have included WNK1 knockdown data from cell lines (Figure S7A-F).

      Whole body WNK1 knockout is embryonically lethal, and we do not have access to brain tissue specific WNK knockout animal models. In addition, knockout of WNKs from brain tissue samples from animals is not very efficient from our experience and therefore, we included data from cell lines. In other non-neuronal cell lines, WNK1 knockdown replicates the effect of WNK463 (Figure S7A-D). However, in SHSY5Y cells, WNK1 knockdown did not replicate the effects of WNK463 on pAKT levels (Figure S7EF). This suggests tissue-specific effects of WNKs, and it also supports our suggestion that cooperativity among WNK family members is required in neuronal cells. This further supports our conclusion that WNK463 is an ideal tool to test our hypothesis in this study as it targets all 4 WNKs (WNK1-4) and furthermore, WNK463 was reported in the literature to inhibit only the four WNKs out of more than 400 kinases tested, indicating more selectivity than many small molecules used to target other enzymes.

      (2) The authors do not report any data on whether the global inhibition of WNKs affects insulin levels as such. Since the authors demonstrate the synergistic effect of simultaneous insulin treatment and WNK1-4 inhibition, such data are missing.

      Revisions 1: The authors added Fig. S5A to address this point. It is appreciated that authors performed the needed experiment. Unfortunately, no significant change was found, therefore, the authors still cannot conclude that they demonstrate a synergistic effect of simultaneous insulin treatment and WNT1-4 inhibition. It is a missed opportunity that the authors did not measure insulin in the CSF or tissue lysate to support the data.

      Thank you for the comment. As suggested, we tried to measure insulin in mouse hippocampal tissue lysate, and the levels fell way below the detectable range (78 - 5000 pg/mL) of the mouse insulin detection kit (Abcam: AB285341) used.

      (3) The study discovered that the Sortilin receptor binds to OSR1, leading the authors to speculate that Sortilin may be involved in the insulin-dependent GLUT4 surface trafficking. The authors conclude in the result section that "WNK/OSR1/SPAK influences insulin-sensitive GLUT4 trafficking by balancing GLUT4 sequestration in the TGN via regulation of Sortilin with GLUT4 release from these vesicles upon insulin stimulation via regulation of AS160." However, the authors do not provide any evidence supporting Sortilin's involvement in such regulation, thus, this conclusion should be removed from the section. Accordingly, the first paragraph of the discussion should be also rephrased or removed.

      Revisions 1: The authors added Fig. 5M-N to address this point. The new experiment is appreciated. However, the authors still do not show that sortilin is involved in insulin or WNK-dependent GLUT4 trafficking in their set up since the authors do not demonstrate any changes in GLUT4 sorting or binding. The conclusions should therefore be rephrased or included purely in the discussion. Moreover, the discussion was not adjusted either, leading to over interpretation based on the available data.

      Thank you for the suggestion. The conclusion has been rephrased as suggested.

      (4) The background relevant to Figure 5, as well as the results and conclusions presented in Figure 5 are quite challenging to follow due to the lack of a clear introduction to the signaling pathways. Consequently, understanding the conclusions drawn from the data is also difficult. It would be beneficial if the authors addressed this issue with either reformulations or additional sections in the introduction. Furthermore, the pulldown experiments in this figure lack some of the necessary controls.

      Revisions 1: The Authors insufficiently addressed this point during the revisions and did not rewrite the introduction as suggested.

      The background information related to figure 5 has been simplified as suggested. Response regarding the controls used is provided in the response to critique 5 as below.

      (5) The authors lack proper independent loading controls (e.g. GAPDH levels) in their immunoblots throughout the paper, and thus their quantifications lack this important normalization step. The authors also did not add knock-out or knock-down controls in their co-IPs. This is disappointing since these improvements were central and suggested during the revision process.

      GAPDH has been used as a loading control wherever applicable for Western blots on lysates (see Figures: 5D, 3F, 4E, 4C) In other cases, such as in Figure 5E, the analysis of pAS160 is normalized to total AS160 as this is more appropriate compared to GAPDH. For IP experiments such as Figure 5G, GAPDH is not an applicable control as we are using purified protein fragments in this case. For IP experiments (Figure 5K, 5L, 5C), IP proteins have been normalized to the input protein levels serving as a loading control for the IP because GAPDH is an intracellular protein which necessarily is not pulled down along with the proteins being IP’ed. Therefore, in this case, GAPDH is not a valid loading control. The choice of our loading controls used are very well supported by previous publications from our lab and other labs working on WNK pathways.

      (6) The schemes that represent only hypotheses (Fig. 1K, 4A) are unnecessary and confusing and thus should be omitted or placed at the end of each figure if the conclusions align.

      Thank you for the suggestion. Figure 1K is already at the end of the figure 1 and it shows the conclusion of that figure. Figure 4A have been placed at the end of the figures as suggested. Other schemes are only added at the end of the figures as suggested.

      (7) Low-quality images, such as Fig. 5H should be replaced with high-resolution photos, moved to the supplementary, or omitted.

      Thank you for your comment. The suggested images have been replaced with higher resolution ones.

      Reviewer #2 (Public review):

      This study by Jaykumar and colleagues seeks to expand the field's appreciation of insulin responses in the brain, specifically by implicating WNK kinase function in various neuronal responses, ranging from behavioral / memory changes to GLUT4 trafficking to the cell surface with subsequent glucose uptake. This revised study is now comprehensive and presents a logical and reasonably documented cascade of molecular interactions responsible in part for GLUT4 trafficking under the regulation of WKK and insulin. Additional data allow the authors to dissect a plausible WNK/OSR1/SPAK-sortilin pathway for the modulation of GLUT4 trafficking, in part by capitalizing on a overlay of various techniques and systems. The data - much of it in vivo or ex vivo - showing a potential role for WNK function in brain glucose utilization remains a compelling part of the story, with the dissection of the signaling cascade and a potential role for sortilin in mediating WNK function via effects on GLUT4 cellular localization now more convincing.

      Initially, the group shows that oral WNK463 treatment - an inhibitor of WNKs broadly - in mice augments a number of memory readouts. These findings fit within the context of the overall story the authors present: that WNK function is critical to brain glucose utilization, which impacts learning. Multiple approaches are used to show that WNK463 treatment, i.e. inhibition of WNKs, increases glucose uptake, including labeled 2deoxyglucose uptake in vivo in the brain and in isolated synaptosome, and uptake in ex vivo hippocampal slices. These findings are solid and consistent. With the exception of some relatively minor comments regarding the data presentation made to the authors and now fully addressed, the findings showing that WNK463 treatment increases GLUT4-mediated glucose uptake and surface localization of GLUT4 are reasonable, with the hippocampal slice data being particularly relevant.

      While the details of the WNK signaling cascade is dense, in the revised application one clearly appreciates the molecular interrogation and interactions the group is dissecting, supported by the use of multiple models. With the additional findings, these systems and the data now reinforce each other, presenting a strongly documented overall story.

      A limitation of the study with the initial submission was the authors' reliance upon a single pharmacological tool (WNK463) to inhibit WNK kinases. WNK463 apparently has substantial specificity for WNKs and WNK463 treatment lessened OSR1 phosphorylation (a WNK substrate). Nevertheless, the cohesiveness of the findings in terms of the broader pathway engagement (GLUT4 trafficking, glucose uptake) is consistent with the author's proposed mechanisms and conclusions. The authors have additionally addressed this concern in the revised manuscript with more information supporting the specificity of WNK463 as well as the multiple approaches to confirm the effect of WNK463 on the WNK signaling pathway of interest.

      The final few paragraphs of the discussion that weave the author's findings into the field more broadly, including Sortilin function and neurological disorders, are appreciated. Additional clarity in the Methods section is also helpful.

      Thank you for the positive response and acknowledging that we have satisfactorily addressed all of your critiques.

    1. eLife Assessment

      This manuscript presents a valuable computational tool for identifying 3-5 gene regulatory network topologies capable of generating oscillatory dynamics. The application of Monte Carlo Tree Search to circuit design is novel and effectively expands the scale at which non-linear behaviours can be explored in silico. The efficiency of the proposed algorithm is convincing, and the work will be of interest to the systems and synthetic biology communities. While the generality of the identified circuit properties is constrained by the simplifying modelling assumptions and parameter choices, the methodological contribution represents a significant advance in the field.